# Stage 4 Pressure Testing in Cangjie-Skill: How '诱饵' (Bait) Questions Validate Skill Triggers

> Validate skill triggers in Stage 4 pressure testing using bait questions. Expose false positives before deployment with blind sub-agent evaluation. Ensure accuracy and reliability.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: testing
- Published: 2026-08-14

---

**Stage 4 pressure testing validates that a skill's trigger fires only in correct situations using blind sub-agent evaluation and mandatory '诱饵' (bait) questions that expose false positives before deployment.**

The `kangarooking/cangjie-skill` repository implements a rigorous four-stage methodology for developing AI skills. Stage 4 serves as the final quality gate, treating each skill as a black-box and subjecting it to model-agnostic validation. This process ensures that triggers—defined in Stage 2 as the A2 (action) component—are precise enough for production use.

## The Four-Step Pressure Testing Workflow

The complete specification lives in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md). The workflow deliberately strips away any information that could bias the evaluator:

1. **Create [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json)** — A Darwin-compatible schema file containing all test cases for the skill.

2. **Execute blind sub-agent evaluations** — An isolated sub-agent receives only:
   - The skill path or description
   - The user prompt
   - An optional list of neighboring skills

   Critically, the sub-agent **does not see** the `type`, `expected_behavior`, or notes fields. It must independently output:
   - `would_trigger` — boolean judgment
   - `reason` — explanation for the decision
   - `if_triggered_action` — action it would take

3. **Compare outputs** against ground-truth expectations in the test file.

4. **Compute pass rate** with strict thresholds:
   - **100%** → Skill accepted
   - **≥80%** → Review failures to determine if trigger or tests need revision
   - **<80%** → Full rejection, return to Stage 2 for A2/E/B rework

## Why '诱饵' (Bait) Questions Are Mandatory

The specification enforces a hard rule: **skills without bait tests are automatically rejected** (see line 66 of [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)). These `should_not_trigger` cases serve two critical purposes.

### Prevent Over-Triggering

Positive-only testing creates dangerous blind spots. A skill scored against only `should_trigger` cases could achieve perfect marks while firing constantly in production. Bait questions expose vague trigger logic by presenting semantically similar but functionally distinct prompts.

### Cross-Skill Discrimination

At minimum, one bait case must target a prompt that **should trigger a different skill from the same book**. This validates that the model can distinguish between related capabilities and prevents the common deployment failure where multiple skills compete for identical user inputs.

## Example Test Configuration

The [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) structure demonstrates all three case types:

```json
{
  "skill": "inversion-thinking",
  "version": "0.1.0",
  "test_cases": [
    {
      "id": "should-trigger-01",
      "type": "should_trigger",
      "prompt": "我要决定要不要接这个新项目,列了一堆好处但还是没底",
      "expected_behavior": "调用 inversion-thinking, 反问'最不希望发生什么'",
      "notes": "正面场景: 决策纠结"
    },
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "帮我查一下这个 API 的参数",
      "expected_behavior": "纯信息查询, 不应调用任何决策 skill",
      "notes": "诱饵: 非决策场景"
    },
    {
      "id": "edge-01",
      "type": "edge_case",
      "prompt": "我在想晚饭吃什么",
      "expected_behavior": "日常琐事, 不应调用 (虽然字面是'决策')",
      "notes": "边界: 区分严肃决策和日常选择"
    }
  ]
}

```

*Source: Excerpt adapted from [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)*

## Implementing Blind Evaluation

The sub-agent evaluation enforces information hiding to prevent cheating:

```python
import json
from sub_agent import SubAgent  # isolated sub-agent class

def blind_test(skill_path, prompt, neighbors=None):
    """Run evaluation with no access to expected answers."""
    ctx = {
        "skill_path": skill_path,
        "prompt": prompt,
        "neighbor_skills": neighbors or []
    }
    # Sub-agent cannot see case type or expected_behavior

    return SubAgent.run(ctx)  # dict with would_trigger, reason, if_triggered_action

# Execution pipeline

with open("test-prompts.json") as f:
    data = json.load(f)

results = []
for case in data["test_cases"]:
    output = blind_test(data["skill"], case["prompt"])
    # Ground truth comparison happens outside sub-agent context

    results.append(validate(output, case))

```

## Pass Rate Arbitration

The final scoring logic implements the repository's strict quality bar:

```python
def arbitrate_skill(pass_rate):
    if pass_rate == 1.0:
        return "ACCEPTED"
    elif pass_rate >= 0.80:
        return "REVIEW_NEEDED"  # Manual triage of failures

    else:
        return "REJECT_AND_REWORK"  # Return to Stage 2

print(f"Skill status: {arbitrate_skill(passed / total)}")

```

## Key Source Files

| File | Purpose |
|------|---------|
| [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) | Complete specification including bait requirements and rejection criteria |
| `templates/test-prompts.json.template` | Starter template for skill test files |
| [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) | Index of all skills with methodology references |
| `extractors/*-extractor.md` | Definitions producing skill descriptions for sub-agent consumption |

## Summary

- Stage 4 pressure testing is the **final validation gate** before skill deployment in the cangjie-skill methodology.
- **Blind sub-agent evaluation** ensures model-agnostic testing by hiding expected answers from the evaluator.
- **'诱饵' (bait) questions** are non-negotiable—skills lacking `should_not_trigger` cases face automatic rejection.
- The **80% pass threshold** creates a clear decision boundary between acceptable variance and fundamental trigger flaws.

## Frequently Asked Questions

### Why does the sub-agent evaluation need to be "blind"?

The sub-agent must not see the `type` or `expected_behavior` fields because any hint about whether a skill should trigger would trivialize the test. Blind evaluation forces the sub-agent to make independent judgments based solely on the skill description and user prompt, mirroring how the skill will actually be used in production.

### What happens if a skill scores exactly 79% on the pressure test?

According to [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md), any score below 80% triggers **REJECT_AND_REWORK** status. The skill must return to Stage 2 for revision of its trigger logic (A2), entry conditions (E), or behavior definition (B). This hard cutoff prevents marginally functional skills from shipping.

### How many bait questions are required per skill?

The specification mandates **at least one** `should_not_trigger` case, with an additional requirement that one bait must specifically test cross-skill discrimination. Developers typically include 2-4 bait cases covering different failure modes: unrelated domains, neighboring skill overlaps, and edge-case semantic similarities.

### Can I reuse bait questions across multiple skills?

No—bait questions must be **skill-specific**. A prompt that correctly avoids triggering Skill A might legitimately trigger Skill B. Each skill's [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) needs tailored bait cases that genuinely test its unique trigger boundaries, particularly for skills within the same book that share thematic overlap.