# How Pressure Testing with Trap Questions Validates cangjie-skill Quality

> Validate cangjie-skill quality with pressure testing. Learn how trap questions ensure trigger accuracy and output quality through blind, automated evaluation for seamless acceptance.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: testing
- Published: 2026-07-19

---

**Pressure testing with trap questions is the decisive fourth stage of the cangjie-skill pipeline that validates trigger accuracy and output quality by submitting skills to blind, automated evaluation against deliberately crafted prompts, requiring a 100% pass rate for final acceptance.**

The cangjie-skill framework (`kangarooking/cangjie-skill`) implements a rigorous four-stage delivery pipeline where the final quality gate subjects every skill to a battery of adversarial test cases. This stage ensures that the skill's trigger description (defined in the A2 section of [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)) is unambiguous and that the skill will not fire erroneously in production environments.

## The Three Categories of Trap Questions

Pressure tests rely on [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json), a Darwin-compatible test suite that defines three mandatory categories of trap questions. These are defined in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md).

### should_trigger

These prompts represent legitimate scenarios where the skill **must** activate. For example, a prompt like "我在考虑是否要换工作，列了很多好处但还在犹豫" should trigger the `inversion-thinking` skill. The pressure test validates that the skill fires and produces the expected action.

### should_not_trigger

These "bait" prompts test **negative precision**—scenarios where the skill must remain dormant despite superficial similarities to valid triggers. A prompt such as "帮我查一下这个 API 的参数" should not trigger a decision-making skill, ensuring zero tolerance for false positives.

### edge_case

Edge cases probe **boundary handling** by presenting ambiguous inputs that sit between clear trigger and non-trigger territories. For instance, "我在想今晚吃什么" tests whether the skill can distinguish serious decision-making from trivial daily choices, validating the nuance of the trigger logic.

## Blind Sub-Agent Evaluation

Each prompt is fed to an **independent sub-agent** that simulates real-world deployment conditions. According to the methodology in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md), the sub-agent operates without knowledge of the expected behavior—the `type`, `expected_behavior`, and `notes` fields remain hidden.

The sub-agent must output three fields:
- `would_trigger` – boolean indicating if it would call the skill
- `reason` – the reasoning behind the decision
- `if_triggered_action` – the specific action it would take

This blind evaluation mirrors production scenarios where the agent sees only the skill description and the user's utterance, ensuring the test results reflect actual runtime behavior.

## Automated Scoring and Pass Thresholds

The main pipeline compares sub-agent responses against expectations using strict scoring logic defined in the "执行流程" section of [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md).

The validation criteria are:
- **should_trigger**: The skill must invoke and the action must match the expected behavior exactly
- **should_not_trigger**: The skill must not invoke under any circumstances
- **edge_case**: The decision must align with the defined boundary rationale

The acceptance thresholds create an uncompromising feedback loop:
- **100% pass**: The skill is accepted for release
- **≥ 80% pass**: Failing cases undergo analysis to determine if the trigger description (A2) or the test itself requires refinement
- **< 80% pass**: The skill is sent back to **stage 2** for a full redesign of its trigger logic

## Cross-Skill Confusion Testing

The methodology mandates **cross-skill confusion tests** to prevent "skill-cannibalisation" after deployment. At least one bait prompt must be a query that *should* trigger a *different* skill from the same book. This forces precise scope definition and ensures that skills with overlapping domains do not interfere with each other's activation logic, as specified in the "跨 skill 混淆测试" clause of [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md).

## Implementing Pressure Tests in Practice

Developers create and execute pressure tests using the repository's standardized tooling and templates.

### Creating the Test Suite

Use the template located at `templates/test-prompts.json.template` to structure your test cases:

```python
import json
import pathlib

test_cases = [
    {
        "id": "should-trigger-01",
        "type": "should_trigger",
        "prompt": "我在考虑是否要换工作，列了很多好处但还在犹豫",
        "expected_behavior": "调用 inversion-thinking, 反问‘最不希望发生什么’",
        "notes": "正面场景: 决策纠结"
    },
    {
        "id": "should-not-trigger-01",
        "type": "should_not_trigger",
        "prompt": "帮我查一下这个 API 的参数",
        "expected_behavior": "纯信息查询, 不应调用任何决策 skill",
        "notes": "诱饵: 非决策场景"
    },
    {
        "id": "edge-01",
        "type": "edge_case",
        "prompt": "我在想今晚吃什么",
        "expected_behavior": "日常琐事, 不应调用 (虽然字面是‘决策’)",
        "notes": "边界: 区分严肃决策和日常选择"
    }
]

path = pathlib.Path("books/my-book/inversion-thinking/test-prompts.json")
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(json.dumps(
    {"skill": "inversion-thinking", "version": "0.1.0", "test_cases": test_cases},
    ensure_ascii=False,
    indent=2
))

```

### Running the Pressure Test

Execute the pressure test using the CLI tool provided in the repository:

```bash
python -m cangjie_skill.run_pressure_test books/my-book/inversion-thinking

```

This script spawns a clean sub-agent for each prompt, hides the expected outcomes to ensure blind evaluation, compares the results against [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json), and generates [`test-results.md`](https://github.com/kangarooking/cangjie-skill/blob/main/test-results.md) alongside the skill's [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md).

### Interpreting Results

The generated [`test-results.md`](https://github.com/kangarooking/cangjie-skill/blob/main/test-results.md) provides a pass-rate summary:

```markdown

## Inversion-thinking – Pressure Test Results

- Total prompts: 3
- Passed: 3 / 3 (100%)
- Verdict: ✅ Skill accepted

```

## Summary

- **Pressure testing** is the fourth and final stage of the cangjie-skill pipeline, serving as the decisive quality gate before release.
- **Trap questions** fall into three categories (`should_trigger`, `should_not_trigger`, `edge_case`) defined in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md).
- **Blind sub-agent evaluation** ensures tests simulate real-world conditions where the agent has no prior knowledge of expected outcomes.
- **Pass thresholds** are strict: 100% for acceptance, ≥80% for potential refinement, and <80% mandates a return to stage 2 for redesign.
- **Cross-skill confusion tests** prevent skills from cannibalizing each other's triggers by testing against related skills in the same book.
- **Artifacts** ([`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) and [`test-results.md`](https://github.com/kangarooking/cangjie-skill/blob/main/test-results.md)) provide Darwin-compatible test suites and automated reporting for continuous validation.

## Frequently Asked Questions

### What happens if a skill scores exactly 85% on the pressure test?

According to the methodology in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md), a score between 80% and 99% triggers a failure analysis workflow. The development team must examine the specific failing cases to determine whether the trigger description (A2) needs tightening or whether the test expectations themselves require adjustment. The skill does not pass automatically until it achieves 100% or the issues are resolved through refinement.

### How does the blind sub-agent evaluation prevent test contamination?

The sub-agent evaluation is designed to mirror production reality by hiding the `type`, `expected_behavior`, and `notes` fields from the evaluating agent. As implemented in the `cangjie_skill.run_pressure_test` module, the sub-agent sees only the skill description and the user prompt, exactly as a real user interaction would occur. This prevents the evaluator from gaming the test based on prior knowledge of expected outcomes.

### What is the purpose of cross-skill confusion testing?

Cross-skill confusion tests ensure that skills within the same book maintain distinct trigger boundaries. By including at least one bait prompt that *should* trigger a *different* skill from the same book, the test validates that the current skill does not over-capture queries intended for other skills. This "anti-cannibalisation" measure is mandatory according to the "跨 skill 混淆测试" clause and prevents deployment collisions where multiple skills compete for the same user utterances.

### Where are pressure test artifacts stored within the repository?

Each skill maintains its own pressure test artifacts within its book directory. The [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) file (containing the test suite) and [`test-results.md`](https://github.com/kangarooking/cangjie-skill/blob/main/test-results.md) (containing the evaluation report) live alongside the skill's [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) definition. This co-location ensures that quality assurance data travels with the skill through version control, as outlined in the directory layout specifications of the main [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) file.