How Pressure Testing with Trap Questions Validates cangjie-skill Quality

Pressure testing with trap questions is the decisive fourth stage of the cangjie-skill pipeline that validates trigger accuracy and output quality by submitting skills to blind, automated evaluation against deliberately crafted prompts, requiring a 100% pass rate for final acceptance.

The cangjie-skill framework (kangarooking/cangjie-skill) implements a rigorous four-stage delivery pipeline where the final quality gate subjects every skill to a battery of adversarial test cases. This stage ensures that the skill's trigger description (defined in the A2 section of SKILL.md) is unambiguous and that the skill will not fire erroneously in production environments.

The Three Categories of Trap Questions

Pressure tests rely on test-prompts.json, a Darwin-compatible test suite that defines three mandatory categories of trap questions. These are defined in methodology/06-stage4-pressure-test.md.

should_trigger

These prompts represent legitimate scenarios where the skill must activate. For example, a prompt like "我在考虑是否要换工作,列了很多好处但还在犹豫" should trigger the inversion-thinking skill. The pressure test validates that the skill fires and produces the expected action.

should_not_trigger

These "bait" prompts test negative precision—scenarios where the skill must remain dormant despite superficial similarities to valid triggers. A prompt such as "帮我查一下这个 API 的参数" should not trigger a decision-making skill, ensuring zero tolerance for false positives.

edge_case

Edge cases probe boundary handling by presenting ambiguous inputs that sit between clear trigger and non-trigger territories. For instance, "我在想今晚吃什么" tests whether the skill can distinguish serious decision-making from trivial daily choices, validating the nuance of the trigger logic.

Blind Sub-Agent Evaluation

Each prompt is fed to an independent sub-agent that simulates real-world deployment conditions. According to the methodology in methodology/06-stage4-pressure-test.md, the sub-agent operates without knowledge of the expected behavior—the type, expected_behavior, and notes fields remain hidden.

The sub-agent must output three fields:

  • would_trigger – boolean indicating if it would call the skill
  • reason – the reasoning behind the decision
  • if_triggered_action – the specific action it would take

This blind evaluation mirrors production scenarios where the agent sees only the skill description and the user's utterance, ensuring the test results reflect actual runtime behavior.

Automated Scoring and Pass Thresholds

The main pipeline compares sub-agent responses against expectations using strict scoring logic defined in the "执行流程" section of methodology/06-stage4-pressure-test.md.

The validation criteria are:

  • should_trigger: The skill must invoke and the action must match the expected behavior exactly
  • should_not_trigger: The skill must not invoke under any circumstances
  • edge_case: The decision must align with the defined boundary rationale

The acceptance thresholds create an uncompromising feedback loop:

  • 100% pass: The skill is accepted for release
  • ≥ 80% pass: Failing cases undergo analysis to determine if the trigger description (A2) or the test itself requires refinement
  • < 80% pass: The skill is sent back to stage 2 for a full redesign of its trigger logic

Cross-Skill Confusion Testing

The methodology mandates cross-skill confusion tests to prevent "skill-cannibalisation" after deployment. At least one bait prompt must be a query that should trigger a different skill from the same book. This forces precise scope definition and ensures that skills with overlapping domains do not interfere with each other's activation logic, as specified in the "跨 skill 混淆测试" clause of methodology/06-stage4-pressure-test.md.

Implementing Pressure Tests in Practice

Developers create and execute pressure tests using the repository's standardized tooling and templates.

Creating the Test Suite

Use the template located at templates/test-prompts.json.template to structure your test cases:

import json
import pathlib

test_cases = [
    {
        "id": "should-trigger-01",
        "type": "should_trigger",
        "prompt": "我在考虑是否要换工作,列了很多好处但还在犹豫",
        "expected_behavior": "调用 inversion-thinking, 反问‘最不希望发生什么’",
        "notes": "正面场景: 决策纠结"
    },
    {
        "id": "should-not-trigger-01",
        "type": "should_not_trigger",
        "prompt": "帮我查一下这个 API 的参数",
        "expected_behavior": "纯信息查询, 不应调用任何决策 skill",
        "notes": "诱饵: 非决策场景"
    },
    {
        "id": "edge-01",
        "type": "edge_case",
        "prompt": "我在想今晚吃什么",
        "expected_behavior": "日常琐事, 不应调用 (虽然字面是‘决策’)",
        "notes": "边界: 区分严肃决策和日常选择"
    }
]

path = pathlib.Path("books/my-book/inversion-thinking/test-prompts.json")
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(json.dumps(
    {"skill": "inversion-thinking", "version": "0.1.0", "test_cases": test_cases},
    ensure_ascii=False,
    indent=2
))

Running the Pressure Test

Execute the pressure test using the CLI tool provided in the repository:

python -m cangjie_skill.run_pressure_test books/my-book/inversion-thinking

This script spawns a clean sub-agent for each prompt, hides the expected outcomes to ensure blind evaluation, compares the results against test-prompts.json, and generates test-results.md alongside the skill's SKILL.md.

Interpreting Results

The generated test-results.md provides a pass-rate summary:


## Inversion-thinking – Pressure Test Results

- Total prompts: 3
- Passed: 3 / 3 (100%)
- Verdict: ✅ Skill accepted

Summary

  • Pressure testing is the fourth and final stage of the cangjie-skill pipeline, serving as the decisive quality gate before release.
  • Trap questions fall into three categories (should_trigger, should_not_trigger, edge_case) defined in methodology/06-stage4-pressure-test.md.
  • Blind sub-agent evaluation ensures tests simulate real-world conditions where the agent has no prior knowledge of expected outcomes.
  • Pass thresholds are strict: 100% for acceptance, ≥80% for potential refinement, and <80% mandates a return to stage 2 for redesign.
  • Cross-skill confusion tests prevent skills from cannibalizing each other's triggers by testing against related skills in the same book.
  • Artifacts (test-prompts.json and test-results.md) provide Darwin-compatible test suites and automated reporting for continuous validation.

Frequently Asked Questions

What happens if a skill scores exactly 85% on the pressure test?

According to the methodology in methodology/06-stage4-pressure-test.md, a score between 80% and 99% triggers a failure analysis workflow. The development team must examine the specific failing cases to determine whether the trigger description (A2) needs tightening or whether the test expectations themselves require adjustment. The skill does not pass automatically until it achieves 100% or the issues are resolved through refinement.

How does the blind sub-agent evaluation prevent test contamination?

The sub-agent evaluation is designed to mirror production reality by hiding the type, expected_behavior, and notes fields from the evaluating agent. As implemented in the cangjie_skill.run_pressure_test module, the sub-agent sees only the skill description and the user prompt, exactly as a real user interaction would occur. This prevents the evaluator from gaming the test based on prior knowledge of expected outcomes.

What is the purpose of cross-skill confusion testing?

Cross-skill confusion tests ensure that skills within the same book maintain distinct trigger boundaries. By including at least one bait prompt that should trigger a different skill from the same book, the test validates that the current skill does not over-capture queries intended for other skills. This "anti-cannibalisation" measure is mandatory according to the "跨 skill 混淆测试" clause and prevents deployment collisions where multiple skills compete for the same user utterances.

Where are pressure test artifacts stored within the repository?

Each skill maintains its own pressure test artifacts within its book directory. The test-prompts.json file (containing the test suite) and test-results.md (containing the evaluation report) live alongside the skill's SKILL.md definition. This co-location ensures that quality assurance data travels with the skill through version control, as outlined in the directory layout specifications of the main SKILL.md file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →