Stage 4 Pressure Testing in Cangjie-Skill: How '诱饵' (Bait) Questions Validate Skill Triggers

Stage 4 pressure testing validates that a skill's trigger fires only in correct situations using blind sub-agent evaluation and mandatory '诱饵' (bait) questions that expose false positives before deployment.

The kangarooking/cangjie-skill repository implements a rigorous four-stage methodology for developing AI skills. Stage 4 serves as the final quality gate, treating each skill as a black-box and subjecting it to model-agnostic validation. This process ensures that triggers—defined in Stage 2 as the A2 (action) component—are precise enough for production use.

The Four-Step Pressure Testing Workflow

The complete specification lives in methodology/06-stage4-pressure-test.md. The workflow deliberately strips away any information that could bias the evaluator:

  1. Create test-prompts.json — A Darwin-compatible schema file containing all test cases for the skill.

  2. Execute blind sub-agent evaluations — An isolated sub-agent receives only:

    • The skill path or description
    • The user prompt
    • An optional list of neighboring skills

    Critically, the sub-agent does not see the type, expected_behavior, or notes fields. It must independently output:

    • would_trigger — boolean judgment
    • reason — explanation for the decision
    • if_triggered_action — action it would take
  3. Compare outputs against ground-truth expectations in the test file.

  4. Compute pass rate with strict thresholds:

    • 100% → Skill accepted
    • ≥80% → Review failures to determine if trigger or tests need revision
    • <80% → Full rejection, return to Stage 2 for A2/E/B rework

Why '诱饵' (Bait) Questions Are Mandatory

The specification enforces a hard rule: skills without bait tests are automatically rejected (see line 66 of methodology/06-stage4-pressure-test.md). These should_not_trigger cases serve two critical purposes.

Prevent Over-Triggering

Positive-only testing creates dangerous blind spots. A skill scored against only should_trigger cases could achieve perfect marks while firing constantly in production. Bait questions expose vague trigger logic by presenting semantically similar but functionally distinct prompts.

Cross-Skill Discrimination

At minimum, one bait case must target a prompt that should trigger a different skill from the same book. This validates that the model can distinguish between related capabilities and prevents the common deployment failure where multiple skills compete for identical user inputs.

Example Test Configuration

The test-prompts.json structure demonstrates all three case types:

{
  "skill": "inversion-thinking",
  "version": "0.1.0",
  "test_cases": [
    {
      "id": "should-trigger-01",
      "type": "should_trigger",
      "prompt": "我要决定要不要接这个新项目,列了一堆好处但还是没底",
      "expected_behavior": "调用 inversion-thinking, 反问'最不希望发生什么'",
      "notes": "正面场景: 决策纠结"
    },
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "帮我查一下这个 API 的参数",
      "expected_behavior": "纯信息查询, 不应调用任何决策 skill",
      "notes": "诱饵: 非决策场景"
    },
    {
      "id": "edge-01",
      "type": "edge_case",
      "prompt": "我在想晚饭吃什么",
      "expected_behavior": "日常琐事, 不应调用 (虽然字面是'决策')",
      "notes": "边界: 区分严肃决策和日常选择"
    }
  ]
}

Source: Excerpt adapted from methodology/06-stage4-pressure-test.md

Implementing Blind Evaluation

The sub-agent evaluation enforces information hiding to prevent cheating:

import json
from sub_agent import SubAgent  # isolated sub-agent class

def blind_test(skill_path, prompt, neighbors=None):
    """Run evaluation with no access to expected answers."""
    ctx = {
        "skill_path": skill_path,
        "prompt": prompt,
        "neighbor_skills": neighbors or []
    }
    # Sub-agent cannot see case type or expected_behavior

    return SubAgent.run(ctx)  # dict with would_trigger, reason, if_triggered_action

# Execution pipeline

with open("test-prompts.json") as f:
    data = json.load(f)

results = []
for case in data["test_cases"]:
    output = blind_test(data["skill"], case["prompt"])
    # Ground truth comparison happens outside sub-agent context

    results.append(validate(output, case))

Pass Rate Arbitration

The final scoring logic implements the repository's strict quality bar:

def arbitrate_skill(pass_rate):
    if pass_rate == 1.0:
        return "ACCEPTED"
    elif pass_rate >= 0.80:
        return "REVIEW_NEEDED"  # Manual triage of failures

    else:
        return "REJECT_AND_REWORK"  # Return to Stage 2

print(f"Skill status: {arbitrate_skill(passed / total)}")

Key Source Files

File Purpose
methodology/06-stage4-pressure-test.md Complete specification including bait requirements and rejection criteria
templates/test-prompts.json.template Starter template for skill test files
SKILL.md Index of all skills with methodology references
extractors/*-extractor.md Definitions producing skill descriptions for sub-agent consumption

Summary

  • Stage 4 pressure testing is the final validation gate before skill deployment in the cangjie-skill methodology.
  • Blind sub-agent evaluation ensures model-agnostic testing by hiding expected answers from the evaluator.
  • '诱饵' (bait) questions are non-negotiable—skills lacking should_not_trigger cases face automatic rejection.
  • The 80% pass threshold creates a clear decision boundary between acceptable variance and fundamental trigger flaws.

Frequently Asked Questions

Why does the sub-agent evaluation need to be "blind"?

The sub-agent must not see the type or expected_behavior fields because any hint about whether a skill should trigger would trivialize the test. Blind evaluation forces the sub-agent to make independent judgments based solely on the skill description and user prompt, mirroring how the skill will actually be used in production.

What happens if a skill scores exactly 79% on the pressure test?

According to methodology/06-stage4-pressure-test.md, any score below 80% triggers REJECT_AND_REWORK status. The skill must return to Stage 2 for revision of its trigger logic (A2), entry conditions (E), or behavior definition (B). This hard cutoff prevents marginally functional skills from shipping.

How many bait questions are required per skill?

The specification mandates at least one should_not_trigger case, with an additional requirement that one bait must specifically test cross-skill discrimination. Developers typically include 2-4 bait cases covering different failure modes: unrelated domains, neighboring skill overlaps, and edge-case semantic similarities.

Can I reuse bait questions across multiple skills?

No—bait questions must be skill-specific. A prompt that correctly avoids triggering Skill A might legitimately trigger Skill B. Each skill's test-prompts.json needs tailored bait cases that genuinely test its unique trigger boundaries, particularly for skills within the same book that share thematic overlap.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →