How Pressure Testing with Darwin-Skill Compatible test-prompts.json Works in Cangjie-Skill

Pressure testing validates that a skill's trigger logic fires only when appropriate by using a Darwin-compatible test-prompts.json file and blind sub-agent evaluation before delivery.

Pressure testing is the fourth stage of the cangjie-skill development pipeline. It ensures a skill's trigger logic (defined in the A2 field) behaves correctly under realistic conditions. The process depends on a Darwin-compatible test-prompts.json structure and independent sub-agents that evaluate test cases without prior knowledge of expected outcomes.

What Makes a test-prompts.json Darwin-Skill Compatible

The Darwin-compatible format follows a strict schema designed for blind evaluation. You create this file from the template at templates/test-prompts.json.template.

A valid test-prompts.json contains:

  • skill – the skill's slug identifier
  • version – semantic version of the test suite
  • test_cases – an array of evaluation scenarios, each with:
    • id – unique case identifier
    • type – one of should_trigger, should_not_trigger, or edge_case
    • prompt – the user utterance to test
    • expected_behavior – what correct behavior looks like
    • notes – context for human reviewers (hidden from sub-agents)

The schema is documented in methodology/06-stage4-pressure-test.md.

The Five Steps of Pressure Testing

1. Create the Test File

Fill the template with three categories of test cases:

  • Should-trigger cases – typical user utterances that must activate the skill
  • Should-not-trigger cases – "bait" prompts that test for false positives
  • Edge cases – ambiguous inputs where the decision requires nuanced reasoning

Include cross-skill prompts if sibling skills exist in the same book to prevent activation conflicts.

2. Run Blind Sub-Agent Evaluations

For each test_case, spin up a clean sub-agent with strictly limited information:

def evaluate_prompt(skill_path, prompt, sibling_skills=None):
    # Load skill description but hide trigger details

    skill_meta = load_skill_meta(skill_path, hide=['type','expected_behavior','notes'])
    # Run the sub-agent with only the prompt and optional sibling list

    result = sub_agent.run(
        skill_meta=skill_meta,
        user_prompt=prompt,
        sibling_skills=sibling_skills,
    )
    return {
        "would_trigger": result.would_trigger,
        "reason": result.reason,
        "if_triggered_action": result.action,
    }

Critical constraint: The sub-agent receives only:

  • The skill's path and content (metadata without evaluation fields)
  • The user prompt
  • Optional sibling skill list

The type, expected_behavior, and notes fields are intentionally withheld to eliminate bias.

3. Compare Results to Expected Behavior

The main workflow matches sub-agent output against test-prompts.json expectations:

Case Type Passing Criteria
should_trigger Sub-agent triggers skill and performs expected action
should_not_trigger Sub-agent does not trigger (zero tolerance for false positives)
edge_case Decision aligns with reasoning in expected_behavior

4. Calculate Pass Rate and Decide

  • 100% pass → Skill accepted for delivery
  • ≥80% pass → Investigate failures; decide whether to refine A2 trigger description or adjust test cases
  • <80% pass → Return to stage 2 for complete trigger logic redesign

5. Iterate Until Threshold Met

After any fix, re-run blind tests. Update test-results.md in the skill directory to track progress.

Example test-prompts.json Structure

{
  "skill": "inversion-thinking",
  "version": "0.1.0",
  "test_cases": [
    {
      "id": "should-trigger-01",
      "type": "should_trigger",
      "prompt": "我要决定要不要接这个新项目,列了一堆好处但还是没底",
      "expected_behavior": "调用 inversion‑thinking, 反问 '最不希望发生什么'",
      "notes": "正面场景: 决策纠结"
    },
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "帮我查一下这个 API 的参数",
      "expected_behavior": "纯信息查询, 不应调用任何决策 skill",
      "notes": "诱饵: 非决策场景"
    },
    {
      "id": "edge-01",
      "type": "edge_case",
      "prompt": "我在想晚饭吃什么",
      "expected_behavior": "日常琐事, 不应调用 (虽然字面是'决策')",
      "notes": "边界: 区分严肃决策和日常选择"
    }
  ]
}

This example demonstrates defense in depth: positive testing, negative testing, and boundary analysis in a single suite.

Why Blind Sub-Agent Evaluation Matters

The trigger (A2) is the most fragile component of any skill. Without pressure testing, a skill can appear functional in development while silently mis-triggering in production.

Independent sub-agents mimic real-world deployment where:

  • No prior knowledge of "correct" answers exists
  • Users provide unpredictable, context-light prompts
  • Similar skills compete for activation

Including should-not-trigger cases prevents "always-trigger" shortcuts in the A2 description. Cross-skill cases ensure activation discrimination when multiple skills from the same book are available.

Key Files in the Pressure Testing Pipeline

File Path Purpose
methodology/06-stage4-pressure-test.md Complete methodology, pass criteria, and workflow documentation
templates/test-prompts.json.template Boilerplate schema for creating skill-specific test files
SKILL.md Per-skill location where test-prompts.json resides and test-results.md is generated
methodology/07-stage5-deliver.md Post-testing stage for final digest generation and deployment

Summary

  • Pressure testing with Darwin-skill compatible test-prompts.json validates trigger logic through blind sub-agent evaluation before production.
  • The three test case types (should_trigger, should_not_trigger, edge_case) provide comprehensive coverage of activation scenarios.
  • Zero tolerance for false positives in should_not_trigger cases prevents over-triggering.
  • The 80% threshold determines whether a skill proceeds, iterates, or returns for redesign.
  • All test artifacts follow schemas defined in methodology/06-stage4-pressure-test.md and templates/test-prompts.json.template.

Frequently Asked Questions

What information does the sub-agent receive during pressure testing?

The sub-agent receives only the skill's metadata (with evaluation fields hidden), the user prompt, and an optional sibling skill list. The type, expected_behavior, and notes fields from test-prompts.json are deliberately excluded to ensure unbiased evaluation that mirrors real-world usage.

Why is the pass threshold set at 80% rather than 100%?

The 100% threshold is required for immediate acceptance. The 80% threshold allows for investigation and iteration—some failures may indicate test case flaws rather than skill defects. Below 80%, the pattern of failures suggests fundamental trigger logic problems requiring stage 2 redesign.

How does pressure testing differ from earlier verification stages?

Earlier stages (stage 1-3) verify functional correctness and output quality with full knowledge of expected behavior. Pressure testing in stage 4 specifically validates trigger discrimination using blind evaluation—the sub-agent has no access to expected answers, preventing confirmation bias in the assessment.

Where does the Darwin-skill compatibility requirement come from?

The test-prompts.json format is Darwin-compatible, meaning it adheres to schema conventions used across the Darwin skill ecosystem. This ensures interoperability with shared tooling and evaluation infrastructure. The template at templates/test-prompts.json.template enforces this compatibility for all cangjie-skill implementations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →