How Pressure Testing with Darwin-Skill Compatible test-prompts.json Works in Cangjie-Skill
Pressure testing validates that a skill's trigger logic fires only when appropriate by using a Darwin-compatible test-prompts.json file and blind sub-agent evaluation before delivery.
Pressure testing is the fourth stage of the cangjie-skill development pipeline. It ensures a skill's trigger logic (defined in the A2 field) behaves correctly under realistic conditions. The process depends on a Darwin-compatible test-prompts.json structure and independent sub-agents that evaluate test cases without prior knowledge of expected outcomes.
What Makes a test-prompts.json Darwin-Skill Compatible
The Darwin-compatible format follows a strict schema designed for blind evaluation. You create this file from the template at templates/test-prompts.json.template.
A valid test-prompts.json contains:
skill– the skill's slug identifierversion– semantic version of the test suitetest_cases– an array of evaluation scenarios, each with:id– unique case identifiertype– one ofshould_trigger,should_not_trigger, oredge_caseprompt– the user utterance to testexpected_behavior– what correct behavior looks likenotes– context for human reviewers (hidden from sub-agents)
The schema is documented in methodology/06-stage4-pressure-test.md.
The Five Steps of Pressure Testing
1. Create the Test File
Fill the template with three categories of test cases:
- Should-trigger cases – typical user utterances that must activate the skill
- Should-not-trigger cases – "bait" prompts that test for false positives
- Edge cases – ambiguous inputs where the decision requires nuanced reasoning
Include cross-skill prompts if sibling skills exist in the same book to prevent activation conflicts.
2. Run Blind Sub-Agent Evaluations
For each test_case, spin up a clean sub-agent with strictly limited information:
def evaluate_prompt(skill_path, prompt, sibling_skills=None):
# Load skill description but hide trigger details
skill_meta = load_skill_meta(skill_path, hide=['type','expected_behavior','notes'])
# Run the sub-agent with only the prompt and optional sibling list
result = sub_agent.run(
skill_meta=skill_meta,
user_prompt=prompt,
sibling_skills=sibling_skills,
)
return {
"would_trigger": result.would_trigger,
"reason": result.reason,
"if_triggered_action": result.action,
}
Critical constraint: The sub-agent receives only:
- The skill's path and content (metadata without evaluation fields)
- The user prompt
- Optional sibling skill list
The type, expected_behavior, and notes fields are intentionally withheld to eliminate bias.
3. Compare Results to Expected Behavior
The main workflow matches sub-agent output against test-prompts.json expectations:
| Case Type | Passing Criteria |
|---|---|
should_trigger |
Sub-agent triggers skill and performs expected action |
should_not_trigger |
Sub-agent does not trigger (zero tolerance for false positives) |
edge_case |
Decision aligns with reasoning in expected_behavior |
4. Calculate Pass Rate and Decide
- 100% pass → Skill accepted for delivery
- ≥80% pass → Investigate failures; decide whether to refine A2 trigger description or adjust test cases
- <80% pass → Return to stage 2 for complete trigger logic redesign
5. Iterate Until Threshold Met
After any fix, re-run blind tests. Update test-results.md in the skill directory to track progress.
Example test-prompts.json Structure
{
"skill": "inversion-thinking",
"version": "0.1.0",
"test_cases": [
{
"id": "should-trigger-01",
"type": "should_trigger",
"prompt": "我要决定要不要接这个新项目,列了一堆好处但还是没底",
"expected_behavior": "调用 inversion‑thinking, 反问 '最不希望发生什么'",
"notes": "正面场景: 决策纠结"
},
{
"id": "should-not-trigger-01",
"type": "should_not_trigger",
"prompt": "帮我查一下这个 API 的参数",
"expected_behavior": "纯信息查询, 不应调用任何决策 skill",
"notes": "诱饵: 非决策场景"
},
{
"id": "edge-01",
"type": "edge_case",
"prompt": "我在想晚饭吃什么",
"expected_behavior": "日常琐事, 不应调用 (虽然字面是'决策')",
"notes": "边界: 区分严肃决策和日常选择"
}
]
}
This example demonstrates defense in depth: positive testing, negative testing, and boundary analysis in a single suite.
Why Blind Sub-Agent Evaluation Matters
The trigger (A2) is the most fragile component of any skill. Without pressure testing, a skill can appear functional in development while silently mis-triggering in production.
Independent sub-agents mimic real-world deployment where:
- No prior knowledge of "correct" answers exists
- Users provide unpredictable, context-light prompts
- Similar skills compete for activation
Including should-not-trigger cases prevents "always-trigger" shortcuts in the A2 description. Cross-skill cases ensure activation discrimination when multiple skills from the same book are available.
Key Files in the Pressure Testing Pipeline
| File Path | Purpose |
|---|---|
methodology/06-stage4-pressure-test.md |
Complete methodology, pass criteria, and workflow documentation |
templates/test-prompts.json.template |
Boilerplate schema for creating skill-specific test files |
SKILL.md |
Per-skill location where test-prompts.json resides and test-results.md is generated |
methodology/07-stage5-deliver.md |
Post-testing stage for final digest generation and deployment |
Summary
- Pressure testing with Darwin-skill compatible
test-prompts.jsonvalidates trigger logic through blind sub-agent evaluation before production. - The three test case types (
should_trigger,should_not_trigger,edge_case) provide comprehensive coverage of activation scenarios. - Zero tolerance for false positives in
should_not_triggercases prevents over-triggering. - The 80% threshold determines whether a skill proceeds, iterates, or returns for redesign.
- All test artifacts follow schemas defined in
methodology/06-stage4-pressure-test.mdandtemplates/test-prompts.json.template.
Frequently Asked Questions
What information does the sub-agent receive during pressure testing?
The sub-agent receives only the skill's metadata (with evaluation fields hidden), the user prompt, and an optional sibling skill list. The type, expected_behavior, and notes fields from test-prompts.json are deliberately excluded to ensure unbiased evaluation that mirrors real-world usage.
Why is the pass threshold set at 80% rather than 100%?
The 100% threshold is required for immediate acceptance. The 80% threshold allows for investigation and iteration—some failures may indicate test case flaws rather than skill defects. Below 80%, the pattern of failures suggests fundamental trigger logic problems requiring stage 2 redesign.
How does pressure testing differ from earlier verification stages?
Earlier stages (stage 1-3) verify functional correctness and output quality with full knowledge of expected behavior. Pressure testing in stage 4 specifically validates trigger discrimination using blind evaluation—the sub-agent has no access to expected answers, preventing confirmation bias in the assessment.
Where does the Darwin-skill compatibility requirement come from?
The test-prompts.json format is Darwin-compatible, meaning it adheres to schema conventions used across the Darwin skill ecosystem. This ensures interoperability with shared tooling and evaluation infrastructure. The template at templates/test-prompts.json.template enforces this compatibility for all cangjie-skill implementations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →