What Happens When a Skill Fails Pressure Testing in Cangjie-Skill: A Complete Breakdown
When a Cangjie-skill fails pressure testing, it is immediately sent back to Stage 2 for complete reconstruction ("回炉") until it achieves at least an 80% pass rate on blind evaluation.
Pressure testing is the fourth and final quality gate in the RIA-TV++ workflow used by kangarooking/cangjie-skill. This stage validates whether a generated skill correctly triggers (would_trigger) and behaves as expected when faced with realistic user prompts. Understanding what happens on failure is critical for anyone building or debugging skills in this framework.
The Pressure Testing Failure Sequence
When a skill does not meet the threshold, the pipeline executes a structured recovery protocol rather than allowing partial deployments.
1. Failure Detection and Threshold Evaluation
Each skill is evaluated against a test-prompts.json file containing should_trigger, should_not_trigger, and edge_case scenarios. An independent sub-agent (or the main process if sub-agents are unavailable) blindly evaluates each case.
The evaluation compares the agent's output—specifically would_trigger, reason, and if_triggered_action—against the expected_behavior defined in the test file.
If the aggregate pass rate falls below 80%, the skill is marked as failed. This threshold is defined in methodology/06-stage4-pressure-test.md at lines 78-82.
// tests/example-skill/test-prompts.json
{
"skill": "example-skill",
"version": "0.1.0",
"test_cases": [
{
"id": "should-trigger-01",
"type": "should_trigger",
"prompt": "我在决定是否接受一个新项目,犹豫不决",
"expected_behavior": "调用 example-skill, 询问“最不希望发生什么”",
"notes": "正面决策场景"
},
{
"id": "should-not-trigger-01",
"type": "should_not_trigger",
"prompt": "帮我查一下这个 API 的参数",
"expected_behavior": "不调用任何决策 skill",
"notes": "信息查询诱饵"
},
{
"id": "edge-01",
"type": "edge_case",
"prompt": "今晚吃什么好",
"expected_behavior": "不调用 (日常琐事)",
"notes": "边界模糊"
}
]
}
2. Complete Reconstruction ("回炉")
Failed skills trigger Stage 2 reconstruction—not a superficial patch. As documented at lines 7-8 of the pressure-test methodology: "不通过的必须回炉 … 重做阶段 2 的 A2 / E / B" (Those that do not pass must be re-melted … redo Stage 2's A2 / E / B).
This means the entire RIA++ construction is regenerated:
- R (Role)
- I (Instruction)
- A1 (Agent 1 prompt)
- A2 (Agent 2 / Trigger description)
- E (Evaluation criteria)
- B (Behavior/Tool binding)
# Illustrative workflow
cangjie-skill rebuild --skill example-skill # triggers Stage 2 rebuild
cangjie-skill pressure-test --skill example-skill # re-runs test suite
3. Root-Cause Diagnosis
Before reconstruction, the pipeline applies a decision matrix to determine the specific fix strategy:
| Failure Pattern | Action Taken |
|---|---|
| Ambiguous trigger description (A2) | Refine the skill's trigger logic |
| Unforeseen but valid user scenario | Expand skill coverage |
| Unrealistic "bait" prompt | Revise the test case instead |
This matrix appears at lines 84-89 of methodology/06-stage4-pressure-test.md. The distinction between skill deficiency and test deficiency prevents wasted rebuilding cycles.
4. Iterative Re-Testing
After reconstruction, a fresh test-prompts.json is generated (or the existing file is reused) and pressure testing executes again. This loop continues until the skill achieves:
- Minimum threshold: ≥ 80% pass rate
- Target goal: 100% pass rate (ideal)
5. Deployment Gate: Failed Skills Never Ship
The final stage enforces strict quality control. Per methodology/07-stage5-deliver.md at lines 52-54, only skills passing pressure testing are copied or symlinked into the user's skills directory.
# Final installation only for passed skills
cangjie-skill install --skill example-skill --target ~/.claude/skills/
Failed skills remain isolated in the build folder, awaiting further iteration. They are never exposed to production environments.
Key Source Files for Understanding Pressure Test Failures
| File | Purpose | Critical Lines |
|---|---|---|
methodology/06-stage4-pressure-test.md |
Full testing protocol and failure handling | 7-8 (回炉 policy), 78-82 (threshold), 84-89 (diagnosis) |
SKILL.md |
RIA-TV++ pipeline overview | Complete workflow description |
methodology/07-stage5-deliver.md |
Deployment restrictions | 52-54 (installation gate) |
templates/test-prompts.json.template |
Test case authoring template | Schema reference |
Summary
- 80% pass rate is the hard floor for pressure testing in Cangjie-skill
- Failure triggers "回炉" (re-melting): complete Stage 2 reconstruction, not incremental fixes
- Root-cause diagnosis distinguishes skill flaws from test flaws before rebuilding
- Failed skills are deployment-blocked and remain in build folder until passing
- Iteration continues until the skill meets quality thresholds
Frequently Asked Questions
What is the exact pass rate threshold for pressure testing in Cangjie-skill?
The documented threshold is 80% as defined in methodology/06-stage4-pressure-test.md lines 78-82. The methodology notes that 100% is the ideal target, but 80% is the minimum bar for proceeding to deployment.
Does a failed pressure test trigger automatic fixes or manual intervention?
The failure triggers automated reconstruction ("回炉") back to Stage 2, but the pipeline first applies a decision matrix to diagnose whether the skill itself or the test case requires adjustment. This diagnosis logic appears at lines 84-89 of the pressure-test methodology.
Can a partially passing skill be deployed with warnings?
No. According to methodology/07-stage5-deliver.md lines 52-54, the delivery stage explicitly excludes failed skills from installation. Only skills that have fully passed pressure testing are copied or symlinked to the target skills directory.
How are test cases structured for pressure testing?
Test cases follow the test-prompts.json schema with three types: should_trigger (positive cases), should_not_trigger (negative/bait cases), and edge_case (boundary scenarios). Each case specifies prompt, expected_behavior, and optional notes for evaluation context.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →