What Is Pressure Testing in Cangjie-Skill and How to Perform It
Pressure testing is the mandatory fourth stage of the Cangjie-skill development pipeline that verifies trigger accuracy and output quality before shipping, using independent sub-agent blind testing to catch the most fragile failure modes in skill activation.
In the Cangjie-skill framework—an open-source methodology for building reliable AI skills—pressure testing serves as the final gate before release. The kangarooking/cangjie-skill repository implements this as a rigorous, protocol-driven validation system designed to solve the hardest problem in skill development: ensuring your skill actually activates when it should and stays silent when it shouldn't.
The Core Problem: Why Pressure Testing Is Non-Negotiable
According to the Cangjie-skill source methodology in [methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md), the A2 trigger (the decision logic that determines whether a skill should activate) is the most fragile component of any skill. An inaccurate trigger doesn't just produce wrong answers—it renders the skill invisible to users entirely.
Pressure testing exists because:
- Trigger bugs are silent failures — unlike output errors, missed activations go unreported
- No other stage catches this — unit tests, integration tests, and validation stages don't simulate real user-facing trigger decisions
- Pre-release is your only chance — once shipped, trigger misbehavior erodes user trust permanently
How Pressure Testing Works: The Sub-Agent Blind Test Architecture
Cangjie-skill's pressure testing follows a strict independent sub-agent blind testing principle. Here's how the mechanism operates:
Testing Principle: Zero Knowledge Simulation
Each prompt evaluation spawns a fresh sub-agent with strictly limited context:
{
"skill_description": "<only the public description>",
"user_prompt": "<the test input>"
}
The sub-agent receives no expected answer, no notes, no type fields, and no hints about whether the prompt should trigger the skill. This architecture simulates genuine user-facing conditions where the system must make an autonomous activation decision.
Fallback Mode: Self-Testing with Confidence Labels
When sub-agent capability is unavailable, the main process performs a self-test and records results in test-results.md with an explicit lower confidence label. This ensures testing continuity across environments while maintaining honest result grading.
Required Test Case Types and Quantities
The pressure testing protocol mandates three mandatory prompt categories, with rejection enforced if any class is missing:
| Category | Quantity | Purpose |
|---|---|---|
should_trigger |
3-5 cases | Verify the skill activates when appropriate |
should_not_trigger |
2-3 cases ("bait") | Verify the skill stays silent for unrelated prompts |
edge_case |
1-3 cases | Test boundary conditions and ambiguous inputs |
Cross-Skill Confusion Prevention
At least one should_not_trigger case must specifically target a sibling skill from the same book—another skill that shares domain overlap. This prevents trigger collisions where multiple skills fight for the same user intent.
Step-by-Step Pressure Testing Execution
Based on the execution flow defined in the Cangjie-skill source:
Step 1: Prepare Test Prompts
Create test-prompts.json for each skill following the mandatory category distribution:
{
"skill_name": "weather_lookup",
"test_cases": [
{
"type": "should_trigger",
"prompt": "What's the weather in Beijing tomorrow?",
"expected": "activation"
},
{
"type": "should_trigger",
"prompt": "Will it rain this weekend in Shanghai?",
"expected": "activation"
},
{
"type": "should_not_trigger",
"prompt": "Book a flight to Tokyo next Tuesday",
"expected": "no_activation",
"notes": "bait: travel booking, not weather"
},
{
"type": "should_not_trigger",
"prompt": "Convert 100 USD to EUR",
"expected": "no_activation",
"notes": "bait: currency conversion, sibling skill collision test"
},
{
"type": "edge_case",
"prompt": "Weather?",
"expected": "activation",
"notes": "minimal viable query"
}
]
}
Step 2: Execute Blind Sub-Agent Tests
Run the pressure test harness which:
- Spawns isolated sub-agents per test case
- Injects only
skill_descriptionanduser_prompt - Captures the activation decision and any output
Step 3: Compare Against Expectations
The harness evaluates sub-agent responses against the ground truth:
- Correct activation (
should_trigger→ triggered ✓) - Correct silence (
should_not_trigger→ not triggered ✓) - Correct edge handling (
edge_case→ appropriate behavior ✓)
Step 4: Compute Pass Rate and Decide
| Pass Rate | Action |
|---|---|
| 100% | Accept — skill proceeds to release |
| ≥80% | Review — analyze failures to determine if the skill (A2 trigger) or test cases need correction |
| <80% | Reject — skill requires fundamental trigger redesign before retesting |
Why Blind Testing Beats Transparent Validation
The Cangjie-skill methodology explicitly rejects "transparent" testing (where the evaluator knows expected outcomes) because:
- Real users don't provide labels — the production system must infer relevance from raw prompts alone
- Hinted evaluation produces false confidence — knowing the "right answer" biases judgment
- Sub-agent independence prevents contamination — each test runs in cognitive isolation, eliminating cross-case pattern matching
Summary
- Pressure testing in Cangjie-skill is the fourth-stage, pre-release validation that verifies trigger accuracy and output quality through independent sub-agent blind testing
- The A2 trigger is the most fragile skill component, making pressure testing the only reliable pre-shipping detection method for activation failures
- Three mandatory prompt types (
should_trigger,should_not_trigger,edge_case) must all be present, with specific quantity requirements - Cross-sibling bait cases prevent skill collisions within the same book
- Pass criteria are strict: 100% for automatic acceptance, ≥80% for manual review, below 80% for rejection
- Sub-agent architecture simulates genuine user conditions; fallback self-testing degrades gracefully with confidence labeling
Frequently Asked Questions
What makes pressure testing different from other testing stages in Cangjie-skill?
Pressure testing is the only stage that validates the A2 trigger decision boundary using fully blind, user-simulating conditions. Earlier stages verify output quality given correct activation; pressure testing verifies that activation happens correctly in the first place. The independent sub-agent architecture ensures no hint contamination skews results.
Can pressure testing run without sub-agent support, and how reliable is it?
Yes. When sub-agent capability is unavailable, the system falls back to self-testing with explicit lower confidence labeling in test-results.md. This degraded mode preserves testing continuity but marks results as less authoritative. Production releases should prefer sub-agent execution where possible.
Why must should_not_trigger cases include sibling skill collisions?
Sibling skills from the same book often share semantic domains, creating natural trigger ambiguity. Without explicit cross-skill confusion testing, multiple skills might simultaneously claim the same user intent, causing unpredictable routing. The mandatory sibling bait case ensures trigger boundaries are intentionally designed, not accidentally overlapping.
What happens if I submit pressure tests missing one of the three required prompt types?
The Cangjie-skill validation layer rejects the submission entirely. The protocol treats category completeness as a hard requirement—partial test coverage provides false assurance about trigger robustness. All three classes (activation, silence, and boundary behavior) must be explicitly exercised and documented.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →