How Cross-Skill Confusion Testing Works with Sibling Skills in Cangjie-Skill
Cross-skill confusion testing in Cangjie-Skill uses deliberately ambiguous prompts to verify that sibling skills (skills sharing overlapping concepts) remain uniquely identifiable and do not interfere with each other during LLM execution.
The Cangjie-Skill framework implements rigorous quality assurance through its pressure-testing stage, with cross-skill confusion testing serving as a critical mechanism to ensure discrete, unambiguous skill boundaries. This process guarantees that when an AI agent processes loosely worded user requests, it reliably selects the correct skill from among related siblings.
What Are Sibling Skills?
Sibling skills are distilled capabilities that live in the same skill pack or share overlapping conceptual domains. For example, investment_decision and product_strategy might both respond to prompts about evaluating tech startups.
Without explicit boundary enforcement, these skills can bleed into each other, causing unpredictable agent behavior. The confusion test systematically exposes and eliminates these vulnerabilities.
The Cross-Skill Confusion Test Pipeline
Step 1: Trigger and Boundary Definition
Every skill in Cangjie-Skill defines its activation criteria in SKILL.md:
- Trigger field: The exact phrasing that should invoke this skill
- Boundary field: Contexts where the skill must not apply
These definitions are reinforced by the INDEX.md linking graph, which records sibling relationships across the skill pack. According to the repository's README.en.md at line 36, this structure forms the foundation for all confusion testing.
Step 2: Bait and Confusion Prompt Generation
The test harness generates two prompt types for each skill, using templates/test-prompts.json.template as its source:
| Prompt Type | Purpose | Example |
|---|---|---|
| Bait | Verify clean activation | Clear trigger phrase within skill scope |
| Confusion | Test boundary separation | Ambiguous wording shared with siblings |
A minimal confusion test definition appears as:
{
"skill": "investment_decision",
"prompt": "When evaluating a tech startup, should I focus on market size or team experience?",
"expectedSkill": "investment_decision"
},
{
"skill": "product_strategy",
"prompt": "When evaluating a tech startup, should I focus on market size or team experience?",
"expectedSkill": "product_strategy"
}
The identical natural-language prompt targets two sibling skills. The test passes only if each execution returns its expected skill's RIA-TV++ structure.
Step 3: Execution and Verification
The generated prompt flows through the skill execution pipeline:
def run_confusion_test(prompt, expected_skill):
response = execute_skill(prompt) # Calls Claude via the skill runner
skill_id = extract_skill_id(response) # Parses the RIA-TV++ header
assert skill_id == expected_skill, "Cross-skill confusion detected"
The response validation checks the RIA-TV++ structure:
- R: Quote block with source attribution
- I: Reconstruction of core methodology
- A: Two contrasting examples (A1/A2)
- T: Execution steps
- V++: Enhanced verification with boundaries
If skill_id does not match expected_skill, the test fails—the prompt triggered the wrong sibling.
Step 4: Failure Handling and Iteration
A failed confusion test triggers immediate remediation. Per README.en.md at line 37, the skill returns to Triple Verification for reconstruction:
- Trigger refinement: Sharpen activation phrasing
- Boundary tightening: Explicitly exclude sibling domains
- Skill splitting: Decompose overlapping capabilities into granular units
This loop repeats for every sibling pair until all confusion scenarios resolve correctly.
Key Implementation Files
The cross-skill confusion testing system spans these core files in kangarooking/cangjie-skill:
README.en.md— RIA-TV++ pipeline overview and pressure-testing documentationSKILL.md— Per-skill trigger, boundary, and execution definitionsmethodology/06-stage4-pressure-test.md— Detailed pressure-testing methodologytemplates/test-prompts.json.template— Prompt generation template for automated testingtemplates/SKILL.md.template— Skill definition skeleton with trigger/boundary sectionstemplates/INDEX.md.template— Sibling relationship linking graph
Why Cross-Skill Confusion Testing Matters
| Without Testing | With Confusion Testing |
|---|---|
| Agent randomly selects between sibling skills | Deterministic skill selection based on precise boundaries |
| Overlapping capabilities cause inconsistent outputs | Each skill delivers predictable, scoped behavior |
| User prompts with ambiguous wording fail unpredictably | System handles loose phrasing through refined triggers |
By enforcing discrete skill boundaries through systematic confusion exposure, Cangjie-Skill ensures production-ready reliability for agent-driven execution.
Summary
- Cross-skill confusion testing deliberately creates ambiguous prompts shared between sibling skills to verify boundary enforcement.
- Each skill's
SKILL.mddefines triggers (activation) and boundaries (exclusion), reinforced byINDEX.mdrelationship graphs. - The test harness in
templates/test-prompts.json.templategenerates bait and confusion prompts, executed against Claude through the skill pipeline. - RIA-TV++ structure validation confirms correct skill selection; failures return skills to Triple Verification for refinement.
- This iterative process, documented in
methodology/06-stage4-pressure-test.md, guarantees unambiguous, deterministic skill behavior in production environments.
Frequently Asked Questions
How does Cangjie-Skill define which skills are "siblings"?
Sibling relationships are recorded in the INDEX.md linking graph, implemented via templates/INDEX.md.template. This graph tracks skills sharing conceptual overlap or residing in the same skill pack, enabling the test harness to generate targeted confusion prompts between related capabilities.
What happens when a confusion test fails?
The failing skill returns to Triple Verification for reconstruction. According to README.en.md, developers refine trigger phrasing, tighten boundary definitions, or split the skill into more granular units until the ambiguity resolves. This ensures no sibling interference reaches production.
Why use identical prompts for different expected skills?
Identical prompts create maximum ambiguity—if the system can still route to the correct skill despite identical input, the boundaries are sufficiently robust. This represents a worst-case scenario that validates real-world performance on loosely worded user requests.
What is RIA-TV++ and why does it matter for confusion testing?
RIA-TV++ is Cangjie-Skill's structured output format containing Reconstruction, Interpretation, A1/A2 examples, Tasks, Verification, and enhanced V++ boundary specifications. The confusion test parses this structure to confirm the actual invoked skill matches the expected skill, providing objective pass/fail criteria.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →