How Cross-Skill Confusion Testing Works with Sibling Skills in Cangjie-Skill

Cross-skill confusion testing in Cangjie-Skill uses deliberately ambiguous prompts to verify that sibling skills (skills sharing overlapping concepts) remain uniquely identifiable and do not interfere with each other during LLM execution.

The Cangjie-Skill framework implements rigorous quality assurance through its pressure-testing stage, with cross-skill confusion testing serving as a critical mechanism to ensure discrete, unambiguous skill boundaries. This process guarantees that when an AI agent processes loosely worded user requests, it reliably selects the correct skill from among related siblings.

What Are Sibling Skills?

Sibling skills are distilled capabilities that live in the same skill pack or share overlapping conceptual domains. For example, investment_decision and product_strategy might both respond to prompts about evaluating tech startups.

Without explicit boundary enforcement, these skills can bleed into each other, causing unpredictable agent behavior. The confusion test systematically exposes and eliminates these vulnerabilities.

The Cross-Skill Confusion Test Pipeline

Step 1: Trigger and Boundary Definition

Every skill in Cangjie-Skill defines its activation criteria in SKILL.md:

  • Trigger field: The exact phrasing that should invoke this skill
  • Boundary field: Contexts where the skill must not apply

These definitions are reinforced by the INDEX.md linking graph, which records sibling relationships across the skill pack. According to the repository's README.en.md at line 36, this structure forms the foundation for all confusion testing.

Step 2: Bait and Confusion Prompt Generation

The test harness generates two prompt types for each skill, using templates/test-prompts.json.template as its source:

Prompt Type Purpose Example
Bait Verify clean activation Clear trigger phrase within skill scope
Confusion Test boundary separation Ambiguous wording shared with siblings

A minimal confusion test definition appears as:

{
  "skill": "investment_decision",
  "prompt": "When evaluating a tech startup, should I focus on market size or team experience?",
  "expectedSkill": "investment_decision"
},
{
  "skill": "product_strategy",
  "prompt": "When evaluating a tech startup, should I focus on market size or team experience?",
  "expectedSkill": "product_strategy"
}

The identical natural-language prompt targets two sibling skills. The test passes only if each execution returns its expected skill's RIA-TV++ structure.

Step 3: Execution and Verification

The generated prompt flows through the skill execution pipeline:

def run_confusion_test(prompt, expected_skill):
    response = execute_skill(prompt)               # Calls Claude via the skill runner

    skill_id = extract_skill_id(response)          # Parses the RIA-TV++ header

    assert skill_id == expected_skill, "Cross-skill confusion detected"

The response validation checks the RIA-TV++ structure:

  • R: Quote block with source attribution
  • I: Reconstruction of core methodology
  • A: Two contrasting examples (A1/A2)
  • T: Execution steps
  • V++: Enhanced verification with boundaries

If skill_id does not match expected_skill, the test fails—the prompt triggered the wrong sibling.

Step 4: Failure Handling and Iteration

A failed confusion test triggers immediate remediation. Per README.en.md at line 37, the skill returns to Triple Verification for reconstruction:

  1. Trigger refinement: Sharpen activation phrasing
  2. Boundary tightening: Explicitly exclude sibling domains
  3. Skill splitting: Decompose overlapping capabilities into granular units

This loop repeats for every sibling pair until all confusion scenarios resolve correctly.

Key Implementation Files

The cross-skill confusion testing system spans these core files in kangarooking/cangjie-skill:

  • README.en.md — RIA-TV++ pipeline overview and pressure-testing documentation
  • SKILL.md — Per-skill trigger, boundary, and execution definitions
  • methodology/06-stage4-pressure-test.md — Detailed pressure-testing methodology
  • templates/test-prompts.json.template — Prompt generation template for automated testing
  • templates/SKILL.md.template — Skill definition skeleton with trigger/boundary sections
  • templates/INDEX.md.template — Sibling relationship linking graph

Why Cross-Skill Confusion Testing Matters

Without Testing With Confusion Testing
Agent randomly selects between sibling skills Deterministic skill selection based on precise boundaries
Overlapping capabilities cause inconsistent outputs Each skill delivers predictable, scoped behavior
User prompts with ambiguous wording fail unpredictably System handles loose phrasing through refined triggers

By enforcing discrete skill boundaries through systematic confusion exposure, Cangjie-Skill ensures production-ready reliability for agent-driven execution.

Summary

  • Cross-skill confusion testing deliberately creates ambiguous prompts shared between sibling skills to verify boundary enforcement.
  • Each skill's SKILL.md defines triggers (activation) and boundaries (exclusion), reinforced by INDEX.md relationship graphs.
  • The test harness in templates/test-prompts.json.template generates bait and confusion prompts, executed against Claude through the skill pipeline.
  • RIA-TV++ structure validation confirms correct skill selection; failures return skills to Triple Verification for refinement.
  • This iterative process, documented in methodology/06-stage4-pressure-test.md, guarantees unambiguous, deterministic skill behavior in production environments.

Frequently Asked Questions

How does Cangjie-Skill define which skills are "siblings"?

Sibling relationships are recorded in the INDEX.md linking graph, implemented via templates/INDEX.md.template. This graph tracks skills sharing conceptual overlap or residing in the same skill pack, enabling the test harness to generate targeted confusion prompts between related capabilities.

What happens when a confusion test fails?

The failing skill returns to Triple Verification for reconstruction. According to README.en.md, developers refine trigger phrasing, tighten boundary definitions, or split the skill into more granular units until the ambiguity resolves. This ensures no sibling interference reaches production.

Why use identical prompts for different expected skills?

Identical prompts create maximum ambiguity—if the system can still route to the correct skill despite identical input, the boundaries are sufficiently robust. This represents a worst-case scenario that validates real-world performance on loosely worded user requests.

What is RIA-TV++ and why does it matter for confusion testing?

RIA-TV++ is Cangjie-Skill's structured output format containing Reconstruction, Interpretation, A1/A2 examples, Tasks, Verification, and enhanced V++ boundary specifications. The confusion test parses this structure to confirm the actual invoked skill matches the expected skill, providing objective pass/fail criteria.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →