How to Design Cross-Skill Confusion Tests in test-prompts.json

Cross-skill confusion tests are purposely ambiguous prompts added to templates/test-prompts.json.template that verify the LLM routes overlapping queries to the correct skill rather than triggering false positives.

The cangjie-skill repository uses a JSON-based testing framework to validate skill selection accuracy. By designing strategic confusion tests in the prompt template, you can ensure that ambiguous user queries—such as those overlapping between Cangjie typing and Pinyin conversion—are handled by the appropriate skill handler.

Understanding the Test-Prompt Template Structure

The test harness reads from templates/test-prompts.json.template, which expands into the runtime test-prompts.json during CI execution. Each entry in this JSON array follows a strict schema designed to evaluate skill routing decisions.

The template accepts objects with these core fields:

  • prompt — The user-facing text fed to the LLM (e.g., "Give me the Cangjie code for the character that spells 'ma'.")
  • expectedSkill — The skill identifier that should handle the request (e.g., "cangjie-skill")
  • tags — Categorical metadata for test filtering, typically including ["confusion", "cross-skill"] for ambiguity cases
  • metadata — Optional configuration such as timeout values or priority weights

When the test runner executes, it transmits each prompt to the LLM, captures the returned skill ID, and asserts equality against expectedSkill. Any mismatch indicates a confusion failure requiring heuristic adjustment.

Designing Effective Cross-Skill Confusion Tests

Creating robust confusion tests requires identifying semantic overlap between skills and crafting prompts that genuinely test the decision boundaries.

Identify Overlapping Skill Domains

Begin by mapping domains where vocabulary intersects. In the cangjie-skill ecosystem, common overlaps occur between:

  • Cangjie typing — Character decomposition and code generation
  • Chinese-dictionary lookup — General character information retrieval
  • Pinyin-to-Zhuyin conversion — Phonetic transcription services

These skills share terminology like "character", "code", and "spelling", making them prime candidates for confusion testing.

Craft Ambiguous Prompts

Construct prompts that could logically resolve to multiple skills, forcing the selection logic to disambiguate based on subtle cues. The prompt should contain keywords attractive to incorrect skills while explicitly requiring the target skill's specific capability.

{
  "prompt": "Give me the Cangjie code for the character that spells 'ma'.",
  "expectedSkill": "cangjie-skill",
  "tags": ["confusion", "cross-skill"]
}

Here, the word "code" might attract a dictionary lookup skill, while "character" suggests general Chinese processing. The test validates that the system recognizes the specific request for Cangjie codes rather than generic encoding.

Add Negative Control Cases

Include mirror prompts that appear similar but legitimately belong to different skills. These negative controls verify the system does not over-generalize Cangjie handling to all character-related queries.

{
  "prompt": "What is the Pinyin for the character '马'?",
  "expectedSkill": "pinyin-to-zhuyin-skill",
  "tags": ["confusion", "cross-skill"]
}

This case tests that pure phonetic requests route to the conversion skill rather than defaulting to Cangjie processing.

Implementing Tests in test-prompts.json.template

Append new confusion test cases to the JSON array in templates/test-prompts.json.template. Ensure valid JSON syntax—trailing commas are not permitted in the final entry of each object or the array itself.

[
  {
    "prompt": "Give me the Cangjie code for the character that spells 'ma'.",
    "expectedSkill": "cangjie-skill",
    "tags": ["confusion", "cross-skill"],
    "metadata": { "priority": "high" }
  },
  {
    "prompt": "What is the Pinyin for the character '马'?",
    "expectedSkill": "pinyin-to-zhuyin-skill",
    "tags": ["confusion", "cross-skill"],
    "metadata": { "priority": "high" }
  }
]

According to the source code in methodology/06-stage4-pressure-test.md, these tests form part of the Stage 4 pressure testing suite designed to validate skill isolation under ambiguous conditions.

Validating Tests via CI Pipeline

Once committed, the repository's continuous integration workflow defined in .github/workflows/update-star-history.yml regenerates test-prompts.json from the template and executes the full test harness. Failures appear in CI logs when the LLM selects a skill ID differing from the expectedSkill field.

Review failures in the context of SKILL.md, which documents the intended behavior boundaries of the Cangjie skill. Adjust the skill's selection heuristics or examples if confusion tests consistently fail, then re-run the pipeline to verify fixes.

Summary

  • Cross-skill confusion tests reside in templates/test-prompts.json.template and validate that ambiguous prompts route to correct skills.
  • Each test requires a prompt, expectedSkill, and tags including "confusion" to enable categorization.
  • Effective tests target overlapping domains like Cangjie typing versus Pinyin conversion using semantically ambiguous language.
  • Negative controls prevent false positives by testing similar prompts that belong to different skills.
  • The CI pipeline in .github/workflows/update-star-history.yml automatically evaluates these tests during template expansion.

Frequently Asked Questions

What makes a prompt "confusing" enough for cross-skill testing?

A qualifying prompt contains vocabulary or concepts shared between multiple skills while requiring a specific skill's unique capability to answer correctly. For example, asking for a "code" regarding a Chinese character could imply either Cangjie input methods or general Unicode information, creating valid confusion between cangjie-skill and dictionary skills.

How does the test runner evaluate confusion test results?

The runner transmits the prompt value to the LLM and compares the returned skill identifier against the expectedSkill field defined in the JSON template. If the LLM selects any skill other than the expected one, the test fails and logs the mismatch for analysis.

Can I include metadata to adjust how confusion tests execute?

Yes. The metadata field accepts configuration objects such as {"timeout": 5000} or priority flags. These values inform the test harness about resource allocation or execution order without affecting the core validation logic comparing actual versus expected skill selection.

Where should I document the reasoning behind specific confusion tests?

Document the semantic overlap and expected routing logic in SKILL.md or adjacent methodology files like methodology/06-stage4-pressure-test.md. This documentation helps future maintainers understand why certain prompts should trigger specific skills despite surface-level ambiguity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →