How to Design Decoy Test Prompts for Cross-Skill Confusion in Cangjie-Skill

Decoy test prompts are synthetic utterances marked with "decoy": true that exercise a skill's intent-recognition pipeline without triggering side effects, enabling automated detection of cross-skill activation collisions before production deployment.

Cross-skill confusion occurs when a user's input could plausibly match the activation phrase of more than one skill, causing the platform to route the request to the wrong handler. The kangarooking/cangjie-skill repository implements a robust defense against this by embedding decoy test prompts—realistic-looking synthetic requests that validate disambiguation logic. When you design decoy test prompts for cross-skill confusion systematically, you create a safety net that catches routing errors during automated testing rather than in live user interactions.

Understanding Cross-Skill Confusion

Cross-skill confusion arises when overlapping activation phrases or similar slot patterns cause the platform's dispatcher to misroute requests. For example, a user saying "Translate hello to Japanese" might trigger a translation skill or a language-learning skill depending on subtle phrasing differences. Without systematic testing, these collisions surface only in production, degrading user trust and skill reliability.

The Four-Stage Decoy Implementation Strategy

The cangjie-skill repository implements decoy testing through a structured pipeline that generates, tags, isolates, and continuously verifies synthetic prompts.

1. Prompt Generation Using templates/test-prompts.json.template

The file templates/test-prompts.json.template defines a JSON schema for creating realistic-looking prompts that include variations in wording, intent, and slot values. This template supports both legitimate test cases and synthetic decoys, ensuring that generated utterances mimic natural language edge cases that appear in real-world usage but are unlikely to appear in standard unit tests.

2. Decoy Tagging and Quiet-Mode Execution

Each generated prompt can carry a "decoy": true flag. According to the request handler logic documented in SKILL.md, the skill checks this flag early in the processing chain and routes the request through a quiet-mode path. This path exercises the full intent-recognition and parsing logic while skipping user-visible side effects such as sending messages, modifying persistent state, or triggering external API calls.

3. Cross-Skill Isolation via Pressure Testing

During the pressure-test stage documented in methodology/06-stage4-pressure-test.md, decoy prompts are dispatched alongside real prompts to a sandbox environment. The platform's dispatcher records any collisions—instances where a decoy prompt intended for cangjie-skill is claimed by another skill. When collisions are detected, the system logs the specific utterance and triggers refinement of the skill's activation phrases.

4. Continuous Verification with Triple-Verify

The triple-verify stage defined in methodology/03-stage1.5-triple-verify.md re-runs the full suite of decoy prompts after every code change. This ensures that new intents, slot expansions, or activation phrase modifications do not unintentionally overlap with other skills in the ecosystem. The CI workflow defined in .github/workflows/update-star-history.yml automatically posts collision reports to the issue tracker, creating an auditable trail of disambiguation improvements.

Why Decoy Prompts Work for Cross-Skill Confusion Prevention

Decoy test prompts provide specific advantages that traditional testing methods cannot replicate:

  • Synthetic but realistic: They mimic natural language variations, exposing edge-case parsing bugs that only appear in the wild.
  • Flagged as decoy: The "decoy": true property prevents side-effects while still exercising the complete processing chain.
  • Run in parallel with real tests: This guarantees that any regression in disambiguation logic is caught immediately, not after deployment.
  • Automated collision detection: The methodology leverages the platform's routing logs to spot cross-skill matches without requiring manual monitoring of production traffic.

Designing Effective Decoy Test Prompts: A 5-Step Process

Follow this systematic approach to build a decoy test suite that eliminates cross-skill ambiguity:

  1. Identify Overlap Zones: Review your skill's activation phrases and compare them against other skills using the platform's skill-registry or public skill list. Look for shared verbs, nouns, or slot patterns.

  2. Create Variant Templates: In templates/test-prompts.json.template, add synonym, typo, and phrasing variations that sit just inside the gray-area of those overlaps. Include utterances that are grammatically correct but semantically ambiguous.

  3. Mark as Decoy: Add "decoy": true to each variant. The handler in SKILL.md detects this flag and routes to a no-op branch, ensuring safe execution.

  4. Run the Pressure-Test Suite: Execute scripts/generate_star_history.py or trigger the CI job; the workflow automatically posts any collisions to the issue tracker with detailed routing diagnostics.

  5. Iterate: When a collision is reported, adjust the activation phrase or add a disambiguation prompt (e.g., "Did you mean X or Y?"). Re-run the suite until the collision report returns empty.

Code Implementation Examples

The following snippets demonstrate how decoy prompts are defined and handled in the cangjie-skill codebase.

Defining a Decoy Prompt Template

{
  "prompt": "Translate {text} to {language}",
  "slots": {
    "text": "Hello, world!",
    "language": "Japanese"
  },
  "decoy": true
}

Handling Decoys in the Request Processor

def handle_request(request):
    if request.get("decoy"):
        # Run through the normal parsing pipeline but skip any side‑effects

        result = process_intent(request)
        logger.debug("Decoy processed: %s", result)
        return {"status": "ok", "decoy": True}
    # … normal handling for real user requests …

CI Configuration for Collision Detection

- name: Run decoy prompt suite
  run: |
    python -m cangjie.test_suite --decoy-only
- name: Detect collisions
  run: |
    python scripts/detect_collisions.py --output collisions.json

Critical Files in the Cangjie-Skill Repository

  • templates/test-prompts.json.template: JSON template used to generate both real and decoy test prompts with variable slot filling.
  • SKILL.md: Core documentation defining the decoy-aware request flow and quiet-mode branching logic.
  • methodology/06-stage4-pressure-test.md: Describes the pressure-testing stage where decoy prompts are executed against the platform's sandbox dispatcher.
  • methodology/03-stage1.5-triple-verify.md: Explains the triple-verify process that re-runs decoy prompts after each code change to prevent regression.
  • scripts/generate_star_history.py: CI helper that aggregates test results, including decoy collision statistics and routing logs.

Summary

  • Decoy test prompts use a "decoy": true flag to trigger quiet-mode processing, allowing safe testing of intent recognition without side effects.
  • The templates/test-prompts.json.template file structures synthetic utterances that mimic ambiguous real-world inputs.
  • methodology/06-stage4-pressure-test.md defines the sandbox environment where collisions with other skills are automatically detected.
  • Continuous verification via methodology/03-stage1.5-triple-verify.md ensures that new code changes do not reintroduce cross-skill confusion.
  • Implementing this strategy in the kangarooking/cangjie-skill repository prevents routing errors from reaching production users.

Frequently Asked Questions

What is the difference between a decoy test prompt and a regular unit test?

A regular unit test typically verifies internal function logic with controlled inputs, while a decoy test prompt is a synthetic utterance that travels through the full platform dispatcher and intent-recognition pipeline. Decoys specifically test cross-skill boundary conditions and collision scenarios that unit tests cannot simulate because they require the actual routing infrastructure.

How does the cangjie-skill repository detect cross-skill collisions automatically?

The repository uses the pressure-test stage defined in methodology/06-stage4-pressure-test.md to dispatch decoy prompts to a sandbox environment. The platform's dispatcher logs which skill claims each prompt; if a decoy intended for cangjie-skill is claimed by another skill, the scripts/detect_collisions.py utility records the collision and the CI pipeline posts an alert to the issue tracker.

Can decoy prompts prevent all types of skill routing errors?

Decoy prompts effectively prevent cross-skill confusion caused by overlapping activation phrases and ambiguous slot patterns. However, they cannot catch errors stemming from platform-level dispatcher bugs or skills that are not yet published in the skill-registry. They are most effective when combined with the triple-verify process to catch regressions in your own skill's logic.

Where should I define the decoy flag in my skill's test suite?

Define the decoy flag in your JSON prompt templates, specifically in a file like templates/test-prompts.json.template. Each test case object should include the "decoy": true property. Your skill's request handler, as documented in SKILL.md, must check for this property at the entry point to route the request through the quiet-mode processing branch.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →