How to Design Effective Test Prompts with Decoy Scenarios for cangjie‑skill

Effective test prompts with decoy scenarios in cangjie‑skill use a three‑category JSON schema (should_trigger, should_not_trigger, edge_case) to pressure‑test AI Skills and prevent false positives through deliberately misleading prompts that share keywords with sibling skills.

The cangjie‑skill framework converts extracted knowledge‑units—principles, frameworks, and cases—into AI Skills that agents can invoke. To ensure each skill fires only in its intended context, the project implements a rigorous pressure‑test stage based on the RIA‑TV++ methodology. Central to this validation is a structured JSON test‑prompt file that embeds decoy scenarios designed to trip up imprecise trigger logic.

Understanding the Test‑Prompt Schema

The canonical template resides at templates/test‑prompts.json.template. Every generated skill pack includes a populated version of this file with three distinct test families:

Test Type Purpose Example Prompt Pattern
should_trigger Confirm positive activation on legitimate user intent "{{用户的典型话语 1 — 正面场景}}"
should_not_trigger Decoy scenarios that resemble valid triggers but must not fire "{{诱饵 1 — 看似相关但实际不该调用}}"
edge_case Borderline inputs forcing explicit decision explanations "{{边界模糊场景}}"

This schema enforces zero tolerance for decoy failures while allowing an 80% minimum pass rate for edge cases, as documented in [methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md).

Why Decoy Scenarios Are Critical

Decoy scenarios serve two essential functions during Stage 4 – Pressure Test:

  • Prevent false positives — They replicate phrasing patterns that could activate sibling skills, forcing the selector’s trigger logic to distinguish subtle intent differences
  • Enforce skill isolation — By requiring minimum_pass_rate ≥ 0.8 with absolute decoy compliance, the system prevents skill hijacking in multi‑agent deployments

This approach aligns with the cangjie‑skill design principle: "only expose what truly belongs to the skill." The methodology explicitly references cross‑skill discrimination through sibling‑skill slugs in decoy definitions.

Building Robust Test Prompts: A 5‑Step Framework

Follow this structured workflow when creating test-prompts.json for a new skill:

  1. Extract core trigger phrases from the source material (the R component in RIA‑TV++)
  2. Draft three positive examples — Vary vocabulary while preserving underlying user intent
  3. Construct two strategic decoys:
    • Decoy 1: Shares keywords but maps to a different concept within the same source book
    • Decoy 2: Would correctly trigger a sibling skill, testing cross‑skill boundary precision
  4. Add one edge‑case — Force explanatory reasoning (e.g., conditional applicability queries)
  5. Configure minimum_pass_rate — The standard 0.8 threshold balances strictness with practical variance

Complete Test‑Prompt Example

Below is a fully populated test-prompts.json for a hypothetical example-skill:

{
  "skill": "example-skill",
  "version": "0.1.0",
  "source_book": "The Example Book — Jane Doe",
  "darwin_compatible": true,
  "test_cases": [
    {
      "id": "should-trigger-01",
      "type": "should_trigger",
      "prompt": "请告诉我如何在实际项目中使用《示例方法》",
      "expected_behavior": "应激活 example-skill, 并提供实施步骤",
      "notes": "正面场景:直接请求使用方法"
    },
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "我想了解《示例方法》中的相关案例",
      "expected_behavior": "不应激活本 skill, 因为该请求应由 example-case-skill 处理",
      "notes": "跨 skill 混淆诱饵"
    },
    {
      "id": "edge-01",
      "type": "edge_case",
      "prompt": "如果项目规模很小,是否仍需要使用《示例方法》?",
      "expected_behavior": "说明方法在小规模项目中的适用性或限制",
      "notes": "边界模糊场景"
    }
  ],
  "minimum_pass_rate": 0.8,
  "notes": "3 条 should_trigger + 1 条 should_not_trigger + 1 条 edge_case"
}

Validating Test Results Programmatically

While the cangjie‑skill pipeline automates evaluation, this Python snippet illustrates the validation logic from [SKILL.md](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) specifications:

import json

def evaluate_results(results, tmpl):
    """Check if test outcomes meet minimum_pass_rate threshold."""
    passed = sum(1 for r in results if r['passed'])
    rate = passed / len(results)
    return rate >= tmpl['minimum_pass_rate']

with open('test-prompts.json') as f:
    tmpl = json.load(f)

# Mock results for illustration (actual pipeline queries LLM)

mock_results = [{'id': tc['id'], 'passed': True} for tc in tmpl['test_cases']]

if evaluate_results(mock_results, tmpl):
    print("All tests passed – skill is ready for deployment")
else:
    print("Some tests failed – revisit prompts or trigger logic")

Key Implementation Files

File Path Role in Test‑Prompt Design
templates/test‑prompts.json.template Master schema defining all three test categories with decoy requirements
[methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) Pressure‑test stage documentation, acceptance criteria, and zero‑tolerance rationale
[SKILL.md](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) Execution spec linking test‑prompt validation to deployment gates
buffett-letters-skill/ (example pack) Production skill demonstrating populated test prompts in context

Summary

  • Decoy scenarios in should_not_trigger entries are non‑negotiable for preventing skill collisions
  • The minimum_pass_rate of 0.8 applies globally, but decoys must achieve 100% negative accuracy
  • Test‑prompt files follow a strict JSON schema from templates/test‑prompts.json.template
  • Cross‑skill discrimination is validated through sibling‑skill decoys per RIA‑TV++ Stage 4
  • Real‑world examples exist in the Buffett Letters skill pack and other repository references

Frequently Asked Questions

What makes a decoy scenario effective in cangjie‑skill?

An effective decoy shares surface‑level keywords with the target skill while representing a genuinely different user intent. The best decoys reference sibling skills explicitly—for example, a prompt about "cases" when the current skill handles "methodology," ensuring the selector must distinguish conceptual boundaries rather than merely pattern‑match on vocabulary.

How strict is the pass rate requirement for decoy scenarios?

Decoy scenarios (should_not_trigger) operate under zero‑tolerance enforcement: any activation constitutes a failure regardless of the configured minimum_pass_rate. This absolute standard prevents subtle false positives that aggregate into unreliable multi‑skill agents. Only should_trigger and edge_case entries benefit from the 0.8 threshold flexibility.

Can I modify the test‑prompt schema for specialized skills?

The schema in templates/test‑prompts.json.template is version‑locked via "darwin_compatible": true to ensure pipeline compatibility. Custom extensions should preserve the three required categories (should_trigger, should_not_trigger, edge_case) and the minimum_pass_rate field. Consult [SKILL.md](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) for Darwin execution spec compliance before deviation.

Where does the pressure‑test stage fit in the RIA‑TV++ pipeline?

Stage 4 executes after skill extraction and packaging but before deployment eligibility. The pipeline feeds each test_cases entry to the LLM selector, compares actual against expected_behavior, and gates release on the pass‑rate criteria defined in [methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →