# How to Design Effective Test Prompts with Decoy Scenarios for cangjie‑skill

> Design effective test prompts with decoy scenarios for cangjie-skill using a three-category JSON schema. Pressure-test AI skills and prevent false positives with misleading prompts.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: best-practices
- Published: 2026-08-15

---

**Effective test prompts with decoy scenarios in cangjie‑skill use a three‑category JSON schema (`should_trigger`, `should_not_trigger`, `edge_case`) to pressure‑test AI Skills and prevent false positives through deliberately misleading prompts that share keywords with sibling skills.**

The **cangjie‑skill** framework converts extracted knowledge‑units—principles, frameworks, and cases—into **AI Skills** that agents can invoke. To ensure each skill fires only in its intended context, the project implements a rigorous **pressure‑test stage** based on the **RIA‑TV++ methodology**. Central to this validation is a structured JSON test‑prompt file that embeds **decoy scenarios** designed to trip up imprecise trigger logic.

## Understanding the Test‑Prompt Schema

The canonical template resides at **[`templates/test‑prompts.json.template`](https://github.com/kangarooking/cangjie-skill/blob/main/templates/test-prompts.json.template)**. Every generated skill pack includes a populated version of this file with three distinct test families:

| Test Type | Purpose | Example Prompt Pattern |
|-----------|---------|------------------------|
| **`should_trigger`** | Confirm positive activation on legitimate user intent | `"{{用户的典型话语 1 — 正面场景}}"` |
| **`should_not_trigger`** | **Decoy scenarios** that resemble valid triggers but must not fire | `"{{诱饵 1 — 看似相关但实际不该调用}}"` |
| **`edge_case`** | Borderline inputs forcing explicit decision explanations | `"{{边界模糊场景}}"` |

This schema enforces **zero tolerance for decoy failures** while allowing an 80% minimum pass rate for edge cases, as documented in **[[`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)**.

## Why Decoy Scenarios Are Critical

Decoy scenarios serve two essential functions during **Stage 4 – Pressure Test**:

- **Prevent false positives** — They replicate phrasing patterns that could activate sibling skills, forcing the selector’s trigger logic to distinguish subtle intent differences
- **Enforce skill isolation** — By requiring `minimum_pass_rate` ≥ 0.8 with *absolute* decoy compliance, the system prevents skill hijacking in multi‑agent deployments

This approach aligns with the cangjie‑skill design principle: **"only expose what truly belongs to the skill."** The methodology explicitly references cross‑skill discrimination through sibling‑skill slugs in decoy definitions.

## Building Robust Test Prompts: A 5‑Step Framework

Follow this structured workflow when creating **[`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json)** for a new skill:

1. **Extract core trigger phrases** from the source material (the `R` component in RIA‑TV++)
2. **Draft three positive examples** — Vary vocabulary while preserving underlying user intent
3. **Construct two strategic decoys**:
   - **Decoy 1**: Shares keywords but maps to a *different concept* within the same source book
   - **Decoy 2**: Would correctly trigger a *sibling skill*, testing cross‑skill boundary precision
4. **Add one edge‑case** — Force explanatory reasoning (e.g., conditional applicability queries)
5. **Configure `minimum_pass_rate`** — The standard 0.8 threshold balances strictness with practical variance

## Complete Test‑Prompt Example

Below is a fully populated **[`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json)** for a hypothetical `example-skill`:

```json
{
  "skill": "example-skill",
  "version": "0.1.0",
  "source_book": "The Example Book — Jane Doe",
  "darwin_compatible": true,
  "test_cases": [
    {
      "id": "should-trigger-01",
      "type": "should_trigger",
      "prompt": "请告诉我如何在实际项目中使用《示例方法》",
      "expected_behavior": "应激活 example-skill, 并提供实施步骤",
      "notes": "正面场景：直接请求使用方法"
    },
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "我想了解《示例方法》中的相关案例",
      "expected_behavior": "不应激活本 skill, 因为该请求应由 example-case-skill 处理",
      "notes": "跨 skill 混淆诱饵"
    },
    {
      "id": "edge-01",
      "type": "edge_case",
      "prompt": "如果项目规模很小，是否仍需要使用《示例方法》？",
      "expected_behavior": "说明方法在小规模项目中的适用性或限制",
      "notes": "边界模糊场景"
    }
  ],
  "minimum_pass_rate": 0.8,
  "notes": "3 条 should_trigger + 1 条 should_not_trigger + 1 条 edge_case"
}

```

## Validating Test Results Programmatically

While the cangjie‑skill pipeline automates evaluation, this Python snippet illustrates the validation logic from **[[`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)** specifications:

```python
import json

def evaluate_results(results, tmpl):
    """Check if test outcomes meet minimum_pass_rate threshold."""
    passed = sum(1 for r in results if r['passed'])
    rate = passed / len(results)
    return rate >= tmpl['minimum_pass_rate']

with open('test-prompts.json') as f:
    tmpl = json.load(f)

# Mock results for illustration (actual pipeline queries LLM)

mock_results = [{'id': tc['id'], 'passed': True} for tc in tmpl['test_cases']]

if evaluate_results(mock_results, tmpl):
    print("All tests passed – skill is ready for deployment")
else:
    print("Some tests failed – revisit prompts or trigger logic")

```

## Key Implementation Files

| File Path | Role in Test‑Prompt Design |
|-----------|---------------------------|
| [`templates/test‑prompts.json.template`](https://github.com/kangarooking/cangjie-skill/blob/main/templates/test-prompts.json.template) | Master schema defining all three test categories with decoy requirements |
| [[`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) | Pressure‑test stage documentation, acceptance criteria, and zero‑tolerance rationale |
| [[`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) | Execution spec linking test‑prompt validation to deployment gates |
| `buffett-letters-skill/` (example pack) | Production skill demonstrating populated test prompts in context |

## Summary

- **Decoy scenarios** in `should_not_trigger` entries are non‑negotiable for preventing skill collisions
- The **`minimum_pass_rate`** of 0.8 applies globally, but decoys must achieve 100% negative accuracy
- Test‑prompt files follow a strict JSON schema from **`templates/test‑prompts.json.template`**
- Cross‑skill discrimination is validated through sibling‑skill decoys per **RIA‑TV++ Stage 4**
- Real‑world examples exist in the **Buffett Letters** skill pack and other repository references

## Frequently Asked Questions

### What makes a decoy scenario effective in cangjie‑skill?

An effective decoy shares surface‑level keywords with the target skill while representing a genuinely different user intent. The best decoys reference sibling skills explicitly—for example, a prompt about "cases" when the current skill handles "methodology," ensuring the selector must distinguish conceptual boundaries rather than merely pattern‑match on vocabulary.

### How strict is the pass rate requirement for decoy scenarios?

Decoy scenarios (`should_not_trigger`) operate under **zero‑tolerance enforcement**: any activation constitutes a failure regardless of the configured `minimum_pass_rate`. This absolute standard prevents subtle false positives that aggregate into unreliable multi‑skill agents. Only `should_trigger` and `edge_case` entries benefit from the 0.8 threshold flexibility.

### Can I modify the test‑prompt schema for specialized skills?

The schema in **`templates/test‑prompts.json.template`** is version‑locked via `"darwin_compatible": true` to ensure pipeline compatibility. Custom extensions should preserve the three required categories (`should_trigger`, `should_not_trigger`, `edge_case`) and the `minimum_pass_rate` field. Consult **[[`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)** for Darwin execution spec compliance before deviation.

### Where does the pressure‑test stage fit in the RIA‑TV++ pipeline?

Stage 4 executes after skill extraction and packaging but before deployment eligibility. The pipeline feeds each `test_cases` entry to the LLM selector, compares actual against `expected_behavior`, and gates release on the pass‑rate criteria defined in **[[`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)**.