# How to Design Pressure Tests with Decoy Prompts That Should Not Trigger the Skill

> Design effective pressure tests for cangjie-skill using decoy prompts. Learn to create should_not_trigger cases for zero false positives and ensure accurate skill activation.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-08-16

---

**Pressure tests with decoy prompts verify that a cangjie-skill only activates when user intent truly matches the skill's trigger description, using "should_not_trigger" test cases that must have zero tolerance for false positives.**

The **cangjie-skill** framework, hosted at `kangarooking/cangjie-skill`, treats decoy prompts as the primary defense against **A2 trigger description** failures. These decoys—seemingly related prompts designed not to fire the skill—expose when a trigger description is too vague before the skill reaches production.

## Why Decoy Prompts Are Critical

The **A2 (future-trigger) description** is a single point of failure. If it overgeneralizes, the skill activates in unrelated contexts, creating noisy behavior. According to the methodology in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md), **stage 4 is the only pre-delivery mechanism** that can catch trigger-precision problems.

Decoy prompts serve two specific verification roles:

- **False positive prevention**: Ensure superficially similar queries don't erroneously trigger the skill
- **Cross-skill discrimination**: Verify the system correctly routes to sibling skills when context demands

## Core Principles for Decoy Design

Building effective decoys requires strict adherence to six principles defined in the pipeline:

| Principle | Implementation |
|-----------|---------------|
| **Independence** | Execute each test case in an isolated sub-agent with no memory of prior prompts |
| **Hidden Metadata** | Conceal `type`, `expected_behavior`, and `notes` fields from the sub-agent; expose only skill description, user prompt, and optional sibling skill list |
| **Required Output** | Sub-agent returns `would_trigger`, `reason`, and `if_triggered_action` for validation |
| **Decoy Ratio** | Include **2-3** `should_not_trigger` cases per skill |
| **Cross-skill Confusion** | **At least one decoy must trigger a different skill from the same book**—this is a hard requirement |
| **Pass Threshold** | Overall rate ≥ 0.8, but **decoys require 0% tolerance** for any false activation |

### The Zero-Tolerance Rule for Decoys

While the overall pass threshold allows 20% failure on `should_trigger` cases, **any decoy that incorrectly triggers the skill constitutes an automatic rejection**. This asymmetric strictness reflects the higher cost of false positives versus false negatives in production deployments.

## Test-Prompt File Structure

The repository provides a standardized template at `templates/test-prompts.json.template` that enforces decoy inclusion. A minimal decoy configuration demonstrates both required decoy types:

```json
{
  "skill": "inversion-thinking",
  "version": "0.1.0",
  "test_cases": [
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "帮我查一下这个 API 的参数",
      "expected_behavior": "不应激活本 skill, 因为它是纯信息查询",
      "notes": "诱饵: 信息检索类请求"
    },
    {
      "id": "should-not-trigger-02",
      "type": "should_not_trigger",
      "prompt": "我想找一本关于时间管理的书，推荐一本吧",
      "expected_behavior": "不应激活本 skill, 应激活 library-search skill",
      "notes": "跨 skill 混淆诱饵"
    }
  ],
  "minimum_pass_rate": 0.8
}

```

**`should-not-trigger-01`** demonstrates a **pure-information decoy**—topically adjacent but semantically distinct. **`should-not-trigger-02`** demonstrates **cross-skill confusion**, forcing discrimination between `inversion-thinking` and `library-search` within the same book.

### Field Requirements

- `type`: Must be `"should_not_trigger"` for decoys (distinct from `"should_trigger"`)
- `expected_behavior`: Precise description of correct system response for validation
- `notes`: Human-readable categorization aiding failure analysis

## Executing Decoy Pressure Tests

The four-stage execution flow operates as a closed validation loop:

1. **Generate** [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) from the template, ensuring 2-3 decoys with one cross-skill entry
2. **Blind-test** each case through a clean sub-agent with hidden metadata
3. **Compare** sub-agent output (`would_trigger`, `reason`, `if_triggered_action`) against `expected_behavior`
4. **Record** results in `<skill-dir>/test-results.md`; any decoy failure triggers refinement

### Programmatic Test Generation

Automate decoy creation with strict schema compliance:

```python
import json
from pathlib import Path

skill_slug = "inversion-thinking"

test_cases = [
    {
        "id": "should-not-trigger-01",
        "type": "should_not_trigger",
        "prompt": "帮我查一下这个 API 的参数",
        "expected_behavior": "不应激活本 skill, 因为它是纯信息查询",
        "notes": "诱饵: 信息检索类请求"
    },
    {
        "id": "should-not-trigger-02",
        "type": "should_not_trigger",
        "prompt": "我想找一本关于时间管理的书，推荐一本吧",
        "expected_behavior": "不应激活本 skill, 应激活 library-search skill",
        "notes": "跨 skill 混淆诱饵"
    },
    {
        "id": "should-not-trigger-03",
        "type": "should_not_trigger",
        "prompt": "用逆向思维分析一下这个商业案例",
        "expected_behavior": "不应激活本 skill, 应激活 general-analysis skill",
        "notes": "跨 skill 混淆诱饵: 分析方法重叠"
    }
]

payload = {
    "skill": skill_slug,
    "version": "0.1.0",
    "test_cases": test_cases,
    "minimum_pass_rate": 0.8
}

out_path = Path(f"books/sample/{skill_slug}/test-prompts.json")
out_path.parent.mkdir(parents=True, exist_ok=True)
out_path.write_text(json.dumps(payload, ensure_ascii=False, indent=2))
print(f"Created {out_path}")

```

Note: The pipeline requires **at least three `should_trigger` cases** alongside these decoys. Add positive cases identically with `"type": "should_trigger"`.

### Running the Pressure Test

Invoke the blind-test pipeline:

```bash
cangjie-skill run-pressure-test books/sample/inversion-thinking

```

The command processes each test case through an isolated sub-agent, hiding metadata fields as specified in [`06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/06-stage4-pressure-test.md).

## Responding to Decoy Failures

Failure analysis distinguishes three remediation paths:

- **Trigger description ambiguity** → Refine A2 description in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) (skill-level change)
- **Plausible unhandled scenario** → Expand skill capabilities (skill-level change)
- **Unrealistic decoy construction** → Adjust test case (test-level change)

The "回炉淘汰" (reject-and-refine) decision applies when decoy failures indicate fundamental trigger imprecision.

## Summary

- **Decoy prompts** (`should_not_trigger` cases) are mandatory 2-3 per skill with **zero tolerance** for false activation
- **Cross-skill confusion decoys** are required—at least one must correctly trigger a sibling skill instead
- **Metadata hiding** ensures the sub-agent judges purely on description and prompt, not expected outcomes
- **Automatic rejection** occurs if any decoy triggers; overall 0.8 pass rate applies only to positive cases
- **Source files** governing this workflow: [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md), `templates/test-prompts.json.template`, and per-skill [`test-results.md`](https://github.com/kangarooking/cangjie-skill/blob/main/test-results.md)

## Frequently Asked Questions

### What makes a decoy prompt effective versus ineffective?

An effective decoy sits in the **semantic adjacent zone**—superficially related vocabulary and domain, but substantively different user intent. Ineffective decoys are either too obviously unrelated (no discrimination value) or so subtly different that humans disagree on correct classification. The methodology recommends testing decoy clarity with human reviewers before pipeline inclusion.

### Why must cross-skill confusion decoys trigger a different skill rather than trigger nothing?

Requiring activation of a **sibling skill** rather than null response tests the **routing precision** of the entire book. This prevents skills from "shadowing" each other through vague descriptions and validates that the A2 trigger descriptions are mutually exclusive where intended. Per [`06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/06-stage4-pressure-test.md), this is labeled "跨 skill 混淆测试" and constitutes a hard gate.

### How does the sub-agent isolation prevent test contamination?

Each test case executes in a **fresh sub-agent context** with no conversational memory, no access to prior test results, and no visibility of the `type` and `expected_behavior` fields. This design, specified in the independence principle, prevents the sub-agent from learning test patterns or inferring desired outcomes from metadata leakage.

### Can a skill pass with failing decoys if overall accuracy exceeds 80%?

**No.** The `minimum_pass_rate: 0.8` applies exclusively to `should_trigger` cases. Decoys (`should_not_trigger`) operate under **0% tolerance**—any single false positive constitutes skill rejection. This asymmetric threshold reflects the production cost of erroneous skill activation versus missed activation opportunities.