# How Pressure Testing with Darwin-Skill Compatible test-prompts.json Works in Cangjie-Skill

> Learn how pressure testing with Darwin-skill compatible test-prompts.json validates Cangjie-Skill trigger logic. Ensure accuracy before delivery with blind sub-agent evaluation.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-08-13

---

**Pressure testing validates that a skill's trigger logic fires only when appropriate by using a Darwin-compatible [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) file and blind sub-agent evaluation before delivery.**

Pressure testing is the **fourth stage** of the **cangjie-skill** development pipeline. It ensures a skill's **trigger logic** (defined in the A2 field) behaves correctly under realistic conditions. The process depends on a Darwin-compatible [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) structure and **independent sub-agents** that evaluate test cases without prior knowledge of expected outcomes.

## What Makes a test-prompts.json Darwin-Skill Compatible

The Darwin-compatible format follows a strict schema designed for **blind evaluation**. You create this file from the template at `templates/test-prompts.json.template`.

A valid [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) contains:

- **`skill`** – the skill's slug identifier
- **`version`** – semantic version of the test suite
- **`test_cases`** – an array of evaluation scenarios, each with:
  - `id` – unique case identifier
  - `type` – one of `should_trigger`, `should_not_trigger`, or `edge_case`
  - `prompt` – the user utterance to test
  - `expected_behavior` – what correct behavior looks like
  - `notes` – context for human reviewers (hidden from sub-agents)

The schema is documented in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md).

## The Five Steps of Pressure Testing

### 1. Create the Test File

Fill the template with **three categories** of test cases:

- **Should-trigger cases** – typical user utterances that must activate the skill
- **Should-not-trigger cases** – "bait" prompts that test for false positives
- **Edge cases** – ambiguous inputs where the decision requires nuanced reasoning

Include **cross-skill prompts** if sibling skills exist in the same book to prevent activation conflicts.

### 2. Run Blind Sub-Agent Evaluations

For each `test_case`, spin up a **clean sub-agent** with strictly limited information:

```python
def evaluate_prompt(skill_path, prompt, sibling_skills=None):
    # Load skill description but hide trigger details

    skill_meta = load_skill_meta(skill_path, hide=['type','expected_behavior','notes'])
    # Run the sub-agent with only the prompt and optional sibling list

    result = sub_agent.run(
        skill_meta=skill_meta,
        user_prompt=prompt,
        sibling_skills=sibling_skills,
    )
    return {
        "would_trigger": result.would_trigger,
        "reason": result.reason,
        "if_triggered_action": result.action,
    }

```

**Critical constraint:** The sub-agent receives **only**:
- The skill's path and content (metadata without evaluation fields)
- The user prompt
- Optional sibling skill list

The `type`, `expected_behavior`, and `notes` fields are **intentionally withheld** to eliminate bias.

### 3. Compare Results to Expected Behavior

The main workflow matches sub-agent output against [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) expectations:

| Case Type | Passing Criteria |
|-----------|----------------|
| `should_trigger` | Sub-agent triggers skill **and** performs expected action |
| `should_not_trigger` | Sub-agent **does not trigger** (zero tolerance for false positives) |
| `edge_case` | Decision aligns with reasoning in `expected_behavior` |

### 4. Calculate Pass Rate and Decide

- **100% pass** → Skill accepted for delivery
- **≥80% pass** → Investigate failures; decide whether to refine A2 trigger description or adjust test cases
- **<80% pass** → Return to **stage 2** for complete trigger logic redesign

### 5. Iterate Until Threshold Met

After any fix, re-run blind tests. Update [`test-results.md`](https://github.com/kangarooking/cangjie-skill/blob/main/test-results.md) in the skill directory to track progress.

## Example test-prompts.json Structure

```json
{
  "skill": "inversion-thinking",
  "version": "0.1.0",
  "test_cases": [
    {
      "id": "should-trigger-01",
      "type": "should_trigger",
      "prompt": "我要决定要不要接这个新项目,列了一堆好处但还是没底",
      "expected_behavior": "调用 inversion‑thinking, 反问 '最不希望发生什么'",
      "notes": "正面场景: 决策纠结"
    },
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "帮我查一下这个 API 的参数",
      "expected_behavior": "纯信息查询, 不应调用任何决策 skill",
      "notes": "诱饵: 非决策场景"
    },
    {
      "id": "edge-01",
      "type": "edge_case",
      "prompt": "我在想晚饭吃什么",
      "expected_behavior": "日常琐事, 不应调用 (虽然字面是'决策')",
      "notes": "边界: 区分严肃决策和日常选择"
    }
  ]
}

```

This example demonstrates **defense in depth**: positive testing, negative testing, and boundary analysis in a single suite.

## Why Blind Sub-Agent Evaluation Matters

The trigger (A2) is the **most fragile component** of any skill. Without pressure testing, a skill can appear functional in development while silently **mis-triggering in production**.

Independent sub-agents mimic real-world deployment where:
- No prior knowledge of "correct" answers exists
- Users provide unpredictable, context-light prompts
- Similar skills compete for activation

Including **should-not-trigger** cases prevents "always-trigger" shortcuts in the A2 description. Cross-skill cases ensure **activation discrimination** when multiple skills from the same book are available.

## Key Files in the Pressure Testing Pipeline

| File Path | Purpose |
|-----------|---------|
| [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) | Complete methodology, pass criteria, and workflow documentation |
| `templates/test-prompts.json.template` | Boilerplate schema for creating skill-specific test files |
| [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) | Per-skill location where [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) resides and [`test-results.md`](https://github.com/kangarooking/cangjie-skill/blob/main/test-results.md) is generated |
| [`methodology/07-stage5-deliver.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md) | Post-testing stage for final digest generation and deployment |

## Summary

- **Pressure testing with Darwin-skill compatible [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json)** validates trigger logic through blind sub-agent evaluation before production.
- The **three test case types** (`should_trigger`, `should_not_trigger`, `edge_case`) provide comprehensive coverage of activation scenarios.
- **Zero tolerance** for false positives in `should_not_trigger` cases prevents over-triggering.
- The **80% threshold** determines whether a skill proceeds, iterates, or returns for redesign.
- All test artifacts follow schemas defined in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) and `templates/test-prompts.json.template`.

## Frequently Asked Questions

### What information does the sub-agent receive during pressure testing?

The sub-agent receives **only** the skill's metadata (with evaluation fields hidden), the user prompt, and an optional sibling skill list. The `type`, `expected_behavior`, and `notes` fields from [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) are **deliberately excluded** to ensure unbiased evaluation that mirrors real-world usage.

### Why is the pass threshold set at 80% rather than 100%?

The 100% threshold is required for **immediate acceptance**. The 80% threshold allows for **investigation and iteration**—some failures may indicate test case flaws rather than skill defects. Below 80%, the pattern of failures suggests fundamental trigger logic problems requiring stage 2 redesign.

### How does pressure testing differ from earlier verification stages?

Earlier stages (stage 1-3) verify **functional correctness** and **output quality** with full knowledge of expected behavior. Pressure testing in **stage 4** specifically validates **trigger discrimination** using **blind evaluation**—the sub-agent has no access to expected answers, preventing confirmation bias in the assessment.

### Where does the Darwin-skill compatibility requirement come from?

The [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) format is **Darwin-compatible**, meaning it adheres to schema conventions used across the Darwin skill ecosystem. This ensures interoperability with shared tooling and evaluation infrastructure. The template at `templates/test-prompts.json.template` enforces this compatibility for all cangjie-skill implementations.