# How Pressure Testing with Darwin-Skill Compatibility Works in Stage 4 of cangjie-skill

> Discover how Stage 4 of cangjie-skill utilizes pressure testing with darwin-skill compatibility to ensure skills activate accurately from public descriptions alone.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-08-15

---

**Pressure testing in Stage 4 uses a blind sub‑agent to verify that a skill activates correctly from its public description alone, requiring 100% pass rate for acceptance or ≥80% with documented fixes.**

Stage 4 — **Pressure Testing with Darwin‑Skill Compatibility** — is the final quality gate in the cangjie‑skill methodology before a skill ships. This stage validates **trigger precision** and **output quality** through a deliberately model‑agnostic approach that forces each skill to prove its activatability without relying on internal design knowledge.

## Why Darwin‑Skill Compatibility Testing Exists

The **trigger (A2)** is the most fragile component of any skill. Even a perfectly implemented skill fails if its trigger description is ambiguous or overlaps with sibling skills. According to the cangjie‑skill source code in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) (lines 9‑12), this pressure test is the **only pre‑release method** that surfaces trigger problems before users encounter them.

The core principle: simulate a completely **independent sub‑agent** that has never seen the skill's internal design. This sub‑agent must activate the skill *solely* from its public [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) description and a user prompt — exactly how real Darwin ecosystem agents will encounter it.

## The Test Data Schema: [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json)

Every skill must ship a `test‑prompts.json` file defined by the schema in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) (lines 26‑56) and instantiated from `templates/test-prompts.json.template`. The file contains three mandatory test categories:

| Category | Minimum Cases | Purpose |
|----------|---------------|---------|
| `should_trigger` | 3‑5 | Verify the skill activates on intended prompts (lines 60‑63) |
| `should_not_trigger` | 2‑3 | Verify the skill **does not** activate on distractor prompts (lines 62‑64) |
| `edge_case` | 1‑3 | Validate borderline decisions where activation is ambiguous (lines 62‑64) |

A **cross‑skill confusion test** is mandatory: at least one `should_not_trigger` case must use a prompt that *should* activate a different skill from the same book, preventing "skill‑swapping" bugs (lines 68‑69).

### Example Minimal Test Suite

```json
{
  "skill": "inversion-thinking",
  "version": "0.1.0",
  "test_cases": [
    {
      "id": "should-trigger-01",
      "type": "should_trigger",
      "prompt": "我要决定要不要接这个新项目,列了一堆好处但还是没底",
      "expected_behavior": "调用 inversion-thinking, 反问'最不希望发生什么'",
      "notes": "正面场景: 决策纠结"
    },
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "帮我查一下这个 API 的参数",
      "expected_behavior": "纯信息查询, 不应调用任何决策 skill",
      "notes": "诱饵: 非决策场景"
    },
    {
      "id": "edge-01",
      "type": "edge_case",
      "prompt": "我在想晚饭吃什么",
      "expected_behavior": "日常琐事, 不应调用 (虽然字面是'决策')",
      "notes": "边界: 区分严肃决策和日常选择"
    }
  ]
}

```

The full schema template lives at `templates/test-prompts.json.template` in the kangarooking/cangjie‑skill repository.

## Execution Flow of the Blind Sub‑Agent Test

The pressure test follows a strict five‑step execution sequence defined in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) (lines 72‑83):

### Step 1: Create [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json)

Populate the template with concrete skill name, prompts, expected behaviors, and notes for each test case (lines 72‑74).

### Step 2: Run the Blind Sub‑Agent

For every test case, the sub‑agent receives **only**:

- The skill's public description (from [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md))
- The user prompt
- Optional list of sibling skill names

The sub‑agent **does not see** `type`, `expected_behavior`, or `notes`. This simulates real‑world Darwin ecosystem conditions where agents have no internal knowledge of test expectations (lines 13‑22).

The sub‑agent must output three fields:
- `would_trigger` (yes/no)
- `reason` (explanation for the decision)
- `if_triggered_action` (the action it would take)

### Step 3: Evaluate Results Against Expected Behavior

Pass/fail rules are strict and type‑specific (lines 75‑77):

- **`should_trigger`** — sub‑agent must **explicitly** call the skill and perform the expected action
- **`should_not_trigger`** — sub‑agent must **not** call the skill; any activation is immediate failure (zero tolerance)
- **`edge_case`** — the sub‑agent's decision must align with the rationale in `expected_behavior`

### Step 4: Calculate Pass Rate

The acceptance thresholds are non‑negotiable (lines 80‑82):

- **100%** → skill accepted for delivery
- **≥80%** → analyze failures; decide whether to modify trigger description (A2) or adjust test cases
- **<80%** → **must redo Stage 2** (complete trigger redesign) — not a minor fix

### Step 5: Iterate Until Threshold Met

Fix identified issues and re‑run until the required pass rate is achieved (lines 82‑83).

If the execution environment lacks sub‑agent capability, the system falls back to self‑test mode and records lower‑confidence results in `test‑results.md` (lines 24‑25).

## Implementation: Blind Sub‑Agent Runner

The following Python skeleton demonstrates the **blind evaluation principle** — the runner never exposes test metadata to the agent under evaluation:

```python
import json
import subprocess
import pathlib

def run_sub_agent(skill_dir: str, prompt: str) -> dict:
    """Execute darwin-compatible sub-agent with public description only."""
    desc = pathlib.Path(skill_dir, "SKILL.md").read_text()
    
    result = subprocess.run(
        ["darwin-subagent", "--description", desc, "--prompt", prompt],
        capture_output=True,
        text=True
    )
    return json.loads(result.stdout)


def evaluate_skill(skill_dir: str) -> float:
    """Run pressure test and return pass rate."""
    test_file = pathlib.Path(skill_dir, "test-prompts.json")
    tests = json.loads(test_file.read_text())["test_cases"]
    
    passes = 0
    for tc in tests:
        out = run_sub_agent(skill_dir, tc["prompt"])
        
        if tc["type"] == "should_trigger":
            ok = out["would_trigger"] and out["if_triggered_action"] == tc["expected_behavior"]
        elif tc["type"] == "should_not_trigger":
            ok = not out["would_trigger"]
        else:  # edge_case

            ok = out["reason"] == tc["expected_behavior"]
        
        passes += ok
    
    return passes / len(tests)

```

The key design constraint: `tc["type"]` and `tc["expected_behavior"]` are **never passed** to the sub‑agent, ensuring genuine model‑agnostic validation.

## Deciding What to Fix

When test results fall below 100%, the methodology in [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) (lines 86‑88) specifies three diagnostic paths:

- **Ambiguous trigger description** → modify the skill's A2 component
- **Uncovered realistic scenario** → expand trigger design or add new test case
- **Over‑engineered bait prompt** → refine the test itself with documented rationale

## Outcome Artifacts

Stage 4 produces two mandatory deliverables per skill (lines 92‑94):

- `<skill-dir>/test-prompts.json` — the darwin‑compatible test suite
- `<skill-dir>/test-results.md` — audit log with pass rate and failure analysis

These artifacts enable reproducible quality verification and continuous integration in Darwin ecosystem deployments.

## Summary

- **Pressure testing with Darwin‑skill compatibility** validates triggers through blind sub‑agent simulation, ensuring skills activate correctly without internal knowledge
- The [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) schema requires **three test categories** with minimum case counts and mandatory cross‑skill confusion testing
- **Pass rate thresholds** are strict: 100% for acceptance, ≥80% with fixes, <80% requires returning to Stage 2
- The **blind evaluation principle** guarantees model‑agnostic quality: sub‑agents see only public [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) descriptions, never test expectations
- Outcome artifacts (`test‑prompts.json`, `test‑results.md`) provide auditable proof of Darwin ecosystem compatibility

## Frequently Asked Questions

### What makes the sub‑agent "blind" in Darwin‑skill pressure testing?

The sub‑agent is **blind** because it receives only the skill's public [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) description and the user prompt. It never sees the `type` field (whether a test case is `should_trigger`, `should_not_trigger`, or `edge_case`), the `expected_behavior`, or any `notes`. This forces genuine skill activation from description alone, exactly as real Darwin ecosystem agents would experience.

### Why is the 80% threshold a hard gate requiring Stage 2 redo?

A pass rate below 80% indicates **fundamental trigger design flaws**, not minor tuning issues. The cangjie‑skill methodology treats this as architectural debt: the A2 trigger component must be redesigned from scratch rather than patched, ensuring robust Darwin‑skill compatibility before release.

### How does cross‑skill confusion testing prevent deployment failures?

Cross‑skill confusion testing requires at least one `should_not_trigger` case that *would* activate a **different** skill from the same book. This catches "skill‑swapping" bugs where vague descriptions cause multiple skills to compete for the same prompt — a critical failure mode in multi‑skill Darwin ecosystem deployments.

### What happens when no sub‑agent capability exists in the environment?

The system falls back to **self‑test mode**, where the skill evaluates itself against its own test cases. This produces lower‑confidence results logged in `test‑results.md` with appropriate caveats. Full Darwin‑skill compatibility validation still requires eventual execution with an independent sub‑agent.