How Pressure Testing with Darwin-Skill Compatibility Works in Stage 4 of cangjie-skill

Pressure testing in Stage 4 uses a blind sub‑agent to verify that a skill activates correctly from its public description alone, requiring 100% pass rate for acceptance or ≥80% with documented fixes.

Stage 4 — Pressure Testing with Darwin‑Skill Compatibility — is the final quality gate in the cangjie‑skill methodology before a skill ships. This stage validates trigger precision and output quality through a deliberately model‑agnostic approach that forces each skill to prove its activatability without relying on internal design knowledge.

Why Darwin‑Skill Compatibility Testing Exists

The trigger (A2) is the most fragile component of any skill. Even a perfectly implemented skill fails if its trigger description is ambiguous or overlaps with sibling skills. According to the cangjie‑skill source code in methodology/06-stage4-pressure-test.md (lines 9‑12), this pressure test is the only pre‑release method that surfaces trigger problems before users encounter them.

The core principle: simulate a completely independent sub‑agent that has never seen the skill's internal design. This sub‑agent must activate the skill solely from its public SKILL.md description and a user prompt — exactly how real Darwin ecosystem agents will encounter it.

The Test Data Schema: test-prompts.json

Every skill must ship a test‑prompts.json file defined by the schema in methodology/06-stage4-pressure-test.md (lines 26‑56) and instantiated from templates/test-prompts.json.template. The file contains three mandatory test categories:

Category Minimum Cases Purpose
should_trigger 3‑5 Verify the skill activates on intended prompts (lines 60‑63)
should_not_trigger 2‑3 Verify the skill does not activate on distractor prompts (lines 62‑64)
edge_case 1‑3 Validate borderline decisions where activation is ambiguous (lines 62‑64)

A cross‑skill confusion test is mandatory: at least one should_not_trigger case must use a prompt that should activate a different skill from the same book, preventing "skill‑swapping" bugs (lines 68‑69).

Example Minimal Test Suite

{
  "skill": "inversion-thinking",
  "version": "0.1.0",
  "test_cases": [
    {
      "id": "should-trigger-01",
      "type": "should_trigger",
      "prompt": "我要决定要不要接这个新项目,列了一堆好处但还是没底",
      "expected_behavior": "调用 inversion-thinking, 反问'最不希望发生什么'",
      "notes": "正面场景: 决策纠结"
    },
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "帮我查一下这个 API 的参数",
      "expected_behavior": "纯信息查询, 不应调用任何决策 skill",
      "notes": "诱饵: 非决策场景"
    },
    {
      "id": "edge-01",
      "type": "edge_case",
      "prompt": "我在想晚饭吃什么",
      "expected_behavior": "日常琐事, 不应调用 (虽然字面是'决策')",
      "notes": "边界: 区分严肃决策和日常选择"
    }
  ]
}

The full schema template lives at templates/test-prompts.json.template in the kangarooking/cangjie‑skill repository.

Execution Flow of the Blind Sub‑Agent Test

The pressure test follows a strict five‑step execution sequence defined in methodology/06-stage4-pressure-test.md (lines 72‑83):

Step 1: Create test-prompts.json

Populate the template with concrete skill name, prompts, expected behaviors, and notes for each test case (lines 72‑74).

Step 2: Run the Blind Sub‑Agent

For every test case, the sub‑agent receives only:

  • The skill's public description (from SKILL.md)
  • The user prompt
  • Optional list of sibling skill names

The sub‑agent does not see type, expected_behavior, or notes. This simulates real‑world Darwin ecosystem conditions where agents have no internal knowledge of test expectations (lines 13‑22).

The sub‑agent must output three fields:

  • would_trigger (yes/no)
  • reason (explanation for the decision)
  • if_triggered_action (the action it would take)

Step 3: Evaluate Results Against Expected Behavior

Pass/fail rules are strict and type‑specific (lines 75‑77):

  • should_trigger — sub‑agent must explicitly call the skill and perform the expected action
  • should_not_trigger — sub‑agent must not call the skill; any activation is immediate failure (zero tolerance)
  • edge_case — the sub‑agent's decision must align with the rationale in expected_behavior

Step 4: Calculate Pass Rate

The acceptance thresholds are non‑negotiable (lines 80‑82):

  • 100% → skill accepted for delivery
  • ≥80% → analyze failures; decide whether to modify trigger description (A2) or adjust test cases
  • <80% → must redo Stage 2 (complete trigger redesign) — not a minor fix

Step 5: Iterate Until Threshold Met

Fix identified issues and re‑run until the required pass rate is achieved (lines 82‑83).

If the execution environment lacks sub‑agent capability, the system falls back to self‑test mode and records lower‑confidence results in test‑results.md (lines 24‑25).

Implementation: Blind Sub‑Agent Runner

The following Python skeleton demonstrates the blind evaluation principle — the runner never exposes test metadata to the agent under evaluation:

import json
import subprocess
import pathlib

def run_sub_agent(skill_dir: str, prompt: str) -> dict:
    """Execute darwin-compatible sub-agent with public description only."""
    desc = pathlib.Path(skill_dir, "SKILL.md").read_text()
    
    result = subprocess.run(
        ["darwin-subagent", "--description", desc, "--prompt", prompt],
        capture_output=True,
        text=True
    )
    return json.loads(result.stdout)


def evaluate_skill(skill_dir: str) -> float:
    """Run pressure test and return pass rate."""
    test_file = pathlib.Path(skill_dir, "test-prompts.json")
    tests = json.loads(test_file.read_text())["test_cases"]
    
    passes = 0
    for tc in tests:
        out = run_sub_agent(skill_dir, tc["prompt"])
        
        if tc["type"] == "should_trigger":
            ok = out["would_trigger"] and out["if_triggered_action"] == tc["expected_behavior"]
        elif tc["type"] == "should_not_trigger":
            ok = not out["would_trigger"]
        else:  # edge_case

            ok = out["reason"] == tc["expected_behavior"]
        
        passes += ok
    
    return passes / len(tests)

The key design constraint: tc["type"] and tc["expected_behavior"] are never passed to the sub‑agent, ensuring genuine model‑agnostic validation.

Deciding What to Fix

When test results fall below 100%, the methodology in methodology/06-stage4-pressure-test.md (lines 86‑88) specifies three diagnostic paths:

  • Ambiguous trigger description → modify the skill's A2 component
  • Uncovered realistic scenario → expand trigger design or add new test case
  • Over‑engineered bait prompt → refine the test itself with documented rationale

Outcome Artifacts

Stage 4 produces two mandatory deliverables per skill (lines 92‑94):

  • <skill-dir>/test-prompts.json — the darwin‑compatible test suite
  • <skill-dir>/test-results.md — audit log with pass rate and failure analysis

These artifacts enable reproducible quality verification and continuous integration in Darwin ecosystem deployments.

Summary

  • Pressure testing with Darwin‑skill compatibility validates triggers through blind sub‑agent simulation, ensuring skills activate correctly without internal knowledge
  • The test-prompts.json schema requires three test categories with minimum case counts and mandatory cross‑skill confusion testing
  • Pass rate thresholds are strict: 100% for acceptance, ≥80% with fixes, <80% requires returning to Stage 2
  • The blind evaluation principle guarantees model‑agnostic quality: sub‑agents see only public SKILL.md descriptions, never test expectations
  • Outcome artifacts (test‑prompts.json, test‑results.md) provide auditable proof of Darwin ecosystem compatibility

Frequently Asked Questions

What makes the sub‑agent "blind" in Darwin‑skill pressure testing?

The sub‑agent is blind because it receives only the skill's public SKILL.md description and the user prompt. It never sees the type field (whether a test case is should_trigger, should_not_trigger, or edge_case), the expected_behavior, or any notes. This forces genuine skill activation from description alone, exactly as real Darwin ecosystem agents would experience.

Why is the 80% threshold a hard gate requiring Stage 2 redo?

A pass rate below 80% indicates fundamental trigger design flaws, not minor tuning issues. The cangjie‑skill methodology treats this as architectural debt: the A2 trigger component must be redesigned from scratch rather than patched, ensuring robust Darwin‑skill compatibility before release.

How does cross‑skill confusion testing prevent deployment failures?

Cross‑skill confusion testing requires at least one should_not_trigger case that would activate a different skill from the same book. This catches "skill‑swapping" bugs where vague descriptions cause multiple skills to compete for the same prompt — a critical failure mode in multi‑skill Darwin ecosystem deployments.

What happens when no sub‑agent capability exists in the environment?

The system falls back to self‑test mode, where the skill evaluates itself against its own test cases. This produces lower‑confidence results logged in test‑results.md with appropriate caveats. Full Darwin‑skill compatibility validation still requires eventual execution with an independent sub‑agent.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →