Understanding the darwin-skill Format for test-prompts.json

The darwin-skill format is a standardized JSON schema that defines how Instagit skills are tested, ensuring automated validation through structured test cases with specific activation and non-activation scenarios.

The test-prompts.json file sits at the heart of the skill validation pipeline in the kangarooking/cangjie-skill repository. This darwin-skill format enables the Darwin test runner to automatically verify that a skill behaves correctly across positive triggers, negative examples, and edge cases before deployment.

Structure of the darwin-skill JSON Schema

The darwin-skill format requires a specific top-level object structure that describes the skill metadata and its validation suite.

Top-Level Metadata Fields

According to the canonical template in templates/test-prompts.json.template, the JSON object must include these fields:

  • skill: The slug identifier for the skill (e.g., {{skill-slug}})
  • version: Semantic version tracking (typically "0.1.0")
  • source_book: Human-readable citation format ({{BOOK_TITLE}} — {{AUTHOR}})
  • darwin_compatible: Boolean flag (must be true) signaling that the file conforms to the Darwin test harness expectations
  • minimum_pass_rate: Decimal threshold (e.g., 0.8) representing the required success ratio across all should_trigger cases
  • notes: Author guidance regarding test case composition requirements
  • test_cases: Array of individual test scenario objects

Test Case Definitions

Each object within the test_cases array must specify:

  • id: Unique identifier for the test case (e.g., should-trigger-01)
  • type: Classification as should_trigger, should_not_trigger, or edge_case
  • prompt: The exact user utterance to feed the skill
  • expected_behavior: Detailed description of required skill response or explicit non-activation
  • notes: Optional rationale explaining the scenario's significance

Why the darwin-skill Format Is Necessary

The strict schema serves five critical functions in the Instagit ecosystem:

Standardized Validation: The darwin_compatible flag tells Instagit's Darwin test harness that the JSON complies with the expected schema, enabling automated loading and execution without manual parsing.

Deterministic Behavior Verification: By explicitly declaring should_trigger versus should_not_trigger cases, the format verifies activation logic against both positive and negative scenarios, preventing false positives in production.

Cross-Skill Guardrails: The template enforces at least one should_not_trigger case referencing a sibling skill (sibling-skill-slug), preventing accidental hijacking of other skills within the same book.

Quantitative Quality Gates: The minimum_pass_rate field provides a concrete metric; skills only pass validation when achieving the configured success threshold across mandatory test types.

Living Documentation: Fields such as source_book, per-case notes, and structured id values create maintainable audit trails for future skill updates.

Implementing the darwin-skill Format

Minimal test-prompts.json Example

{
  "skill": "my-awesome-skill",
  "version": "0.1.0",
  "source_book": "Great Book — Jane Doe",
  "darwin_compatible": true,
  "test_cases": [
    {
      "id": "should-trigger-01",
      "type": "should_trigger",
      "prompt": "给我推荐一本关于时间旅行的小说",
      "expected_behavior": "应激活 my-awesome-skill, 并 返回推荐列表",
      "notes": "正面场景:用户明确请求推荐"
    },
    {
      "id": "should-not-trigger-01",
      "type": "should_not_trigger",
      "prompt": "今天的天气怎么样?",
      "expected_behavior": "不应激活本 skill, 因为它是天气查询,不在本 skill 范畴",
      "notes": "诱饵:与本 skill 无关的对话"
    },
    {
      "id": "edge-01",
      "type": "edge_case",
      "prompt": "我想要一本适合儿童的时间旅行故事,但不要太科幻",
      "expected_behavior": "应返回适合儿童的时间旅行书目,说明为何符合筛选条件",
      "notes": "边界:混合需求需要合理判断"
    }
  ],
  "minimum_pass_rate": 0.8,
  "notes": "至少 2 条 should_trigger + 1 条 should_not_trigger + 1 条 edge_case。"
}

Validation with the Darwin Runner

While the actual Darwin runner is part of Instagit's internal tooling, the expected contract follows this pattern:

from darwin import TestRunner, load_skill_tests

# Load the JSON description

tests = load_skill_tests("path/to/test-prompts.json")

# Execute the suite

runner = TestRunner()
result = runner.run(tests)

print(f"Pass rate: {result.pass_rate:.2%}")
assert result.pass_rate >= tests["minimum_pass_rate"]

Key Files in the Repository

Several files work together to enforce the darwin-skill format:

  • templates/test-prompts.json.template: The canonical schema definition that all skill authors must follow
  • SKILL.md: Human-readable specification referencing the JSON validation requirements
  • README.md: Repository overview and contribution guidelines
  • methodology/07-stage5-deliver.md: Documentation for the final validation stage where Darwin tests execute

Summary

  • The darwin-skill format requires specific metadata fields including darwin_compatible: true and a minimum_pass_rate threshold
  • Test cases must cover three types: should_trigger, should_not_trigger, and edge_case
  • The schema prevents skill conflicts by requiring negative test cases against sibling skills
  • Validation occurs through Instagit's Darwin test harness, which parses the standardized JSON automatically
  • The template file at templates/test-prompts.json.template provides the authoritative reference implementation

Frequently Asked Questions

What is the darwin-skill format?

The darwin-skill format is a JSON schema used in the kangarooking/cangjie-skill repository to define automated test suites for Instagit skills. It specifies how to structure test cases, metadata, and success criteria so that the Darwin test runner can validate skill behavior automatically.

Why is the darwin_compatible field required?

The darwin_compatible boolean flag signals to Instagit's test infrastructure that the JSON file follows the expected schema. Without this field set to true, the Darwin runner will reject the file, preventing execution of the test suite and blocking deployment.

What are the required test case types in the darwin-skill format?

The format recognizes three mandatory test types: should_trigger for positive activation scenarios, should_not_trigger for negative examples (including sibling skill references), and edge_case for boundary testing. Authors must include at least two positive cases, one negative case, and one edge case per skill.

How does the minimum_pass_rate field work?

The minimum_pass_rate field defines a decimal threshold (typically 0.8 or 80%) that the skill must achieve across all should_trigger test cases. The Darwin runner calculates the actual pass rate during validation, and skills failing to meet this threshold are rejected from production deployment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →