Understanding the darwin-skill Format for test-prompts.json
The darwin-skill format is a standardized JSON schema that defines how Instagit skills are tested, ensuring automated validation through structured test cases with specific activation and non-activation scenarios.
The test-prompts.json file sits at the heart of the skill validation pipeline in the kangarooking/cangjie-skill repository. This darwin-skill format enables the Darwin test runner to automatically verify that a skill behaves correctly across positive triggers, negative examples, and edge cases before deployment.
Structure of the darwin-skill JSON Schema
The darwin-skill format requires a specific top-level object structure that describes the skill metadata and its validation suite.
Top-Level Metadata Fields
According to the canonical template in templates/test-prompts.json.template, the JSON object must include these fields:
- skill: The slug identifier for the skill (e.g.,
{{skill-slug}}) - version: Semantic version tracking (typically
"0.1.0") - source_book: Human-readable citation format (
{{BOOK_TITLE}} — {{AUTHOR}}) - darwin_compatible: Boolean flag (must be
true) signaling that the file conforms to the Darwin test harness expectations - minimum_pass_rate: Decimal threshold (e.g.,
0.8) representing the required success ratio across allshould_triggercases - notes: Author guidance regarding test case composition requirements
- test_cases: Array of individual test scenario objects
Test Case Definitions
Each object within the test_cases array must specify:
- id: Unique identifier for the test case (e.g.,
should-trigger-01) - type: Classification as
should_trigger,should_not_trigger, oredge_case - prompt: The exact user utterance to feed the skill
- expected_behavior: Detailed description of required skill response or explicit non-activation
- notes: Optional rationale explaining the scenario's significance
Why the darwin-skill Format Is Necessary
The strict schema serves five critical functions in the Instagit ecosystem:
Standardized Validation: The darwin_compatible flag tells Instagit's Darwin test harness that the JSON complies with the expected schema, enabling automated loading and execution without manual parsing.
Deterministic Behavior Verification: By explicitly declaring should_trigger versus should_not_trigger cases, the format verifies activation logic against both positive and negative scenarios, preventing false positives in production.
Cross-Skill Guardrails: The template enforces at least one should_not_trigger case referencing a sibling skill (sibling-skill-slug), preventing accidental hijacking of other skills within the same book.
Quantitative Quality Gates: The minimum_pass_rate field provides a concrete metric; skills only pass validation when achieving the configured success threshold across mandatory test types.
Living Documentation: Fields such as source_book, per-case notes, and structured id values create maintainable audit trails for future skill updates.
Implementing the darwin-skill Format
Minimal test-prompts.json Example
{
"skill": "my-awesome-skill",
"version": "0.1.0",
"source_book": "Great Book — Jane Doe",
"darwin_compatible": true,
"test_cases": [
{
"id": "should-trigger-01",
"type": "should_trigger",
"prompt": "给我推荐一本关于时间旅行的小说",
"expected_behavior": "应激活 my-awesome-skill, 并 返回推荐列表",
"notes": "正面场景:用户明确请求推荐"
},
{
"id": "should-not-trigger-01",
"type": "should_not_trigger",
"prompt": "今天的天气怎么样?",
"expected_behavior": "不应激活本 skill, 因为它是天气查询,不在本 skill 范畴",
"notes": "诱饵:与本 skill 无关的对话"
},
{
"id": "edge-01",
"type": "edge_case",
"prompt": "我想要一本适合儿童的时间旅行故事,但不要太科幻",
"expected_behavior": "应返回适合儿童的时间旅行书目,说明为何符合筛选条件",
"notes": "边界:混合需求需要合理判断"
}
],
"minimum_pass_rate": 0.8,
"notes": "至少 2 条 should_trigger + 1 条 should_not_trigger + 1 条 edge_case。"
}
Validation with the Darwin Runner
While the actual Darwin runner is part of Instagit's internal tooling, the expected contract follows this pattern:
from darwin import TestRunner, load_skill_tests
# Load the JSON description
tests = load_skill_tests("path/to/test-prompts.json")
# Execute the suite
runner = TestRunner()
result = runner.run(tests)
print(f"Pass rate: {result.pass_rate:.2%}")
assert result.pass_rate >= tests["minimum_pass_rate"]
Key Files in the Repository
Several files work together to enforce the darwin-skill format:
templates/test-prompts.json.template: The canonical schema definition that all skill authors must followSKILL.md: Human-readable specification referencing the JSON validation requirementsREADME.md: Repository overview and contribution guidelinesmethodology/07-stage5-deliver.md: Documentation for the final validation stage where Darwin tests execute
Summary
- The darwin-skill format requires specific metadata fields including
darwin_compatible: trueand aminimum_pass_ratethreshold - Test cases must cover three types:
should_trigger,should_not_trigger, andedge_case - The schema prevents skill conflicts by requiring negative test cases against sibling skills
- Validation occurs through Instagit's Darwin test harness, which parses the standardized JSON automatically
- The template file at
templates/test-prompts.json.templateprovides the authoritative reference implementation
Frequently Asked Questions
What is the darwin-skill format?
The darwin-skill format is a JSON schema used in the kangarooking/cangjie-skill repository to define automated test suites for Instagit skills. It specifies how to structure test cases, metadata, and success criteria so that the Darwin test runner can validate skill behavior automatically.
Why is the darwin_compatible field required?
The darwin_compatible boolean flag signals to Instagit's test infrastructure that the JSON file follows the expected schema. Without this field set to true, the Darwin runner will reject the file, preventing execution of the test suite and blocking deployment.
What are the required test case types in the darwin-skill format?
The format recognizes three mandatory test types: should_trigger for positive activation scenarios, should_not_trigger for negative examples (including sibling skill references), and edge_case for boundary testing. Authors must include at least two positive cases, one negative case, and one edge case per skill.
How does the minimum_pass_rate field work?
The minimum_pass_rate field defines a decimal threshold (typically 0.8 or 80%) that the skill must achieve across all should_trigger test cases. The Darwin runner calculates the actual pass rate during validation, and skills failing to meet this threshold are rejected from production deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →