Cangjie-Skill Distillation Output Structure: A Complete Directory Guide

Cangjie-skill distillation produces a structured directory of agent-ready skill packs containing markdown files, JSON test assets, and organized sub-folders for seamless integration into AI agent systems.

The Cangjie-skill framework transforms entire books into reusable, executable knowledge units through a multi-stage pipeline. According to the kangarooking/cangjie-skill source code, the final output follows a rigorous directory hierarchy designed for immediate deployment in darwin-skill compatible environments. Understanding this structure enables developers to navigate, validate, and integrate distilled book knowledge efficiently.

Root-Level Output Files

After the distillation pipeline completes Stage 5 (delivery), the books/<book-slug>/ directory contains four mandatory markdown documents that provide top-level context and navigation.

BOOK_OVERVIEW.md

BOOK_OVERVIEW.md sits at the repository root and delivers a high-level synthesis of the source material. This file captures the book's structural architecture, main argumentative threads, and critical evaluation notes generated during the distillation process.

The overview serves as the entry point for human reviewers before they deep-dive into specific skills. It references the methodology pipeline documented in [methodology/00-overview.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md).

DIGEST.md

DIGEST.md contains a long-form "essence" article that stitches together vetted skills with contextual commentary. Unlike individual skill files that isolate specific capabilities, the digest presents interconnected knowledge as a cohesive narrative.

The digest follows the template defined in templates/DIGEST.md.template, ensuring consistent formatting across different book distillations.

INDEX.md

INDEX.md provides a navigable directory of all extracted skills, organized thematically with dependency graph visualizations. This file transforms the flat directory structure into a semantic map of the book's actionable knowledge.

The index template at templates/INDEX.md.template includes an "安装使用" (installation and usage) section that specifies how to copy the distillation output into a host agent's skills/ folder.

GLOSSARY.md

GLOSSARY.md contains an alphabetized terminology dictionary extracted by a dedicated glossary extractor during the pipeline. This ensures consistent vocabulary usage across all skills derived from the source book.

Skill Package Directories

The core output of Cangjie-skill distillation resides in individual <skill-slug>/ folders—each representing a discrete, agent-executable capability unit.

SKILL.md Structure

Every skill directory contains a SKILL.md file implementing the RIA-TV++ structure (Reading, Interpretation, Application-1, Application-2, Trigger, Execution, Boundary). This format originates from the template at templates/SKILL.md.template.

The SKILL.md embeds critical audit metadata including:

  • Validation status — whether the skill passed triple-verification
  • Test-prompt coverage — which scenarios have been pressure-tested
  • Reading — source context and prerequisites
  • Interpretation — conceptual analysis of the knowledge
  • Application — two-tier implementation guidance (theory and practice)
  • Trigger — activation conditions for the agent
  • Execution — concrete, step-by-step procedures
  • Boundary — limitations and contraindications

test-prompts.json

Each skill directory includes test-prompts.json, a darwin-skill compatible test-prompt set for automated pressure-testing and evolutionary refinement. These JSON assets enable continuous validation against edge cases and adversarial inputs.

Processing Directories

The output structure preserves visibility into the distillation pipeline's intermediate states through dedicated folders.

candidates/

The candidates/ directory stores the raw extraction pool produced during Stage 1 (parallel extractors). This folder contains unvetted knowledge units before they undergo the triple-verification process described in Stage 1.5.

rejected/

The rejected/ directory holds units that failed verification or were manually flagged as unsuitable. Preserving rejected candidates supports audit trails, iterative improvement, and potential re-evaluation with refined extraction parameters.

Static Assets

The assets/ folder contains supporting resources referenced by markdown files throughout the output structure. This includes:

  • Diagrams and visualizations
  • Font files for consistent rendering
  • Extracted images from the source book
  • Pipeline-generated graphics

Consolidated Test Infrastructure

The root-level test-prompts.json (generated per book) aggregates all individual skill test prompts into a single collection. This consolidation enables:

  • Batch validation across the entire skill set
  • darwin-skill integration for automatic evolution
  • Regression testing when source books are re-distilled

Complete Directory Hierarchy

books/<book-slug>/
├─ BOOK_OVERVIEW.md          # High-level book synthesis

├─ DIGEST.md                 # Long-form essence article

├─ INDEX.md                  # Navigable skill directory

├─ GLOSSARY.md               # Terminology dictionary

├─ test-prompts.json         # Consolidated test collection

├─ candidates/               # Raw extraction pool (Stage 1)

├─ rejected/                 # Failed verification units (Stage 1.5)

├─ <skill-slug-1>/           # Individual skill package

│   ├─ SKILL.md              # RIA-TV++ structured skill

│   └─ test-prompts.json     # darwin-skill compatible tests

├─ <skill-slug-2>/           # Additional skill packages...

│   ├─ SKILL.md
│   └─ test-prompts.json
└─ assets/                   # Static resources

Loading and Deploying Distilled Output

Agent Integration (Python)

import yaml

# Assume distilled book at ./books/lean-startup/

skill_dir = "./books/lean-startup/lean-startup-mvp"
with open(f"{skill_dir}/SKILL.md") as f:
    skill_yaml = yaml.safe_load(f)

# Access structured skill components

description = skill_yaml["description"]
execution_steps = skill_yaml["E — 可执行步骤"]
trigger_conditions = skill_yaml["T — 触发条件"]

Full Deployment (Bash)


# From cangjie-skill repository root

cp -r books/lean-startup/* ~/.claude/skills/

# Verify darwin-skill compatibility

darwin-skill validate ~/.claude/skills/lean-startup/

Source File References

The output structure implementation spans these key repository files:

  • methodology/00-overview.md — Visual pipeline and stage definitions
  • templates/SKILL.md.template — Canonical skill file structure
  • templates/DIGEST.md.template — Long-form digest layout
  • templates/INDEX.md.template — Thematic index with dependency graphs

Summary

  • Cangjie-skill distillation produces a books/<book-slug>/ directory containing four root markdown files, processing folders, and individual skill packages
  • SKILL.md files implement the RIA-TV++ structure with embedded validation metadata and audit fields
  • test-prompts.json assets provide darwin-skill compatible automated testing at both skill and book levels
  • candidates/ and rejected/ folders preserve pipeline transparency for quality assurance and iteration
  • The complete output can be copied directly to ~/.claude/skills/ for immediate agent deployment according to the installation instructions in INDEX.md.template

Frequently Asked Questions

What does the RIA-TV++ structure in SKILL.md contain?

The RIA-TV++ structure organizes executable knowledge into six sections: Reading (source context), Interpretation (conceptual analysis), Application (two-tier theory/practice guidance), Trigger (activation conditions), Execution (step-by-step procedures), and Boundary (limitations). This format ensures agents receive complete contextual grounding before executing actions, with explicit constraints preventing inappropriate application.

How does Cangjie-skill handle extraction units that fail verification?

Failed units move to the rejected/ subdirectory after Stage 1.5 triple-verification. This preserves auditability and enables potential reprocessing with refined parameters, rather than permanent deletion. The rejection reasons and original extraction metadata remain accessible for pipeline improvement analysis.

What makes the output compatible with darwin-skill?

Each skill's test-prompts.json follows darwin-skill's schema for automated pressure-testing and evolutionary refinement. The root-level consolidated test-prompts.json enables batch validation, while the structured markdown format allows darwin-skill parsers to ingest RIA-TV++ sections directly for agent runtime execution.

Where do I install the distillation output for Claude-style agents?

Copy the entire books/<book-slug>/ contents to your agent's skills directory—typically ~/.claude/skills/ as documented in the "安装使用" section of templates/INDEX.md.template. The INDEX.md file within the output provides specific integration instructions for your target agent platform.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →