How to Create a New Skill from a Book Using cangjie-skill: Complete RIA-TV++ Pipeline Guide
cangjie-skill transforms books into reusable AI skills through the RIA-TV++ pipeline, a seven-stage automated workflow that extracts, verifies, and packages knowledge for AI agents.
The kangarooking/cangjie-skill repository provides a systematic methodology for converting long-form content into structured, testable skills that Claude, Cursor, and other AI agents can invoke. This guide walks you through the complete pipeline from raw book text to delivered skill.
The RIA-TV++ Pipeline Overview
According to the cangjie-skill source code, knowledge extraction follows seven well-defined stages documented in [methodology/00-overview.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md). Each stage produces specific artifacts and enforces quality gates.
Stage 0: Adler Analysis
The pipeline begins with structural deconstruction of the entire book. You create BOOK_OVERVIEW.md using the template at templates/BOOK_OVERVIEW.md.template, filling four Adler components:
- Structure — How the book is organized
- Explanation — Core arguments and frameworks
- Critique — Strengths and limitations
- Application — Where the knowledge applies
This stage establishes the foundation for all subsequent extraction.
Stage 1: Parallel Extraction
Five specialized LLM-driven extractors run simultaneously against the raw text:
| Extractor | Purpose | Output Location |
|---|---|---|
| principle-extractor | Core methods and mental models | candidates/principle/ |
| framework-extractor | Structured thinking systems | candidates/framework/ |
| case-extractor | Concrete examples and stories | candidates/case/ |
| counter-example-extractor | Failures and boundary conditions | candidates/counter/ |
| glossary-extractor | Key terminology and definitions | candidates/glossary/ |
Extractors are defined in the extractors/ directory. Each writes raw candidate method-logic units to the candidates/ folder.
Stage 1.5: Triple Verification
Every candidate must survive three independent checks before advancing:
- V1 — Cross-domain evidence: Does this principle work outside its original context?
- V2 — Predictive power: Can it forecast outcomes in new scenarios?
- V3 — Uniqueness: Is it distinguishable from existing skills?
Passed candidates move to skill/; failed ones are archived in rejected/. The verification logic is detailed in [methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md).
Stage 2: RIA++ Construction
Accepted candidates become full skill definitions following the R-I-A1-A2-E-B schema (Rule, Input, Action1, Action2, Example, Boundary). You render templates/SKILL.md.template into skill/<skill-slug>/SKILL.md.
Stage 3: Zettelkasten Linking
Skills are connected into a knowledge graph through INDEX.md (templates/INDEX.md.template). Dependencies, contrasts, and compositions between skills are explicitly mapped.
Stage 4: Pressure Testing
Each skill receives a test-prompts.json file containing validated test prompts. The repository's test runner evaluates pass-rate; skills below threshold return to Stage 2 for refinement.
Stage 5: Delivery
A concise DIGEST.md is generated from templates/DIGEST.md.template, and the final skill package is copied to the target agent's skill directory.
Step-by-Step Workflow to Create a New Skill from a Book
1. Prepare Your Source Material
Export the complete book as plain-text UTF-8. OCR quality, ebook conversions, or transcript exports all work—consistency matters more than source format.
2. Run Adler Analysis
Populate the template fields in BOOK_OVERVIEW.md:
# BOOK_OVERVIEW.md for "The Lean Startup"
## Structure
Part I: Vision → Part II: Steer → Part III: Accelerate
## Explanation
Build-Measure-Learn feedback loop with validated learning as the metric
## Critique
Weak on regulatory environments; assumes rapid iteration is always possible
## Application
Startups, new product development, innovation teams
3. Execute Parallel Extractors
Run the five extractor prompts located in extractors/:
# Typical invocation pattern (exact CLI varies by setup)
for extractor in extractors/*.md; do
llm process --prompt "$extractor" --input book.txt --output "candidates/$(basename $extractor .md).jsonl"
done
4. Apply Triple Verification
For each candidate in candidates/, evaluate against V1/V2/V3 criteria. Move passing items:
mkdir -p skill rejected
mv candidates/principle/mvp-passed-v123.jsonl skill/lean-startup-mvp/candidate.jsonl
mv candidates/principle/mvp-failed-v2.jsonl rejected/
5. Generate the Skill Definition
Render the skill template with verified candidate data:
import pathlib, jinja2, json, datetime
# Load the SKILL template
template_path = pathlib.Path('templates/SKILL.md.template')
template = jinja2.Environment(
loader=jinja2.FileSystemLoader(template_path.parent)
).get_template(template_path.name)
# Payload from verification step
payload = {
"skill-slug": "lean-startup-mvp",
"BOOK_TITLE": "The Lean Startup",
"AUTHOR": "Eric Ries",
"章节": "Chapter 3 – Validated Learning",
"tag1": "lean", "tag2": "startup",
"Skill Title": "Minimum Viable Product (MVP)",
"原文引用": "The only way to win is to learn faster than your competitors.",
"作者": "Eric Ries", "CHAPTER": "3.2",
"用自己的话重写": "An MVP is the smallest product that can be released to test a hypothesis about the market.",
"案例名": "Dropbox early prototype",
"作者遇到了什么": "Need to validate demand for file-sync service",
"作者怎么用这个方法论思考": "Built a simple video demo before writing any code",
"得出了什么": "Users were eager to sign-up",
"实际发生了什么": "Dropbox launched the service six months later",
"场景 1": "Founder wants to test a new SaaS idea with minimal resources",
"场景 2": "Product team needs a quick way to validate UI concepts",
"场景 3": "Investors ask for evidence of market interest",
"典型措辞 1": "How can I test my idea cheaply?",
"典型措辞 2": "What's the smallest version I can launch?",
"related-skill-a": "lean-startup-customer-development",
"related-skill-b": "lean-startup-pivot",
"步骤 1": "Define a clear hypothesis",
"如何判断这一步已完成": "Hypothesis written in one sentence",
"步骤 2": "Identify the minimum set of features",
"步骤 3": "Release to a test audience",
"反场景 1": "When the product concept is already fully defined",
"反场景 2": "When the market is regulated and cannot be tested quickly",
"%": "92",
"DATE": datetime.date.today().isoformat()
}
# Render and write
skill_md = template.render(payload)
skill_dir = pathlib.Path('skill') / payload["skill-slug"]
skill_dir.mkdir(parents=True, exist_ok=True)
(skill_dir / 'SKILL.md').write_text(skill_md, encoding='utf-8')
print(f'✅ Skill written to {skill_dir / "SKILL.md"}')
6. Update the Skill Index
Add your new skill to INDEX.md:
#!/usr/bin/env bash
SKILL_SLUG="lean-startup-mvp"
SKILL_TITLE="Minimum Viable Product (MVP)"
INDEX_FILE="INDEX.md"
printf "\n- [%s](skill/%s/SKILL.md)\n" "$SKILL_TITLE" "$SKILL_SLUG" >> "$INDEX_FILE"
echo "✅ $SKILL_TITLE added to $INDEX_FILE"
Then execute the Zettelkasten linking step per [methodology/05-stage3-zettelkasten.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md).
7. Write and Run Pressure Tests
Create skill/<skill-slug>/test-prompts.json following the schema in templates/DIGEST.md.template, then validate:
./run_test.sh --skill-dir skill/lean-startup-mvp
# Expected output: Pass rate: 94% (47/50) — Above threshold, proceed to delivery
If pass-rate falls below threshold, refine the skill and repeat Stage 2.
8. Deliver to Target Agent
Generate the digest and install:
# Generate DIGEST.md from template
python render_digest.py --skill-dir skill/lean-startup-mvp
# Copy to Claude Code skills directory
cp -r skill/lean-startup-mvp ~/.claude-code/skills/
# Or Cursor
cp -r skill/lean-startup-mvp ~/.cursor/skills/
Key Template Files in cangjie-skill
| File | Purpose |
|---|---|
templates/SKILL.md.template |
Master template for R-I-A1-A2-E-B skill structure |
templates/BOOK_OVERVIEW.md.template |
Adler analysis output format |
templates/INDEX.md.template |
Skill registry and relationship graph |
templates/DIGEST.md.template |
Human-readable summary and test schema |
SKILL.md (root) |
Global pipeline specification |
Automation vs. Human Judgment
The cangjie-skill pipeline is fully automated for stages 1-5—all extractors and verification checks run through LLM sub-agents. However, you must provide:
- The raw book text as input
- Confirmation of which candidates advance past triple verification
- Approval of final skill content before delivery
This human-in-the-loop design ensures quality while scaling throughput.
Summary
- cangjie-skill converts books to AI skills via the RIA-TV++ pipeline with seven stages (Adler analysis → parallel extraction → triple verification → RIA++ construction → Zettelkasten linking → pressure testing → delivery)
- Five parallel extractors (principle, framework, case, counter-example, glossary) generate candidate method-logic units from raw text
- Triple verification (V1/V2/V3) filters candidates for cross-domain validity, predictive power, and uniqueness
- All templates live in
templates/and follow strict schemas: R-I-A1-A2-E-B for skills, Adler format for book overviews - Pressure testing via
test-prompts.jsonensures skills meet quality thresholds before delivery to agent directories
Frequently Asked Questions
What file format should the source book be in?
Plain-text UTF-8 is required. cangjie-skill processes raw character sequences without formatting, so PDFs, EPUBs, or other formats must be converted first. The repository does not provide conversion tools—you must export or extract clean text before Stage 0.
How does triple verification prevent low-quality skills from advancing?
Each candidate must independently satisfy three checks: V1 requires evidence the principle works outside its original domain; V2 demands predictive questions that the skill answers correctly; V3 confirms uniqueness against existing skills. A candidate failing any check is routed to rejected/ rather than skill/, as implemented in the verification logic of methodology/06-stage4-pressure-test.md.
Can I customize the extractor prompts for specific book genres?
Yes. The extractor definitions in extractors/*.md are standard markdown prompt files. You can modify principle-extractor.md, framework-extractor.md, and others to emphasize domain-specific patterns—for example, adding legal case structures for law books or experimental methods for scientific texts. The pipeline architecture remains unchanged.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →