How the RIA-TV++ 7-Stage Pipeline Works in cangjie-skill: A Complete Technical Breakdown

The RIA-TV++ pipeline converts raw text into agent-ready skills through seven sequential stages: Adler comprehension, parallel extraction, triple verification, RIA++ authoring, Zettelkasten linking, pressure testing, and final delivery.

The cangjie-skill repository implements RIA-TV++ – a rigorous methodology for distilling books, transcripts, and podcasts into modular, executable skills. Unlike simple summarization tools, this system enforces quality through multi-stage verification and produces structured outputs that downstream AI agents can invoke directly. The pipeline is orchestrated through SKILL.md and documented across seven methodology files in the methodology/ directory.

Stage 0: Adler Whole-Book Understanding

Source: [methodology/01-stage0-adler.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md)

Before any extraction begins, the system performs structural, interpretive, critical, and applicability analysis following Mortimer Adler's analytical reading framework. This stage ensures comprehensive understanding rather than surface-level skimming.

  • Output: books/<slug>/BOOK_OVERVIEW.md (generated from templates/BOOK_OVERVIEW.md.template)
  • Key deliverable: The Boundary (B) field used in later skill construction
  • Critical function: Prevents context fragmentation by anchoring all subsequent extractions to a holistic grasp of the source material

The BOOK_OVERVIEW.md serves as the single source of truth fed into all five parallel extractors in Stage 1.

Stage 1: Five Parallel Sub-Agents Extract Candidates

Source: [methodology/02-stage1-parallel-extract.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md)

Five independent extractors operate concurrently (with serial fallback), each targeting distinct knowledge types:

Extractor Target Knowledge Output File
framework-extractor Decision frameworks, mental models candidates/frameworks.md
principle-extractor Rules, checklists, operating principles candidates/principles.md
case-extractor Author's real-world examples candidates/cases.md
counter-example-extractor Failure modes, warnings, anti-patterns candidates/counter-examples.md
glossary-extractor Key terminology and concepts candidates/glossary.md

Each extractor receives three inputs: BOOK_OVERVIEW.md, the raw source text, and its specialized prompt from extractors/<type>-extractor.md. Every candidate output must include mandatory YAML frontmatter fields: id, title, type, source, quote, summary, and tags.

Stage 1.5: Triple Verification (TV)

Source: [methodology/03-stage1.5-triple-verify.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md)

This quality gate applies three independent validation criteria:

  • V1 – Cross-Domain: The unit appears in at least two independent contexts within the source
  • V2 – Predictive Power: The unit enables reasoning about novel problems beyond the text's explicit coverage
  • V3 – Exclusivity: The insight is non-obvious to knowledgeable practitioners

Passed candidates write to books/<slug>/verified.md. Failures route to books/<slug>/rejected/<id>.md with full audit trails. Line 50 of the methodology file specifies a lightweight user confirmation step following automated verification.

Stage 2: RIA++ Construction (Skill Authoring)

Source: [methodology/04-stage2-ria-plus.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md)

Verified units expand into complete SKILL.md files through the RIA++ structure:

Field Content
R (Reading) Original insight preserved verbatim
I (Interpretation) Re-expression in the author's own words
A1/A2 (Appropriation) Two distinct application scenarios
E (Execution) Agent-runnable action steps
B (Boundary) Scope limits and contextual requirements from Stage 0

Each SKILL.md is self-contained and ready for deployment to the skills/ directory.

Stage 3: Zettelkasten Linking (Knowledge Graph)

Source: [methodology/05-stage3-zettelkasten.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md)

Skills gain discoverability and composability through [[wiki-style links]] in INDEX.md. This graph structure enables:

  • Cross-skill navigation for compound reasoning
  • Integration with downstream agents like darwin-skill
  • Emergent knowledge networks from accumulated skill libraries

Stage 4: Pressure Test (Robustness)

Source: [methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)

The pipeline generates test-prompts.json compatible with darwin-skill, then executes:

  1. Blind testing against held-out scenarios
  2. Failure analysis and iteration
  3. Refinement cycles until automated test suite passes

This stage eliminates brittle skills that fail under edge-case conditions.

Stage 5: Delivery (Digest)

Source: [methodology/07-stage5-deliver.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md)

Final outputs include:

  • DIGEST.md: Human-readable long-form summary of the complete skill set
  • Installed skills: SKILL.md files moved to skills/ for direct agent invocation

Running the Pipeline: Code Example

The following Python script executes the complete RIA-TV++ pipeline in sequence, invoking the repository's agent-driven modules:

import subprocess
import pathlib

BOOK = pathlib.Path("books/my-book")

# Stage 0: Adler overview (manual or UI-generated)

# Stage 1: Parallel extraction across five extractors

subprocess.run(
    ["python", "-m", "cangjie_skill.extract", str(BOOK)], 
    check=True
)

# Stage 1.5: Triple verification with interactive confirmation

subprocess.run(
    ["python", "-m", "cangjie_skill.verify", str(BOOK)], 
    check=True
)

# Stage 2: Build RIA++ SKILL.md files from verified units

subprocess.run(
    ["python", "-m", "cangjie_skill.build_skills", str(BOOK)], 
    check=True
)

# Stage 3: Create Zettelkasten links and INDEX.md

subprocess.run(
    ["python", "-m", "cangjie_skill.link", str(BOOK)], 
    check=True
)

# Stage 4: Generate test-prompts.json and execute pressure tests

subprocess.run(
    ["python", "-m", "cangjie_skill.pressure_test", str(BOOK)], 
    check=True
)

# Stage 5: Produce DIGEST.md and install to skills/

subprocess.run(
    ["python", "-m", "cangjie_skill.deliver", str(BOOK)], 
    check=True
)

Note: The cangjie_skill.* modules represent the agent-tool interface described in the methodology documentation; actual execution follows the ordered seven-stage sequence regardless of implementation details.

Key Source Files Reference

Purpose Path
Pipeline overview [methodology/00-overview.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md)
Adler analysis spec [methodology/01-stage0-adler.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md)
Parallel extraction design [methodology/02-stage1-parallel-extract.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md)
Triple verification criteria [methodology/03-stage1.5-triple-verify.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md)
RIA++ skill construction [methodology/04-stage2-ria-plus.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md)
Zettelkasten linking [methodology/05-stage3-zettelkasten.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md)
Pressure testing protocol [methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)
Delivery and digest [methodology/07-stage5-deliver.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md)
User-facing documentation [SKILL.md](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)
English quick-start [README.en.md](https://github.com/kangarooking/cangjie-skill/blob/main/README.en.md)

Summary

  • RIA-TV++ transforms unstructured text into agent-executable skills through seven documented stages
  • Stage 0 establishes grounded comprehension via Adler's analytical reading methodology
  • Stage 1 deploys five parallel extractors targeting distinct knowledge types with standardized YAML output
  • Stage 1.5 enforces quality through cross-domain, predictive, and exclusivity verification
  • Stage 2 structures validated content into the R-I-A-E-B framework of SKILL.md
  • Stage 3 enables graph-based skill navigation through Zettelkasten linking
  • Stage 4 validates robustness via test-prompts.json and automated blind testing
  • Stage 5 delivers human-readable digests and installs skills for immediate agent use

Frequently Asked Questions

What does RIA-TV++ stand for?

RIA-TV++ abbreviates the pipeline's core components: Reading-Interpretation-Appropriation (the knowledge structuring framework), Triple Verification (the three-candidate validation gate), and the ++ suffix indicating the extended Execution and Boundary fields added to the classic RIA structure.

How does Triple Verification differ from simple fact-checking?

Triple Verification in methodology/03-stage1.5-triple-verify.md evaluates epistemic value, not just accuracy. V1 requires cross-contextual presence, V2 demands forward-reasoning utility, and V3 filters for non-obviousness. This eliminates tautologies and common knowledge that would waste agent context windows.

Can the pipeline handle non-book sources?

Yes. While the BOOK_OVERVIEW.md naming reflects literary origins, the extractors/ architecture and SKILL.md format are media-agnostic. Podcast transcripts, academic papers, video transcripts, and technical documentation all process through identical stages, with source-specific prompts substituted in Stage 1.

What makes a skill "agent-ready"?

Per methodology/04-stage2-ria-plus.md, agent-ready skills contain executable steps in field E (Execution) that require no external interpretation. The B (Boundary) field constrains applicability, preventing hallucinated invocation. Combined with Zettelkasten links for composability and pressure-test validation, these properties enable autonomous agent orchestration without human-in-the-loop intervention.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →