How the RIA-TV++ 7-Stage Pipeline Works in cangjie-skill: A Complete Technical Breakdown
The RIA-TV++ pipeline converts raw text into agent-ready skills through seven sequential stages: Adler comprehension, parallel extraction, triple verification, RIA++ authoring, Zettelkasten linking, pressure testing, and final delivery.
The cangjie-skill repository implements RIA-TV++ – a rigorous methodology for distilling books, transcripts, and podcasts into modular, executable skills. Unlike simple summarization tools, this system enforces quality through multi-stage verification and produces structured outputs that downstream AI agents can invoke directly. The pipeline is orchestrated through SKILL.md and documented across seven methodology files in the methodology/ directory.
Stage 0: Adler Whole-Book Understanding
Source: [methodology/01-stage0-adler.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md)
Before any extraction begins, the system performs structural, interpretive, critical, and applicability analysis following Mortimer Adler's analytical reading framework. This stage ensures comprehensive understanding rather than surface-level skimming.
- Output:
books/<slug>/BOOK_OVERVIEW.md(generated fromtemplates/BOOK_OVERVIEW.md.template) - Key deliverable: The Boundary (B) field used in later skill construction
- Critical function: Prevents context fragmentation by anchoring all subsequent extractions to a holistic grasp of the source material
The BOOK_OVERVIEW.md serves as the single source of truth fed into all five parallel extractors in Stage 1.
Stage 1: Five Parallel Sub-Agents Extract Candidates
Source: [methodology/02-stage1-parallel-extract.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md)
Five independent extractors operate concurrently (with serial fallback), each targeting distinct knowledge types:
| Extractor | Target Knowledge | Output File |
|---|---|---|
framework-extractor |
Decision frameworks, mental models | candidates/frameworks.md |
principle-extractor |
Rules, checklists, operating principles | candidates/principles.md |
case-extractor |
Author's real-world examples | candidates/cases.md |
counter-example-extractor |
Failure modes, warnings, anti-patterns | candidates/counter-examples.md |
glossary-extractor |
Key terminology and concepts | candidates/glossary.md |
Each extractor receives three inputs: BOOK_OVERVIEW.md, the raw source text, and its specialized prompt from extractors/<type>-extractor.md. Every candidate output must include mandatory YAML frontmatter fields: id, title, type, source, quote, summary, and tags.
Stage 1.5: Triple Verification (TV)
Source: [methodology/03-stage1.5-triple-verify.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md)
This quality gate applies three independent validation criteria:
- V1 – Cross-Domain: The unit appears in at least two independent contexts within the source
- V2 – Predictive Power: The unit enables reasoning about novel problems beyond the text's explicit coverage
- V3 – Exclusivity: The insight is non-obvious to knowledgeable practitioners
Passed candidates write to books/<slug>/verified.md. Failures route to books/<slug>/rejected/<id>.md with full audit trails. Line 50 of the methodology file specifies a lightweight user confirmation step following automated verification.
Stage 2: RIA++ Construction (Skill Authoring)
Source: [methodology/04-stage2-ria-plus.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md)
Verified units expand into complete SKILL.md files through the RIA++ structure:
| Field | Content |
|---|---|
| R (Reading) | Original insight preserved verbatim |
| I (Interpretation) | Re-expression in the author's own words |
| A1/A2 (Appropriation) | Two distinct application scenarios |
| E (Execution) | Agent-runnable action steps |
| B (Boundary) | Scope limits and contextual requirements from Stage 0 |
Each SKILL.md is self-contained and ready for deployment to the skills/ directory.
Stage 3: Zettelkasten Linking (Knowledge Graph)
Source: [methodology/05-stage3-zettelkasten.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md)
Skills gain discoverability and composability through [[wiki-style links]] in INDEX.md. This graph structure enables:
- Cross-skill navigation for compound reasoning
- Integration with downstream agents like
darwin-skill - Emergent knowledge networks from accumulated skill libraries
Stage 4: Pressure Test (Robustness)
Source: [methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)
The pipeline generates test-prompts.json compatible with darwin-skill, then executes:
- Blind testing against held-out scenarios
- Failure analysis and iteration
- Refinement cycles until automated test suite passes
This stage eliminates brittle skills that fail under edge-case conditions.
Stage 5: Delivery (Digest)
Source: [methodology/07-stage5-deliver.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md)
Final outputs include:
- DIGEST.md: Human-readable long-form summary of the complete skill set
- Installed skills:
SKILL.mdfiles moved toskills/for direct agent invocation
Running the Pipeline: Code Example
The following Python script executes the complete RIA-TV++ pipeline in sequence, invoking the repository's agent-driven modules:
import subprocess
import pathlib
BOOK = pathlib.Path("books/my-book")
# Stage 0: Adler overview (manual or UI-generated)
# Stage 1: Parallel extraction across five extractors
subprocess.run(
["python", "-m", "cangjie_skill.extract", str(BOOK)],
check=True
)
# Stage 1.5: Triple verification with interactive confirmation
subprocess.run(
["python", "-m", "cangjie_skill.verify", str(BOOK)],
check=True
)
# Stage 2: Build RIA++ SKILL.md files from verified units
subprocess.run(
["python", "-m", "cangjie_skill.build_skills", str(BOOK)],
check=True
)
# Stage 3: Create Zettelkasten links and INDEX.md
subprocess.run(
["python", "-m", "cangjie_skill.link", str(BOOK)],
check=True
)
# Stage 4: Generate test-prompts.json and execute pressure tests
subprocess.run(
["python", "-m", "cangjie_skill.pressure_test", str(BOOK)],
check=True
)
# Stage 5: Produce DIGEST.md and install to skills/
subprocess.run(
["python", "-m", "cangjie_skill.deliver", str(BOOK)],
check=True
)
Note: The cangjie_skill.* modules represent the agent-tool interface described in the methodology documentation; actual execution follows the ordered seven-stage sequence regardless of implementation details.
Key Source Files Reference
Summary
- RIA-TV++ transforms unstructured text into agent-executable skills through seven documented stages
- Stage 0 establishes grounded comprehension via Adler's analytical reading methodology
- Stage 1 deploys five parallel extractors targeting distinct knowledge types with standardized YAML output
- Stage 1.5 enforces quality through cross-domain, predictive, and exclusivity verification
- Stage 2 structures validated content into the R-I-A-E-B framework of
SKILL.md - Stage 3 enables graph-based skill navigation through Zettelkasten linking
- Stage 4 validates robustness via
test-prompts.jsonand automated blind testing - Stage 5 delivers human-readable digests and installs skills for immediate agent use
Frequently Asked Questions
What does RIA-TV++ stand for?
RIA-TV++ abbreviates the pipeline's core components: Reading-Interpretation-Appropriation (the knowledge structuring framework), Triple Verification (the three-candidate validation gate), and the ++ suffix indicating the extended Execution and Boundary fields added to the classic RIA structure.
How does Triple Verification differ from simple fact-checking?
Triple Verification in methodology/03-stage1.5-triple-verify.md evaluates epistemic value, not just accuracy. V1 requires cross-contextual presence, V2 demands forward-reasoning utility, and V3 filters for non-obviousness. This eliminates tautologies and common knowledge that would waste agent context windows.
Can the pipeline handle non-book sources?
Yes. While the BOOK_OVERVIEW.md naming reflects literary origins, the extractors/ architecture and SKILL.md format are media-agnostic. Podcast transcripts, academic papers, video transcripts, and technical documentation all process through identical stages, with source-specific prompts substituted in Stage 1.
What makes a skill "agent-ready"?
Per methodology/04-stage2-ria-plus.md, agent-ready skills contain executable steps in field E (Execution) that require no external interpretation. The B (Boundary) field constrains applicability, preventing hallucinated invocation. Combined with Zettelkasten links for composability and pressure-test validation, these properties enable autonomous agent orchestration without human-in-the-loop intervention.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →