How to Apply RIA拆书法 Principles Beyond Books: A Complete Guide to Media-Agnostic Skill Extraction
You can apply RIA拆书法 (Reading-Interpretation-Appropriation) to videos, podcasts, and courses by treating "chapters" as timestamps and running the same seven-stage Cangjie-Skill pipeline on plain-text transcripts.
The Cangjie-Skill repository (kangarooking/cangjie-skill) implements the RIA-TV++ pipeline, a deliberately media-agnostic system that redefines "书" (book) as any long-form content. According to the front-matter in [SKILL.md](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md), this includes video transcripts, podcast scripts, course notes, and interview logs.
Understanding the Media-Agnostic Foundation
The pipeline's versatility stems from its definition of source material. In [methodology/00-overview.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md), the system treats all input as plain text, allowing 章节/页码 (chapters/page numbers) to map to 时间戳 (timestamps) for video or 集数 (episode numbers) for podcasts.
This abstraction means the extraction logic never needs to know whether it's processing a PDF or a YouTube transcript. The markdown headings remain identical, guaranteeing downstream extractors locate source locations uniformly.
Stage 0: Whole-Content Analysis for Any Medium
Before extraction, the system performs Adler-style analysis defined in [methodology/01-stage0-adler.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md). This stage creates a BOOK_OVERVIEW.md that structures any content type.
For non-book sources:
- Input: PDF/EPUB converted to plain text, or subtitle/transcription files (
.srt,.vtt,.txt) - Structure mapping: Video sections become timestamp ranges (e.g.,
00:00:00–00:05:30) - Criticism & Application: The same four-step Adler analysis (structure, explanation, criticism, application) applies regardless of medium
The output is a unified markdown document where temporal locations replace pagination, ensuring the parallel extractors receive standardized input.
Stage 1: Parallel Extraction from Transcripts
Five specialized extractors run simultaneously (or sequentially in resource-constrained environments) to identify candidate skill units. Each extractor reads the plain text from Stage 0 and emits candidates to specific folders:
| Extractor | Prompt File | Output Location |
|---|---|---|
| Framework | [extractors/framework-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md) |
candidates/frameworks.md |
| Principle | extractors/principle-extractor.md |
candidates/principles.md |
| Case | extractors/case-extractor.md |
candidates/cases.md |
| Counter-example | extractors/counter-example-extractor.md |
candidates/counter-examples.md |
| Glossary | extractors/glossary-extractor.md |
candidates/glossary.md |
These files contain media-agnostic cue rules such as "author-named thinking patterns" and "when facing X class of problems..." Because they operate on pure text, a video transcript undergoes identical processing to a book chapter.
Stage 1.5: Triple Verification Across Contexts
Candidate units must pass three independent verification gates defined in [methodology/03-stage1.5-triple-verify.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md):
- Cross-Domain (V1): The unit appears in at least two distinct contexts within the source. For video, this means different timestamps or speaker turns; for podcasts, different episodes.
- Predictive Power (V2): The unit must answer a novel question not explicitly stated in the source material.
- Exclusivity (V3): The unit reflects the author's unique insight, not common knowledge.
The verification script extracts temporal context from surrounding text, ensuring a method mentioned at 00:15:00 and demonstrated at 00:45:00 qualifies as cross-domain.
Stage 2: Building RIA++ Skills with Temporal Anchors
Verified units transform into Claude-Code skills using the RIA++ six-section template documented in [methodology/04-stage2-ria-plus.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md). The template adapts to non-book sources as follows:
| Section | Book Format | Video/Podcast Adaptation |
|---|---|---|
| R – Reading | ≤150-word excerpt + page/chapter | ≤150-word excerpt + timestamp (e.g., 00:12:34–00:13:10) |
| I – Interpretation | Rewrite in own words | References visual cues ("when presenter points to slide 3") |
| A1 – Past Application | Author's concrete case | Video segment where author demonstrates the method |
| A2 – Future Trigger | Trigger phrases for queries | Includes spoken cues listeners might use ("I'm stuck where the speaker says '...'") |
| E – Execution | 1-2-3 actionable steps | References media actions ("pause at 02:45, write down...") |
| B – Boundary | When NOT to use | Lists situations the video doesn't cover |
The final SKILL.md generates from templates/SKILL.md.template, with timestamp fields auto-populated from transcript line numbers.
Stages 3-5: Linking, Testing, and Delivery
Stage 3 – Zettelkasten Linking ([methodology/05-stage3-zettelkasten.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md)): Skill dependencies are discovered automatically using source timestamps stored in front-matter. A graph view can show that Skill A (timestamp 00:05:00) triggers Skill B (timestamp 12:30:00) when both address the same problem domain.
Stage 4 – Pressure Testing ([methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)): Each skill receives test prompts from templates/test-prompts.json.template:
- Positive trigger: Uses exact spoken phrase from transcript
- Negative (bait) trigger: Uses similar phrase from different video segment
- Boundary test: Asks cross-domain questions to verify the B-section constraints
Stage 5 – Delivery ([methodology/07-stage5-deliver.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md)): Produces the standard directory structure:
books/<slug>/
├── BOOK_OVERVIEW.md # Whole-content summary
├── verified.md # Passed candidates
├── INDEX.md # Skill map
├── GLOSSARY.md # Shared terminology
├── DIGEST.md # Highlights without full content
└── <skill-slug>/SKILL.md
└── test-prompts.json
The DIGEST.md template allows users to consume key insights without watching the entire video or listening to the full podcast.
Practical Implementation for Video Content
Below is a conceptual YAML configuration showing how to invoke the pipeline on a YouTube transcript. The actual execution is handled by the skill orchestrator inside Claude Code:
# 1️⃣ Provide the raw transcript (plain-text)
input:
source_type: "video"
source_path: "data/transcripts/my-talk.txt"
metadata:
title: "Design Thinking in Product Teams"
creator: "TechTalks"
published: "2024-03-15"
# 2️⃣ Run the pipeline (calls the cangjie-skill meta-skill)
run: cangjie-skill
# 3️⃣ After completion, the skill directory will be:
output:
skill_dir: "books/design-thinking-in-product-teams/"
The orchestrator maps these steps directly to the methodology files, automatically populating timestamp fields from transcript metadata.
Summary
- RIA拆书法 adapts to any long-form content by converting media to plain text and mapping locations to timestamps or episode numbers.
- The seven-stage pipeline (0-5) runs identically across books, videos, and podcasts because the extractors operate on standardized markdown.
- Triple verification (V1-V3) treats different timestamps as distinct contexts, ensuring video-derived skills meet the same quality standards as book-derived ones.
- The RIA++ template uses six sections (R, I, A1, A2, E, B) with temporal anchors instead of page numbers.
- All core logic resides in the
methodology/andextractors/directories of thekangarooking/cangjie-skillrepository.
Frequently Asked Questions
Can RIA拆书法 work with audio-only podcasts?
Yes. The pipeline processes podcast transcripts exactly like book text. Episode numbers replace chapter numbers, and timestamp ranges replace page numbers. The five extractors ([extractors/framework-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md) and others) identify methods based on linguistic cues ("the host suggests..." or "in this episode...") without requiring visual context.
How do I handle video content without transcripts?
You must generate a plain-text transcript first. The pipeline expects standardized input into Stage 0 ([methodology/01-stage0-adler.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md)). Use automated transcription tools to produce .txt or .srt files, then feed them into the Adler analysis step. The system cannot process raw video files directly.
What's the difference between A1 and A2 in video contexts?
In the RIA++ template ([methodology/04-stage2-ria-plus.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md)), A1 (Past Application) cites the specific video segment where the author demonstrated the method (e.g., "see timestamp 15:30 where the presenter debugs the code"). A2 (Future Trigger) provides conversational phrases a user might speak when seeking that skill, including exact quotes from the video that would activate the skill in Claude Code.
Does the pipeline work for live workshop recordings?
Yes, provided the recording is transcribed. The pressure testing stage ([methodology/06-stage4-pressure-test.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md)) includes specific safeguards for live content: boundary tests verify that skills extracted from Q&A segments don't conflate audience questions with facilitator expertise, and the Exclusivity (V3) filter removes generic advice that often appears in unscripted workshop banter.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →