How cangjie-skill Handles Non-Book Content: A Guide to Processing Video and Podcast Transcripts
cangjie-skill treats any long-form text as a "book," deliberately including long-video transcripts, podcast text, courses, and interviews within its unified five-stage pipeline.
The cangjie-skill repository from kangarooking/cangjie-skill is designed to transform textual content into reusable, structured skills. While the terminology uses "book" throughout the codebase, the project explicitly extends this definition to encompass non-book sources. According to SKILL.md lines 12-13, the term book deliberately includes long-video transcripts, podcast text, courses, interviews, and other extensive documents.
How cangjie-skill Processes Video and Podcast Transcripts
The pipeline for non-book content mirrors the standard book workflow exactly. The only adaptations occur in input validation and metadata mapping—no separate code path exists for transcripts.
Step 1: Input Validation and Metadata Mapping
When you provide a video or podcast, cangjie-skill requests the text source (PDF, EPUB, TXT, subtitles, or transcript) plus metadata (title, author, publish date). In SKILL.md lines 44-48, the documentation specifies that for video/podcast sources, metadata replaces the traditional "book name + author + year" pattern, and the chapter field becomes a timestamp or episode number to maintain traceability.
# Example: Distilling a podcast transcript into reusable skills
transcript_path = "/path/to/podcast_transcript.txt"
# 1️⃣ Provide metadata (title, host, publish date)
metadata = {
"title": "Deep Learning Podcast – Episode 12",
"author": "Jane Doe",
"date": "2023-06-15"
}
# 2️⃣ Invoke the cangjie-skill pipeline (pseudo-CLI call)
# The tool requests the transcript file and metadata, then runs all stages
# $ cangjie --input transcript_path --meta "{\"title\":\"...\",\"author\":\"...\",\"date\":\"...\"}"
# Shell example for a video transcript
cangjie \
--input /videos/learning_python_transcript.srt \
--meta '{"title":"Learning Python – Full Course", "author":"John Smith", "date":"2022-11-02"}'
Step 2: Stage 0 — Adler整书理解 (Whole-Book Comprehension)
The raw transcript undergoes chunking and Adler four-step analysis (structure, explanation, critique, application) exactly as with traditional books. This stage is implemented in SKILL.md lines 74-78.
Step 3: Stage 1 — Parallel Extraction with Chunking
Five sub-agents execute the same extractors on the transcript:
- Framework extractor
- Principle extractor
- Case extractor
- Counter-example extractor
- Glossary extractor
For very long transcripts, the pipeline applies the chunking strategy defined in methodology/02-stage1-parallel-extract.md. This ensures parallel processing without losing contextual coherence across transcript segments.
Step 4: Stage 1.5 — Triple Verification
Extracted units pass through identical verification criteria as books: cross-domain evidence, predictive power, and uniqueness checks (SKILL.md lines 98-105).
Step 5: Stages 2-5 — Skill Construction and Output
Validated units become atomic skills through:
- RIA++ construction (Stage 2)
- Zettelkasten linking (Stage 3)
- Darwin-compatible pressure testing (Stage 4)
- Final digest generation (Stage 5)
The system uses standard templates: templates/SKILL.md.template, templates/INDEX.md.template, and others (SKILL.md lines 22-30). Timestamps or episode identifiers persist in the source_chapter field to preserve source traceability (SKILL.md line 48).
Output structure in books/<slug>/:
BOOK_OVERVIEW.md- Individual
SKILL.mdfiles INDEX.mdGLOSSARY.md- Final
DIGEST.md
Key Design Principles for Non-Book Content
| Principle | Implementation |
|---|---|
| Unified pipeline | No branching logic for transcript types—all content flows through identical stages |
| Metadata flexibility | Title/author/date mapping adapts to content source without schema changes |
| Chunking for scale | Long transcripts split intelligently per methodology/02-stage1-parallel-extract.md |
| Traceability preservation | Timestamps and episode numbers retained in source_chapter field |
Critical Source Files
Understanding these files clarifies how cangjie-skill generalizes across content types:
SKILL.md— Pipeline definition and "non-book" terminology expansionmethodology/00-overview.md— Five-stage RIA-TV++ workflow documentationmethodology/02-stage1-parallel-extract.md— Long-document chunking strategiestemplates/BOOK_OVERVIEW.md.template— Rendering template for all source typesextractors/*-extractor.md— Five extractor agent definitions
Summary
- cangjie-skill does not distinguish between books and transcripts—it processes any long-form text through identical stages.
- Metadata fields adapt to capture video/podcast specifics (timestamps replace chapters).
- Chunking in Stage 1 handles transcript length without algorithmic changes.
- Template-based output remains consistent regardless of input source.
- Source traceability is preserved through dedicated timestamp/episode fields.
Frequently Asked Questions
Does cangjie-skill require different configuration for podcasts versus books?
No. Both use the same CLI invocation and pipeline configuration. Only the metadata values differ—podcasts use episode numbers where books use chapter numbers, and timestamps replace page references in the source_chapter field.
What transcript formats does cangjie-skill accept?
The input validation documented in SKILL.md lines 44-48 supports PDF, EPUB, TXT, subtitles (字幕), and raw transcripts (转写稿). No preprocessing is required beyond ensuring the text is extractable from these containers.
How does cangjie-skill handle extremely long video transcripts?
Stage 1 implements automatic chunking per methodology/02-stage1-parallel-extract.md. The parallel extraction agents process segments independently, then results merge for verification—scaling to multi-hour content without manual intervention.
Can extracted skills from podcasts link to skills derived from books?
Yes. The Zettelkasten linking stage (Stage 3) treats all sources uniformly. Skills from video transcripts receive identical UUID-based identifiers and can form bidirectional links with skills from any other source type in the knowledge graph.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →