How cangjie-skill Handles Non-Book Content: A Guide to Processing Video and Podcast Transcripts

cangjie-skill treats any long-form text as a "book," deliberately including long-video transcripts, podcast text, courses, and interviews within its unified five-stage pipeline.

The cangjie-skill repository from kangarooking/cangjie-skill is designed to transform textual content into reusable, structured skills. While the terminology uses "book" throughout the codebase, the project explicitly extends this definition to encompass non-book sources. According to SKILL.md lines 12-13, the term book deliberately includes long-video transcripts, podcast text, courses, interviews, and other extensive documents.

How cangjie-skill Processes Video and Podcast Transcripts

The pipeline for non-book content mirrors the standard book workflow exactly. The only adaptations occur in input validation and metadata mapping—no separate code path exists for transcripts.

Step 1: Input Validation and Metadata Mapping

When you provide a video or podcast, cangjie-skill requests the text source (PDF, EPUB, TXT, subtitles, or transcript) plus metadata (title, author, publish date). In SKILL.md lines 44-48, the documentation specifies that for video/podcast sources, metadata replaces the traditional "book name + author + year" pattern, and the chapter field becomes a timestamp or episode number to maintain traceability.


# Example: Distilling a podcast transcript into reusable skills

transcript_path = "/path/to/podcast_transcript.txt"

# 1️⃣ Provide metadata (title, host, publish date)

metadata = {
    "title": "Deep Learning Podcast – Episode 12",
    "author": "Jane Doe",
    "date": "2023-06-15"
}

# 2️⃣ Invoke the cangjie-skill pipeline (pseudo-CLI call)

# The tool requests the transcript file and metadata, then runs all stages

# $ cangjie --input transcript_path --meta "{\"title\":\"...\",\"author\":\"...\",\"date\":\"...\"}"

# Shell example for a video transcript

cangjie \
  --input /videos/learning_python_transcript.srt \
  --meta '{"title":"Learning Python – Full Course", "author":"John Smith", "date":"2022-11-02"}'

Step 2: Stage 0 — Adler整书理解 (Whole-Book Comprehension)

The raw transcript undergoes chunking and Adler four-step analysis (structure, explanation, critique, application) exactly as with traditional books. This stage is implemented in SKILL.md lines 74-78.

Step 3: Stage 1 — Parallel Extraction with Chunking

Five sub-agents execute the same extractors on the transcript:

  • Framework extractor
  • Principle extractor
  • Case extractor
  • Counter-example extractor
  • Glossary extractor

For very long transcripts, the pipeline applies the chunking strategy defined in methodology/02-stage1-parallel-extract.md. This ensures parallel processing without losing contextual coherence across transcript segments.

Step 4: Stage 1.5 — Triple Verification

Extracted units pass through identical verification criteria as books: cross-domain evidence, predictive power, and uniqueness checks (SKILL.md lines 98-105).

Step 5: Stages 2-5 — Skill Construction and Output

Validated units become atomic skills through:

  • RIA++ construction (Stage 2)
  • Zettelkasten linking (Stage 3)
  • Darwin-compatible pressure testing (Stage 4)
  • Final digest generation (Stage 5)

The system uses standard templates: templates/SKILL.md.template, templates/INDEX.md.template, and others (SKILL.md lines 22-30). Timestamps or episode identifiers persist in the source_chapter field to preserve source traceability (SKILL.md line 48).

Output structure in books/<slug>/:

Key Design Principles for Non-Book Content

Principle Implementation
Unified pipeline No branching logic for transcript types—all content flows through identical stages
Metadata flexibility Title/author/date mapping adapts to content source without schema changes
Chunking for scale Long transcripts split intelligently per methodology/02-stage1-parallel-extract.md
Traceability preservation Timestamps and episode numbers retained in source_chapter field

Critical Source Files

Understanding these files clarifies how cangjie-skill generalizes across content types:

Summary

  • cangjie-skill does not distinguish between books and transcripts—it processes any long-form text through identical stages.
  • Metadata fields adapt to capture video/podcast specifics (timestamps replace chapters).
  • Chunking in Stage 1 handles transcript length without algorithmic changes.
  • Template-based output remains consistent regardless of input source.
  • Source traceability is preserved through dedicated timestamp/episode fields.

Frequently Asked Questions

Does cangjie-skill require different configuration for podcasts versus books?

No. Both use the same CLI invocation and pipeline configuration. Only the metadata values differ—podcasts use episode numbers where books use chapter numbers, and timestamps replace page references in the source_chapter field.

What transcript formats does cangjie-skill accept?

The input validation documented in SKILL.md lines 44-48 supports PDF, EPUB, TXT, subtitles (字幕), and raw transcripts (转写稿). No preprocessing is required beyond ensuring the text is extractable from these containers.

How does cangjie-skill handle extremely long video transcripts?

Stage 1 implements automatic chunking per methodology/02-stage1-parallel-extract.md. The parallel extraction agents process segments independently, then results merge for verification—scaling to multi-hour content without manual intervention.

Yes. The Zettelkasten linking stage (Stage 3) treats all sources uniformly. Skills from video transcripts receive identical UUID-based identifiers and can form bidirectional links with skills from any other source type in the knowledge graph.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →