# How cangjie-skill Handles Non-Book Content: A Guide to Processing Video and Podcast Transcripts

> Discover how cangjie-skill processes non-book content like video and podcast transcripts through its unified five-stage pipeline. Learn to handle diverse long-form text effectively.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-08-15

---

**cangjie-skill treats any long-form text as a "book," deliberately including long-video transcripts, podcast text, courses, and interviews within its unified five-stage pipeline.**

The cangjie-skill repository from kangarooking/cangjie-skill is designed to transform textual content into reusable, structured skills. While the terminology uses **"book"** throughout the codebase, the project explicitly extends this definition to encompass non-book sources. According to [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) lines 12-13, the term **book** deliberately includes *long-video transcripts, podcast text, courses, interviews, and other extensive documents*.

## How cangjie-skill Processes Video and Podcast Transcripts

The pipeline for non-book content mirrors the standard book workflow exactly. The only adaptations occur in input validation and metadata mapping—no separate code path exists for transcripts.

### Step 1: Input Validation and Metadata Mapping

When you provide a video or podcast, cangjie-skill requests the **text source** (PDF, EPUB, TXT, subtitles, or transcript) plus **metadata** (title, author, publish date). In [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) lines 44-48, the documentation specifies that for video/podcast sources, metadata replaces the traditional "book name + author + year" pattern, and the *chapter* field becomes a **timestamp or episode number** to maintain traceability.

```python

# Example: Distilling a podcast transcript into reusable skills

transcript_path = "/path/to/podcast_transcript.txt"

# 1️⃣ Provide metadata (title, host, publish date)

metadata = {
    "title": "Deep Learning Podcast – Episode 12",
    "author": "Jane Doe",
    "date": "2023-06-15"
}

# 2️⃣ Invoke the cangjie-skill pipeline (pseudo-CLI call)

# The tool requests the transcript file and metadata, then runs all stages

# $ cangjie --input transcript_path --meta "{\"title\":\"...\",\"author\":\"...\",\"date\":\"...\"}"

```

```bash

# Shell example for a video transcript

cangjie \
  --input /videos/learning_python_transcript.srt \
  --meta '{"title":"Learning Python – Full Course", "author":"John Smith", "date":"2022-11-02"}'

```

### Step 2: Stage 0 — Adler整书理解 (Whole-Book Comprehension)

The raw transcript undergoes chunking and **Adler four-step analysis** (structure, explanation, critique, application) exactly as with traditional books. This stage is implemented in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) lines 74-78.

### Step 3: Stage 1 — Parallel Extraction with Chunking

Five sub-agents execute the same extractors on the transcript:

- **Framework extractor**
- **Principle extractor**
- **Case extractor**
- **Counter-example extractor**
- **Glossary extractor**

For very long transcripts, the pipeline applies the chunking strategy defined in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md). This ensures parallel processing without losing contextual coherence across transcript segments.

### Step 4: Stage 1.5 — Triple Verification

Extracted units pass through identical verification criteria as books: **cross-domain evidence**, **predictive power**, and **uniqueness** checks ([`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) lines 98-105).

### Step 5: Stages 2-5 — Skill Construction and Output

Validated units become atomic skills through:

- **RIA++ construction** (Stage 2)
- **Zettelkasten linking** (Stage 3)
- **Darwin-compatible pressure testing** (Stage 4)
- **Final digest generation** (Stage 5)

The system uses standard templates: `templates/SKILL.md.template`, `templates/INDEX.md.template`, and others ([`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) lines 22-30). **Timestamps or episode identifiers persist in the `source_chapter` field** to preserve source traceability ([`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) line 48).

Output structure in `books/<slug>/`:

- [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md)
- Individual [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) files
- [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md)
- [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md)
- Final [`DIGEST.md`](https://github.com/kangarooking/cangjie-skill/blob/main/DIGEST.md)

## Key Design Principles for Non-Book Content

| Principle | Implementation |
|-----------|---------------|
| **Unified pipeline** | No branching logic for transcript types—all content flows through identical stages |
| **Metadata flexibility** | Title/author/date mapping adapts to content source without schema changes |
| **Chunking for scale** | Long transcripts split intelligently per [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) |
| **Traceability preservation** | Timestamps and episode numbers retained in `source_chapter` field |

## Critical Source Files

Understanding these files clarifies how cangjie-skill generalizes across content types:

- [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) — Pipeline definition and "non-book" terminology expansion
- [`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md) — Five-stage RIA-TV++ workflow documentation
- [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) — Long-document chunking strategies
- `templates/BOOK_OVERVIEW.md.template` — Rendering template for all source types
- `extractors/*-extractor.md` — Five extractor agent definitions

## Summary

- **cangjie-skill does not distinguish between books and transcripts**—it processes any long-form text through identical stages.
- **Metadata fields adapt** to capture video/podcast specifics (timestamps replace chapters).
- **Chunking in Stage 1** handles transcript length without algorithmic changes.
- **Template-based output** remains consistent regardless of input source.
- **Source traceability** is preserved through dedicated timestamp/episode fields.

## Frequently Asked Questions

### Does cangjie-skill require different configuration for podcasts versus books?

No. Both use the same CLI invocation and pipeline configuration. Only the metadata values differ—podcasts use episode numbers where books use chapter numbers, and timestamps replace page references in the `source_chapter` field.

### What transcript formats does cangjie-skill accept?

The input validation documented in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) lines 44-48 supports **PDF, EPUB, TXT, subtitles (字幕), and raw transcripts (转写稿)**. No preprocessing is required beyond ensuring the text is extractable from these containers.

### How does cangjie-skill handle extremely long video transcripts?

Stage 1 implements automatic chunking per [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md). The parallel extraction agents process segments independently, then results merge for verification—scaling to multi-hour content without manual intervention.

### Can extracted skills from podcasts link to skills derived from books?

Yes. The Zettelkasten linking stage (Stage 3) treats all sources uniformly. Skills from video transcripts receive identical UUID-based identifiers and can form bidirectional links with skills from any other source type in the knowledge graph.