How Cangjie-Skill Distills Video and Podcast Content: A Complete Workflow Guide

Cangjie-skill handles video and podcast content by processing plain-text transcripts through its five-stage distillation pipeline rather than ingesting raw media files directly.

This open-source skill extraction system treats transcripts from videos, podcasts, courses, or interviews identically to book content. The design intentionally separates media acquisition from knowledge distillation, enabling flexible integration with any speech-to-text or subtitle extraction service.

Pre-Processing: Obtaining the Transcript

Cangjie-skill does not download videos or perform speech-to-text conversion. The repository's [README.md](https://github.com/kangarooking/cangjie-skill/blob/main/README.md) explicitly delegates this step to companion tools.

The recommended approach uses the separate video-downloader skill from the kangarooking-skills collection:


# Step 1: Download video and get transcript (video-downloader skill)

transcript_path = run_skill(
    name="video-downloader",
    inputs={"url": "https://www.youtube.com/watch?v=example"},
    output="transcript_path"
)

# Step 2: Run Cangjie-skill on the transcript

skill_output = run_skill(
    name="cangjie-skill",
    inputs={"source_path": transcript_path},
    output="skill_json"
)

print("Generated skill graph:", skill_output)

The downloader returns a local file path pointing to a plain-text transcript (.txt, .srt, or .vtt format). This decoupling lets teams swap transcription providers—Whisper, Azure Speech, or manual captions—without modifying the core distillation logic.

The Five-Stage Distillation Pipeline for Transcripts

Once a transcript is ready, Cangjie-skill executes the same workflow used for books and long-form articles. Each stage is documented in the methodology/ folder:

Stage 0: Adler

[01-stage0-adler.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md) performs sanity checks and extracts basic metadata. For video/podcast sources, this validates transcript integrity and identifies speaker markers or timestamps.

Stage 1: Parallel Extract

[02-stage1-parallel-extract.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) runs extractor modules simultaneously over the transcript. Key extractors include:

  • principle-extractor.md — isolates core principles and mental models
  • framework-extractor.md — detects methodologies and structured approaches
  • Case extractors — captures examples and counter-examples
  • Glossary extractors — identifies domain-specific terminology

These modules operate on pure text, making no distinction between a book chapter and podcast transcript.

Stage 2: RIA Plus

[04-stage2-ria-plus.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/04-stage2-ria-plus.md) refines raw extracts by removing redundancies, resolving conflicts, and adding source reference links. For video content, this stage preserves timestamps when available for traceability.

Stage 3: Zettelkasten

[05-stage3-zettelkasten.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md) organizes refined extracts into a knowledge-graph-style skill graph, establishing relationships between concepts regardless of original media format.

Stage 4: Pressure Test

[06-stage4-pressure-test.md](https://github.com/kangarocking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md) validates generated skills by simulating realistic use-cases. Video and podcast content often yields lower-density knowledge compared to technical books, so this stage iterates until quality thresholds are met.

Stage 5: Deliver

[07-stage5-deliver.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md) renders the final output as standard Claude-skill JSON with accompanying documentation.

Command-line usage follows this two-step pattern:


# 1️⃣ Get subtitles (requires ffmpeg + whisper or any STT service)

video-downloader https://example.com/podcast.mp3 --output transcript.txt

# 2️⃣ Distill the transcript into Claude skills

cangjie-skill --source transcript.txt --out skills.json

Why Video and Podcast Processing Mirrors Book Workflows

The architecture treats source type as a configuration detail, not a code branch. Once content is in plain-text form, the same extractor modules apply. This design decision in Cangjie-skill delivers three advantages:

  • Modularity — Swap transcription services without touching distillation logic
  • Testability — Unit test extractors against text fixtures, not video files
  • Scalability — Process batch transcripts without GPU-bound media processing

The SKILL.md file confirms this approach at line 44: "视频/播客建议先用 video-downloader 类工具拿到转写文本" ("For video/podcast, first use video-downloader tools to obtain the transcript").

API Integration Example

For programmatic access, the skill accepts a minimal JSON payload:

{
  "skill": "cangjie-skill",
  "input": {
    "source_path": "/tmp/podcast-transcript.txt"
  }
}

The source_path parameter accepts any UTF-8 text file, enabling pipeline integration with cloud transcription services or subtitle repositories.

Summary

  • Cangjie-skill requires pre-processed transcripts rather than raw video/audio files
  • The companion video-downloader skill handles media download and speech-to-text conversion
  • All five distillation stages—Adler, Parallel Extract, RIA Plus, Zettelkasten, Pressure Test, and Deliver—execute identically for transcripts and books
  • Extractor modules in extractors/ operate on plain text without media-specific logic
  • Final output is standard Claude-skill JSON compatible with Anthropic's tool use system

Frequently Asked Questions

Does Cangjie-skill transcribe videos automatically?

No. According to the SKILL.md source file, Cangjie-skill expects a ready-made transcript and delegates media handling to external tools like the video-downloader skill. This separation keeps the codebase focused on knowledge extraction rather than infrastructure for speech-to-text processing.

What transcript formats does Cangjie-skill accept?

The pipeline accepts standard plain-text formats including .txt, .srt, and .vtt files. The Stage 0 Adler module performs initial validation, so any UTF-8 encoded text with readable content will process correctly.

How does video content quality affect skill generation?

Video and podcast transcripts typically yield lower information density than technical books. The Stage 4 Pressure Test module specifically addresses this by simulating realistic use-cases and iterating until generated skills meet quality thresholds, ensuring usable output even from conversational content.

Can I use a different transcription service instead of video-downloader?

Yes. The architecture intentionally decouples transcription from distillation. Any service producing plain-text output—OpenAI Whisper, Azure Speech-to-Text, or manual transcription—can feed into Cangjie-skill as long as the result is written to a file path specified in the source_path parameter.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →