How to Distill Video and Podcast Transcripts into Reusable AI Skills

To distill video and podcast transcripts into reusable AI skills, feed the transcribed text through the RIA-TV++ pipeline implemented in the cangjie-skill repository, which processes any structured textual source through seven stages to generate atomic, callable skills.

The cangjie-skill repository provides a systematic methodology for converting structured text into executable AI capabilities. While the approach is often demonstrated with books, the RIA-TV++ pipeline treats video subtitles and podcast transcripts as equivalent inputs, requiring no structural changes to the distillation process. As implemented in kangaroking/cangjie-skill, the pipeline guarantees that resulting skills are traceable, pressure-tested, and ready for immediate deployment in agent environments.

The RIA-TV++ Pipeline for Transcripts

The pipeline processes transcripts through seven distinct stages defined in the repository's methodology/ directory. When working with video or audio content, you first obtain a transcript file via speech-to-text or subtitle extraction, then invoke the same pipeline used for book distillation by specifying source_type="transcript".

Stage 0 – Adler Overview

Defined in methodology/01-stage0-adler.md, the Adler Overview stage reads the entire transcript to determine content type, identify the main thesis, and outline the logical skeleton. This creates the structural foundation for downstream extraction, treating podcast transcripts and video subtitles identically to book chapters.

Stage 1 – Parallel Extraction

Five specialized extractors defined in the extractors/ folder run concurrently over the transcript text:

  • Principle Extractor (extractors/principle-extractor.md) – Pulls rules, checklists, and maxims
  • Framework Extractor – Identifies thinking models and decision frameworks
  • Case Extractor – Captures concrete examples mentioned by speakers
  • Counter-Example Extractor – Flags warned-against practices and anti-patterns
  • Glossary Extractor – Builds a domain-specific term dictionary

Each extractor scans the transcript independently, creating candidate skill units based on distinct knowledge patterns.

Stage 1.5 – Triple Verification

Every candidate skill must satisfy three criteria before advancing:

  1. Cross-Domain Evidence – At least two independent citations within the transcript
  2. Predictive Power – Ability to answer questions not explicitly asked in the source
  3. Uniqueness – Distinction from common-sense statements

This verification step filters noise from conversational content, ensuring only substantive knowledge becomes encoded skills.

Stage 2 – RIA++ Construction

Verified units are converted into structured SKILL.md files following the template in templates/SKILL.md.template. Each skill contains six mandatory fields: R (raw quote), I (paraphrase), A1 (example), A2 (trigger), E (execution step), and B (boundary). This schema is defined in methodology/02-stage0-adler.md.

Stage 3 – Zettelkasten Linking

Skills are interconnected in an INDEX.md file using atomic note principles outlined in methodology/05-stage3-zettelkasten.md. This enables discovery of dependencies between concepts extracted from different segments of the transcript.

Stage 4 – Pressure Testing

Automated test prompts in test-prompts.json verify that each skill fires correctly and maintains distinction from other skills. Failed skills return to earlier stages for revision.

Stage 5 – Delivery

As specified in methodology/07-stage5-deliver.md, the pipeline generates a human-readable DIGEST.md summary and places skill files into a skills/ directory compatible with Claude Code and Cursor. The templates/DIGEST.md.template and templates/INDEX.md.template standardize these outputs.

Running the Pipeline on Transcript Files

To process a video or podcast transcript, replace the book input with your transcript file path. The pipeline accepts the source_type="transcript" parameter to optimize Stage 0 analysis for conversational content.

from cangjie_skill.pipeline import run_pipeline

transcript_path = "data/podcast_episode_42.md"

pipeline_output = run_pipeline(
    source_type="transcript",
    input_path=transcript_path,
    output_dir="skills/podcast-42"
)

print(f"Generated {len(pipeline_output.skills)} skills")
print(f"Digest available at: {pipeline_output.digest_path}")

For command-line usage, the repository provides a CLI entry point:

cangjie-skill \
  --type transcript \
  --input data/video_subtitles.md \
  --outdir skills/video-123

Both methods invoke the internal stages automatically, producing SKILL.md, INDEX.md, DIGEST.md, and test-prompts.json for the transcript.

Key Source Files and Templates

The following files govern how transcripts transform into skills:

  • methodology/01-stage0-adler.md – Defines Stage 0 logic for analyzing whole-text sources
  • extractors/principle-extractor.md – Configuration for rule extraction patterns
  • templates/SKILL.md.template – Schema template for atomic skill generation
  • templates/INDEX.md.template – Zettelkasten linking structure
  • templates/DIGEST.md.template – Human-readable summary format

Summary

  • The cangjie-skill repository processes video and podcast transcripts through the same RIA-TV++ pipeline used for books, requiring only that you substitute the input text file.
  • The seven-stage pipeline includes parallel extraction by five specialized extractors, triple verification for quality control, and automated pressure testing.
  • Output skills follow the R-I-A1-A2-E-B schema defined in SKILL.md templates, ensuring atomicity and traceability.
  • Generated files include atomic SKILL.md units, an INDEX.md for navigation, a DIGEST.md summary, and test-prompts.json for validation.

Frequently Asked Questions

What file formats are supported for transcript input?

The pipeline accepts any plain-text format, including Markdown files, SRT subtitle files, and raw text exports from speech-to-text services. As noted in methodology/01-stage0-adler.md, the system treats the transcript as structured text regardless of original medium, though Markdown with speaker timestamps often yields better context preservation during extraction.

How does triple verification handle conversational filler content?

The Stage 1.5 verification criteria explicitly exclude common-sense statements through the Uniqueness check, while the Cross-Domain Evidence requirement ensures each skill has at least two independent citations in the transcript. This filters out off-topic banter common in podcasts, allowing only substantive frameworks and principles to advance to SKILL.md generation.

Can the pipeline process live stream transcripts in real-time?

The repository is designed for post-processing complete transcripts rather than real-time streams. The Adler Overview in Stage 0 requires access to the full text to determine the logical skeleton and main thesis. For live content, capture the complete transcript first, then batch-process it through the pipeline using the CLI or Python interface.

What distinguishes RIA-TV++ from standard prompt engineering?

RIA-TV++ enforces atomicity through the six-field schema and guarantees execution readiness via Stage 4 pressure testing with test-prompts.json. Unlike ad-hoc prompts, skills generated from transcripts include explicit boundaries (B field) and triggers (A2 field) parsed from conversational context, making them callable units in agent environments rather than static instructions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →