What Types of Content Can Be Distilled into Skills? Complete Guide to RIA-TV++ Inputs

cangjie-skill can convert any text-based, long-form material containing reusable methodology into atomic, executable Claude skills through its RIA-TV++ pipeline.

Whether you're working with books, video transcripts, podcasts, or interview recordings, the kangarooking/cangjie-skill repository provides a structured system for extracting actionable skills from unstructured content. This guide examines the six validated content categories, the architectural pipeline that processes them, and the specific file requirements needed for successful distillation.

The Six Valid Content Types for Skill Distillation

The pipeline explicitly accepts six content categories, as documented in SKILL.md (lines 41-44) and the project README. Each category must ultimately be provided as plain text for processing.

Books (PDF, EPUB, TXT)

Whole-book text provides the structural depth required for Adler analysis and the five parallel extractors. The pipeline performs best with complete works rather than excerpts, allowing the framework-extractor and principle-extractor agents to identify recurring patterns across chapters.

Source requirements:

  • Unencrypted PDF or EPUB
  • UTF-8 encoded plain text
  • Table of contents intact for structural parsing

Long-Video Transcripts (SRT, VTT, TXT)

Subtitle files or auto-generated transcriptions contain the same narrative flow as books. The extractors locate frameworks embedded in spoken dialogue, even without visual context.

Workflow note: Raw video files require pre-processing. The repository references a separate video-downloader skill (mentioned in README) to generate transcripts before distillation begins.

Podcast Transcriptions

Podcast dialogue is treated as continuous conversation. The case-extractor and counter-example-extractor agents parse turns of speech to identify decision frameworks and mental models embedded in expert discussions.

Courses and Lectures

Lecture notes, slide decks converted to text, or recorded-lecture transcripts map directly onto the R-I-A-E-B schema:

  • R (Requirement): Learning objectives
  • I (Insight): Core concepts taught
  • A1/A2 (Action variants): Exercises and implementations
  • E (Evidence): Grading rubrics or success metrics
  • B (Base): Prerequisites and foundational knowledge

Interview Scripts

Verbatim transcripts expose experts' reasoning patterns. The principle-extractor specifically targets interview content for principles and counter-examples that reveal how experts avoid common mistakes.

Long Articles and Data Collections

Markdown, HTML, or plain-text research reports qualify when they contain cohesive methodologies. The key criterion: the content must encode reusable procedures rather than ephemeral information.

The RIA-TV++ Processing Pipeline

Content undergoes seven structured stages, defined in methodology/00-overview.md. Understanding this flow clarifies why certain content types succeed or fail.

Stage 0: Adler Whole-Book Understanding

The raw text is parsed and summarized into BOOK_OVERVIEW.md. This follows the methodology in methodology/01-stage0-adler.md—a structured analytical reading that identifies the author's core arguments, terminology, and organizational logic.

Stage 1: Parallel Extraction

Five dedicated extractor agents operate simultaneously:

Agent Output Target File
framework-extractor Methodological structures candidates/frameworks/
principle-extractor Universal rules and heuristics candidates/principles/
case-extractor Concrete examples candidates/cases/
counter-example-extractor Failure modes and warnings candidates/counter-examples/
glossary-extractor Domain-specific terminology candidates/glossary/

Each agent reads both the BOOK_OVERVIEW.md and original source text, as specified in their prompt files under extractors/.

Stage 1.5: Triple Verification

Candidates must pass three filters:

  • Cross-domain evidence: Principle applies beyond original context
  • Predictive power: Enables accurate forecasting
  • Uniqueness: Not reducible to existing skills

Stage 2-5: Construction Through Delivery

Verified units become full skill definitions following the R-I-A1-A2-E-B structure in templates/SKILL.md.template, linked via templates/INDEX.md.template, and packaged with templates/DIGEST.md.template for final delivery.

Practical Invocation Examples

Distilling a Book

{
  "skill": "cangjie-skill",
  "input": {
    "source_path": "/path/to/poor-charlies-almanack.pdf",
    "metadata": {
      "title": "穷查理宝典",
      "author": "查理·芒格",
      "year": 2020,
      "type": "book"
    },
    "first_time": true
  }
}

Distilling a Video Transcript

{
  "skill": "cangjie-skill",
  "input": {
    "source_path": "/path/to/ai-for-everyone-transcript.srt",
    "metadata": {
      "title": "AI for Everyone",
      "author": "吴恩达",
      "publish_date": "2023-04-15",
      "type": "video"
    },
    "first_time": false
  }
}

Both invocations trigger identical pipeline behavior, producing this output structure:


books/[normalized-title]/
├── BOOK_OVERVIEW.md
├── INDEX.md
├── GLOSSARY.md
├── DIGEST.md
├── candidates/
│   ├── frameworks/
│   ├── principles/
│   ├── cases/
│   ├── counter-examples/
│   └── glossary/
└── skills/
    └── [skill-name]/
        ├── SKILL.md
        ├── test-prompts.json
        └── test-results.md

Critical Input Requirements

The pipeline enforces strict text-based processing. Content types that cannot be directly distilled include:

  • Raw audio files (MP3, WAV)
  • Raw video files (MP4, MKV)
  • Scanned image PDFs without OCR
  • Proprietary document formats

Pre-processing tools must convert these to plain text before invocation.

Key Source Files Reference

File Purpose
[SKILL.md](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) Meta-skill definition with input requirements (lines 41-44)
[methodology/00-overview.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md) RIA-TV++ architectural overview
[methodology/01-stage0-adler.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md) Adler analytical reading specification
[extractors/framework-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md) Parallel extraction agent prompts
templates/SKILL.md.template Atomic skill structure template

Summary

  • Six validated content types: books, long-video transcripts, podcasts, courses/lectures, interviews, and long articles/data collections
  • Universal requirement: Must be convertible to plain text before pipeline entry
  • Core mechanism: Five parallel extractors (framework-, principle-, case-, counter-example-, glossary-extractor) process source material through seven stages
  • Output format: Executable skills in SKILL.md files with test-prompts.json verification
  • Entry point: JSON invocation with source_path and metadata to cangjie-skill

Frequently Asked Questions

Can I distill a YouTube video directly without downloading it?

No. The pipeline requires a local text file as source_path. You must first use the referenced video-downloader skill or equivalent tool to generate an SRT, VTT, or TXT transcript. The cangjie-skill system does not handle network requests or streaming extraction.

Does content length matter for successful extraction?

Yes. The Adler analysis in Stage 0 requires sufficient material to identify structural patterns. Single short articles often fail verification due to insufficient depth. Books, full course transcripts, or multi-episode podcast series provide better candidates for the five parallel extractors to locate reusable methodology.

How does the system handle non-English content?

The pipeline processes UTF-8 encoded text regardless of language. The extractor agents analyze semantic patterns rather than specific English keywords. However, the output skills currently generate English SKILL.md files regardless of source language, as implemented in templates/SKILL.md.template.

What happens if my content lacks clear frameworks or principles?

Stage 1.5 triple verification rejects candidates lacking cross-domain evidence or predictive power. The pipeline explicitly avoids converting informational content without reusable methodology. If no candidates pass verification, the candidates/ folder remains empty and no skills are constructed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →