# What Types of Content Can Be Distilled into Skills? Complete Guide to RIA-TV++ Inputs

> Learn how cangjie-skill transforms text content into executable Claude skills. Discover the RIA-TV++ pipeline and distill your reusable methodology into atomic skills.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: getting-started
- Published: 2026-08-14

---

**cangjie-skill can convert any text-based, long-form material containing reusable methodology into atomic, executable Claude skills through its RIA-TV++ pipeline.**

Whether you're working with books, video transcripts, podcasts, or interview recordings, the `kangarooking/cangjie-skill` repository provides a structured system for extracting actionable skills from unstructured content. This guide examines the six validated content categories, the architectural pipeline that processes them, and the specific file requirements needed for successful distillation.

## The Six Valid Content Types for Skill Distillation

The pipeline explicitly accepts six content categories, as documented in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) (lines 41-44) and the project README. Each category must ultimately be provided as **plain text** for processing.

### Books (PDF, EPUB, TXT)

Whole-book text provides the structural depth required for **Adler analysis** and the five parallel extractors. The pipeline performs best with complete works rather than excerpts, allowing the `framework-extractor` and `principle-extractor` agents to identify recurring patterns across chapters.

**Source requirements:**
- Unencrypted PDF or EPUB
- UTF-8 encoded plain text
- Table of contents intact for structural parsing

### Long-Video Transcripts (SRT, VTT, TXT)

Subtitle files or auto-generated transcriptions contain the same narrative flow as books. The extractors locate frameworks embedded in spoken dialogue, even without visual context.

**Workflow note:** Raw video files require pre-processing. The repository references a separate **`video-downloader`** skill (mentioned in README) to generate transcripts before distillation begins.

### Podcast Transcriptions

Podcast dialogue is treated as continuous conversation. The `case-extractor` and `counter-example-extractor` agents parse turns of speech to identify decision frameworks and mental models embedded in expert discussions.

### Courses and Lectures

Lecture notes, slide decks converted to text, or recorded-lecture transcripts map directly onto the **R-I-A-E-B schema**:
- **R** (Requirement): Learning objectives
- **I** (Insight): Core concepts taught
- **A1/A2** (Action variants): Exercises and implementations
- **E** (Evidence): Grading rubrics or success metrics
- **B** (Base): Prerequisites and foundational knowledge

### Interview Scripts

Verbatim transcripts expose experts' reasoning patterns. The `principle-extractor` specifically targets interview content for **principles** and **counter-examples** that reveal how experts avoid common mistakes.

### Long Articles and Data Collections

Markdown, HTML, or plain-text research reports qualify when they contain cohesive methodologies. The key criterion: the content must encode **reusable procedures** rather than ephemeral information.

## The RIA-TV++ Processing Pipeline

Content undergoes seven structured stages, defined in [`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md). Understanding this flow clarifies why certain content types succeed or fail.

### Stage 0: Adler Whole-Book Understanding

The raw text is parsed and summarized into [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md). This follows the methodology in [`methodology/01-stage0-adler.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md)—a structured analytical reading that identifies the author's core arguments, terminology, and organizational logic.

### Stage 1: Parallel Extraction

Five dedicated extractor agents operate simultaneously:

| Agent | Output | Target File |
|-------|--------|-------------|
| `framework-extractor` | Methodological structures | `candidates/frameworks/` |
| `principle-extractor` | Universal rules and heuristics | `candidates/principles/` |
| `case-extractor` | Concrete examples | `candidates/cases/` |
| `counter-example-extractor` | Failure modes and warnings | `candidates/counter-examples/` |
| `glossary-extractor` | Domain-specific terminology | `candidates/glossary/` |

Each agent reads both the [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) and original source text, as specified in their prompt files under `extractors/`.

### Stage 1.5: Triple Verification

Candidates must pass three filters:
- **Cross-domain evidence:** Principle applies beyond original context
- **Predictive power:** Enables accurate forecasting
- **Uniqueness:** Not reducible to existing skills

### Stage 2-5: Construction Through Delivery

Verified units become full skill definitions following the **R-I-A1-A2-E-B** structure in `templates/SKILL.md.template`, linked via `templates/INDEX.md.template`, and packaged with `templates/DIGEST.md.template` for final delivery.

## Practical Invocation Examples

### Distilling a Book

```json
{
  "skill": "cangjie-skill",
  "input": {
    "source_path": "/path/to/poor-charlies-almanack.pdf",
    "metadata": {
      "title": "穷查理宝典",
      "author": "查理·芒格",
      "year": 2020,
      "type": "book"
    },
    "first_time": true
  }
}

```

### Distilling a Video Transcript

```json
{
  "skill": "cangjie-skill",
  "input": {
    "source_path": "/path/to/ai-for-everyone-transcript.srt",
    "metadata": {
      "title": "AI for Everyone",
      "author": "吴恩达",
      "publish_date": "2023-04-15",
      "type": "video"
    },
    "first_time": false
  }
}

```

Both invocations trigger identical pipeline behavior, producing this output structure:

```

books/[normalized-title]/
├── BOOK_OVERVIEW.md
├── INDEX.md
├── GLOSSARY.md
├── DIGEST.md
├── candidates/
│   ├── frameworks/
│   ├── principles/
│   ├── cases/
│   ├── counter-examples/
│   └── glossary/
└── skills/
    └── [skill-name]/
        ├── SKILL.md
        ├── test-prompts.json
        └── test-results.md

```

## Critical Input Requirements

The pipeline enforces strict text-based processing. Content types that **cannot** be directly distilled include:
- Raw audio files (MP3, WAV)
- Raw video files (MP4, MKV)
- Scanned image PDFs without OCR
- Proprietary document formats

Pre-processing tools must convert these to plain text before invocation.

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [[`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) | Meta-skill definition with input requirements (lines 41-44) |
| [[`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md)](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md) | RIA-TV++ architectural overview |
| [[`methodology/01-stage0-adler.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md)](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md) | Adler analytical reading specification |
| [[`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md) | Parallel extraction agent prompts |
| [`templates/SKILL.md.template`](https://github.com/kangarooking/cangjie-skill/blob/main/templates/SKILL.md.template) | Atomic skill structure template |

## Summary

- **Six validated content types:** books, long-video transcripts, podcasts, courses/lectures, interviews, and long articles/data collections
- **Universal requirement:** Must be convertible to plain text before pipeline entry
- **Core mechanism:** Five parallel extractors (`framework-`, `principle-`, `case-`, `counter-example-`, `glossary-extractor`) process source material through seven stages
- **Output format:** Executable skills in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) files with [`test-prompts.json`](https://github.com/kangarooking/cangjie-skill/blob/main/test-prompts.json) verification
- **Entry point:** JSON invocation with `source_path` and metadata to `cangjie-skill`

## Frequently Asked Questions

### Can I distill a YouTube video directly without downloading it?

No. The pipeline requires a local text file as `source_path`. You must first use the referenced **`video-downloader`** skill or equivalent tool to generate an SRT, VTT, or TXT transcript. The cangjie-skill system does not handle network requests or streaming extraction.

### Does content length matter for successful extraction?

Yes. The Adler analysis in Stage 0 requires sufficient material to identify structural patterns. Single short articles often fail verification due to insufficient depth. Books, full course transcripts, or multi-episode podcast series provide better candidates for the five parallel extractors to locate reusable methodology.

### How does the system handle non-English content?

The pipeline processes UTF-8 encoded text regardless of language. The extractor agents analyze semantic patterns rather than specific English keywords. However, the output skills currently generate English [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) files regardless of source language, as implemented in `templates/SKILL.md.template`.

### What happens if my content lacks clear frameworks or principles?

Stage 1.5 triple verification rejects candidates lacking cross-domain evidence or predictive power. The pipeline explicitly avoids converting informational content without reusable methodology. If no candidates pass verification, the `candidates/` folder remains empty and no skills are constructed.