How Cangjie-Skill Handles Non-Book Content: Courses, Interviews, and Podcasts Explained

Cangjie-skill treats any long-form material as a "book," applying the same five-stage RIA++ pipeline to courses, interviews, podcasts, and video transcripts with only minor metadata adjustments for source tracking.

The kangarooking/cangjie-skill repository is designed around a deliberately broad definition of 书 (book) that encompasses far more than traditional printed works. According to the core specification in SKILL.md, this abstraction intentionally includes books, long-video transcripts, podcast transcripts, courses, interviews, long articles, and collections【https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md#L12】. This unified design eliminates special-case handling and ensures consistent skill extraction across all content types.

The Unified "Book" Abstraction

Rather than building separate pipelines for different media types, Cangjie-skill normalizes all long-form content into a single abstraction. The identical five-stage pipeline runs regardless of source material:

  1. Adler overview — high-level structural analysis
  2. Parallel extraction — framework, principle, case, counter-example, and glossary mining
  3. Triple verification — cross-checking extracted elements
  4. RIA++ skill construction — building the reusable skill artefact
  5. Zettelkasten linking, pressure testing, and delivery

This pipeline is detailed in methodology/00-overview.md and applies unchanged to every content type【https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md】.

Metadata Mapping for Non-Book Sources

The only meaningful differences between content types appear in source metadata, specifically how the source_chapter field is interpreted. This mapping is explicitly defined in the specification【https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md#L48】:

Content type source_chapter interpretation
Video transcript Time-stamp or part number
Podcast transcript Episode number
Course material Lecture / session number
Interview Segment identifier (e.g., "Q5")

When processing non-book content, the skill first collects content metadata: title, author (or presenter/host), and publish date. These fields drive folder naming and maintain audit trails throughout the pipeline【https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md#L45】.

Stage 0: The Content Type Field

During Stage 0 (Adler overview), the BOOK_OVERVIEW.md template explicitly captures the source type through a dedicated Content Type field. The template line rendering this field appears at line 10 of templates/BOOK_OVERVIEW.md.template【https://github.com/kangarooking/cangjie-skill/blob/main/templates/BOOK_OVERVIEW.md.template#L10】.

This ensures that whether the source is labeled "课程" (course), "访谈" (interview), or "播客" (podcast), the generated overview records this distinction for downstream traceability.

Extractor Agnosticism

All downstream extractors operate on raw text only, making them media-agnostic by design. The extractor specifications in extractors/*.md — including framework-extractor.md, principle-extractor.md, and others — process plain-text transcripts without knowledge of whether the text originated from PDF, video subtitle files, or lecture slide decks【https://github.com/kangarooking/cangjie-skill/tree/main/extractors】.

This design choice eliminates extractor complexity and guarantees consistent output quality across formats.

Skill Artefact Uniformity

Generated skills follow identical structure regardless of source. Every skill artefact — including SKILL.md and test-prompts.json — contains the same R-I-A1-A2-E-B sections (Reading, Interpretation, Appropriation 1, Appropriation 2, Entanglement, Bibliography). Storage uses title-author slugs, not media-type paths, enabling downstream agents like darwin-skill to invoke a course-skill or interview-skill exactly as they would a traditional book-skill.

The skill template blueprint resides at templates/SKILL.md.template【https://github.com/kangarooking/cangjie-skill/blob/main/templates/SKILL.md.template】.

Practical Example: Processing a Course

Below is a minimal JSON payload for initiating Cangjie-skill on course material:

{
  "source_type": "course",
  "title": "Effective Communication",
  "author": "John Doe",
  "publish_date": "2023-05-01",
  "source_file": "effective_communication.txt",
  "source_chapter": "Lecture 4 (12:35-15:20)"
}

A typical invocation prompt:


请把以下课程内容蒸馏成 skill:
{payload_above}

The agent then executes:

  1. Creates books/effective-communication/ (slug from title/author)
  2. Generates BOOK_OVERVIEW.md with Content Type set to "课程"
  3. Runs five parallel extractors on the transcript
  4. Completes verification, RIA++ construction, Zettelkasten linking, pressure testing
  5. Delivers DIGEST.md and the installable skill package

Key Implementation Files

File Purpose
SKILL.md Core specification defining the unified "book" concept and non-book field mapping【https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md】
templates/BOOK_OVERVIEW.md.template Stage-0 overview template with Content Type placeholder【https://github.com/kangarooking/cangjie-skill/blob/main/templates/BOOK_OVERVIEW.md.template】
methodology/00-overview.md RIA-TV++ pipeline description applied universally【https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md】
extractors/*.md Media-agnostic parallel extraction agents【https://github.com/kangarooking/cangjie-skill/tree/main/extractors】
templates/SKILL.md.template Standardized skill structure blueprint【https://github.com/kangarooking/cangjie-skill/blob/main/templates/SKILL.md.template】

Summary

  • Unified abstraction — Cangjie-skill treats all long-form content as "books," eliminating format-specific pipelines
  • Minimal metadata variation — Only source_chapter interpretation changes across content types
  • Pipeline consistency — The same RIA++ extraction, verification, and construction stages run for all sources
  • Agent interoperability — Generated skills are media-agnostic and callable by downstream systems without modification
  • Template-driven traceability — Content type is captured at overview stage for audit purposes without affecting processing

Frequently Asked Questions

Can Cangjie-skill process YouTube video transcripts?

Yes. Video transcripts are treated identically to other long-form content. The source_chapter field typically contains a timestamp or part number, and the extractor pipeline processes the transcribed text without knowing its video origin.

Does interview content require special formatting or segmentation?

No special formatting is required. Segments can be identified using arbitrary identifiers (e.g., "Q5" for question 5) in the source_chapter field. The extractors parse the raw transcript regardless of dialogue structure.

Are skills generated from courses compatible with book-based skill libraries?

Yes. The output structure is identical. Skills are stored under title-author slugs and contain the same R-I-A1-A2-E-B sections, making them fully interchangeable with traditional book skills in any downstream agent system.

What metadata is strictly required for non-book content processing?

The skill requires three core metadata fields: title, author (or presenter/host), and publish date. The source_type and source_chapter fields provide traceability but do not alter the processing pipeline itself.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →