# What Data Sources Does cangjie-skill Support? Input Formats and Architecture Guide

> Explore supported data sources for cangjie-skill, including PDFs EPUBs and text files. Learn about its unified processing pipeline for AI skills.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: architecture
- Published: 2026-07-17

---

**cangjie-skill ingests any plain-text representation of long-form content—including PDFs, EPUBs, raw text files, subtitle tracks, and remote URLs—and transforms them into executable AI skills through a unified, source-agnostic processing pipeline.**

The kangarooking/cangjie-skill project is designed to distill books, video courses, podcasts, and lectures into structured, verifiable AI capabilities. Understanding what data sources cangjie-skill interacts with is essential for configuring the pipeline, as the system normalizes every input into a canonical "book" representation regardless of its original format.

## Supported Input Formats

According to the meta-skill definition in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) (lines 44-46), cangjie-skill explicitly accepts the following content source types:

> "内容文本来源: **PDF / EPUB / TXT / 字幕文件 / 转写稿路径**, 或可访问的纯文本"

This translates to a flexible ingestion model that handles both local files and remote resources.

### Document Files (PDF and EPUB)

**PDF** files undergo conversion to plain text before entering Stage 0 (Adler analysis), while **EPUB** files are unpacked and rendered into raw text streams. Both formats preserve structural metadata—such as chapter boundaries—that the pipeline maps to `source_chapter` identifiers during extraction.

### Plain Text and Transcription Files

**TXT** files are read block-by-block during Stage 0 and by the parallel extractors in Stage 1. The pipeline also accepts **subtitle files** (`.srt`, `.vtt`) and pre-generated **transcription files** containing full video or podcast transcripts. These are treated as the primary textual source for non-book media.

### Remote Plain-Text URLs

Any accessible **remote URL** returning plain text can be fetched and streamed directly into the pipeline. This enables processing of hosted transcripts or chapter files without local storage.

## How cangjie-skill Processes Different Data Sources

The architecture treats *any* non-book content—whether a video series, podcast, or online course—as a "book" by mapping its structural metadata onto the same schema used for printed works (as implemented in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md), lines 48-49). This normalization enables five specialized extractors to operate uniformly across all source types.

### Stage 0: Whole-Book Comprehension

During **Stage 0 – Whole-Book Comprehension**, the pipeline reads the supplied raw text (splitting large files into manageable blocks when necessary) and produces [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md). The logic for this ingestion and chunking is detailed in [`methodology/01-stage0-adler.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md).

### Stage 1: Parallel Extraction

**Stage 1** runs five parallel extractors—**framework**, **principle**, **case**, **counter-example**, and **glossary**—on the same raw text stream. Each extractor tags candidates with `source_chapter` metadata (or analogous video timestamps/episode identifiers), making the extraction logic independent of the original file format.

### Stage 1.5: Triple Verification

The **Triple Verification** stage validates candidates by cross-referencing them within the source text. Because verification operates on the normalized candidate objects carrying source metadata, it remains agnostic to whether the original input was a PDF, subtitle file, or remote URL.

## Configuring Data Sources in cangjie-skill

The following YAML examples illustrate the expected input configuration for the cangjie-skill pipeline:

```yaml

# Example 1 – Distilling a PDF book

cangjie:
  source_path: "/path/to/《穷查理宝典》.pdf"
  metadata:
    title: "穷查理宝典"
    author: "查理·芒格"
    year: 2020

```

```yaml

# Example 2 – Distilling a video transcript

cangjie:
  source_path: "/path/to/lecture_transcript.txt"
  metadata:
    title: "AI Product Design Lecture"
    speaker: "Jane Doe"
    date: "2023-05-12"
    source_chapter: "00:00–15:30"

```

In both configurations, the `source_path` key specifies the location of the raw text, while the `metadata` block provides auditing and naming context. The same structure applies to EPUB files, subtitle files (`.srt`), or remote plain-text URLs.

## Key Architecture Files

The source-agnostic design is implemented across the following files in the kangarooking/cangjie-skill repository:

- **[`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)** – Contains the meta-skill definition and explicit list of accepted source types (lines 44-49).
- **[`README.en.md`](https://github.com/kangarooking/cangjie-skill/blob/main/README.en.md)** – Provides high-level documentation of the pipeline and supported input formats (lines 24-28).
- **[`methodology/01-stage0-adler.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md)** – Defines how raw text is ingested and chunked for analysis.
- **`extractors/*.md`** – Houses the extractor prompts that operate on the generic text stream, independent of source format.
- **`templates/BOOK_OVERVIEW.md.template`** – Receives normalized content from any source type during Stage 0.

## Summary

- **cangjie-skill** accepts **PDF, EPUB, TXT, subtitle files, transcripts, and remote plain-text URLs** as valid inputs.
- The pipeline normalizes all sources into a unified "book" structure, enabling the same extraction logic to work across books, videos, podcasts, and courses.
- **Stage 0** handles ingestion and chunking, while **Stages 1 and 1.5** perform extraction and verification using format-agnostic metadata tagging.
- Configuration requires only a `source_path` and minimal metadata, regardless of the original file type.

## Frequently Asked Questions

### Does cangjie-skill support direct video or audio file processing?

No. The pipeline requires pre-generated textual representations such as subtitle files (`.srt`, `.vtt`) or full transcripts saved as `.txt` files. The audio or video itself must be transcribed before ingestion.

### What happens to chapter structure when processing subtitle files?

When processing subtitle or transcript files, the pipeline maps timestamp ranges (e.g., "00:00–15:30") or episode numbers onto the `source_chapter` field. This allows the five extractors to tag provenance accurately, treating temporal segments analogously to book chapters.

### Can cangjie-skill authenticate against protected URLs?

The analysis of [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) and related configuration files indicates that remote URLs must be publicly accessible plain-text endpoints. Authentication mechanisms or API keys are not documented in the current input requirements.

### How does the pipeline handle extremely large PDF textbooks?

During **Stage 0 (Adler analysis)**, large documents are automatically split into manageable text blocks before processing. The [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) template and subsequent extractor prompts in `extractors/*.md` are designed to handle chunked content while maintaining cross-reference integrity during **Stage 1.5 verification**.