What Data Sources Does cangjie-skill Support? Input Formats and Architecture Guide
cangjie-skill ingests any plain-text representation of long-form content—including PDFs, EPUBs, raw text files, subtitle tracks, and remote URLs—and transforms them into executable AI skills through a unified, source-agnostic processing pipeline.
The kangarooking/cangjie-skill project is designed to distill books, video courses, podcasts, and lectures into structured, verifiable AI capabilities. Understanding what data sources cangjie-skill interacts with is essential for configuring the pipeline, as the system normalizes every input into a canonical "book" representation regardless of its original format.
Supported Input Formats
According to the meta-skill definition in SKILL.md (lines 44-46), cangjie-skill explicitly accepts the following content source types:
"内容文本来源: PDF / EPUB / TXT / 字幕文件 / 转写稿路径, 或可访问的纯文本"
This translates to a flexible ingestion model that handles both local files and remote resources.
Document Files (PDF and EPUB)
PDF files undergo conversion to plain text before entering Stage 0 (Adler analysis), while EPUB files are unpacked and rendered into raw text streams. Both formats preserve structural metadata—such as chapter boundaries—that the pipeline maps to source_chapter identifiers during extraction.
Plain Text and Transcription Files
TXT files are read block-by-block during Stage 0 and by the parallel extractors in Stage 1. The pipeline also accepts subtitle files (.srt, .vtt) and pre-generated transcription files containing full video or podcast transcripts. These are treated as the primary textual source for non-book media.
Remote Plain-Text URLs
Any accessible remote URL returning plain text can be fetched and streamed directly into the pipeline. This enables processing of hosted transcripts or chapter files without local storage.
How cangjie-skill Processes Different Data Sources
The architecture treats any non-book content—whether a video series, podcast, or online course—as a "book" by mapping its structural metadata onto the same schema used for printed works (as implemented in SKILL.md, lines 48-49). This normalization enables five specialized extractors to operate uniformly across all source types.
Stage 0: Whole-Book Comprehension
During Stage 0 – Whole-Book Comprehension, the pipeline reads the supplied raw text (splitting large files into manageable blocks when necessary) and produces BOOK_OVERVIEW.md. The logic for this ingestion and chunking is detailed in methodology/01-stage0-adler.md.
Stage 1: Parallel Extraction
Stage 1 runs five parallel extractors—framework, principle, case, counter-example, and glossary—on the same raw text stream. Each extractor tags candidates with source_chapter metadata (or analogous video timestamps/episode identifiers), making the extraction logic independent of the original file format.
Stage 1.5: Triple Verification
The Triple Verification stage validates candidates by cross-referencing them within the source text. Because verification operates on the normalized candidate objects carrying source metadata, it remains agnostic to whether the original input was a PDF, subtitle file, or remote URL.
Configuring Data Sources in cangjie-skill
The following YAML examples illustrate the expected input configuration for the cangjie-skill pipeline:
# Example 1 – Distilling a PDF book
cangjie:
source_path: "/path/to/《穷查理宝典》.pdf"
metadata:
title: "穷查理宝典"
author: "查理·芒格"
year: 2020
# Example 2 – Distilling a video transcript
cangjie:
source_path: "/path/to/lecture_transcript.txt"
metadata:
title: "AI Product Design Lecture"
speaker: "Jane Doe"
date: "2023-05-12"
source_chapter: "00:00–15:30"
In both configurations, the source_path key specifies the location of the raw text, while the metadata block provides auditing and naming context. The same structure applies to EPUB files, subtitle files (.srt), or remote plain-text URLs.
Key Architecture Files
The source-agnostic design is implemented across the following files in the kangarooking/cangjie-skill repository:
SKILL.md– Contains the meta-skill definition and explicit list of accepted source types (lines 44-49).README.en.md– Provides high-level documentation of the pipeline and supported input formats (lines 24-28).methodology/01-stage0-adler.md– Defines how raw text is ingested and chunked for analysis.extractors/*.md– Houses the extractor prompts that operate on the generic text stream, independent of source format.templates/BOOK_OVERVIEW.md.template– Receives normalized content from any source type during Stage 0.
Summary
- cangjie-skill accepts PDF, EPUB, TXT, subtitle files, transcripts, and remote plain-text URLs as valid inputs.
- The pipeline normalizes all sources into a unified "book" structure, enabling the same extraction logic to work across books, videos, podcasts, and courses.
- Stage 0 handles ingestion and chunking, while Stages 1 and 1.5 perform extraction and verification using format-agnostic metadata tagging.
- Configuration requires only a
source_pathand minimal metadata, regardless of the original file type.
Frequently Asked Questions
Does cangjie-skill support direct video or audio file processing?
No. The pipeline requires pre-generated textual representations such as subtitle files (.srt, .vtt) or full transcripts saved as .txt files. The audio or video itself must be transcribed before ingestion.
What happens to chapter structure when processing subtitle files?
When processing subtitle or transcript files, the pipeline maps timestamp ranges (e.g., "00:00–15:30") or episode numbers onto the source_chapter field. This allows the five extractors to tag provenance accurately, treating temporal segments analogously to book chapters.
Can cangjie-skill authenticate against protected URLs?
The analysis of SKILL.md and related configuration files indicates that remote URLs must be publicly accessible plain-text endpoints. Authentication mechanisms or API keys are not documented in the current input requirements.
How does the pipeline handle extremely large PDF textbooks?
During Stage 0 (Adler analysis), large documents are automatically split into manageable text blocks before processing. The BOOK_OVERVIEW.md template and subsequent extractor prompts in extractors/*.md are designed to handle chunked content while maintaining cross-reference integrity during Stage 1.5 verification.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →