Cangjie-Skill Stage 1 Parallel Extractors: What Are the Five Roles and How They Work
The five parallel extractors in Cangjie-Skill Stage 1 are framework-extractor, principle-extractor, case-extractor, counter-example-extractor, and glossary-extractor—each independently scanning the same book for distinct knowledge types to maximize coverage.
The Cangjie-Skill pipeline transforms books into structured skills through a multi-stage process. In Stage 1, the challenge of "reading the book" is distributed across five specialized sub-agents that operate in parallel. According to the methodology/02-stage1-parallel-extract.md source file, all five extractors receive identical inputs—BOOK_OVERVIEW.md, the raw book text, and their respective extractor prompts—but each applies a unique lens to surface different knowledge artifacts.
The Five Parallel Extractor Roles
Framework-Extractor: Thinking Models and Decision Methods
The framework-extractor identifies mental models, decision frameworks, and reasoning methods the author presents.
This extractor surfaces structured approaches to problem-solving: step-by-step processes, analytical frameworks, and thinking patterns that readers can apply broadly. Its output lands in candidates/frameworks.md as a YAML-structured list of frameworks.
Key prompt file: [extractors/framework-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md)
Principle-Extractor: Rules, Checklists, and Assertions
The principle-extractor captures principles, checklists, rules, and explicit assertions.
These are prescriptive statements—"always do X," "never do Y," numbered lists, and heuristic guidelines that the author treats as foundational truths. Output goes to candidates/principles.md.
Key prompt file: [extractors/principle-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md)
Case-Extractor: Concrete Examples from the Text
The case-extractor pulls concrete examples the author actually uses in the book.
These are specific scenarios, stories, applications, or worked examples that illustrate broader points. Unlike hypothetical illustrations, these are instances explicitly deployed by the author. Results saved to candidates/cases.md.
Key prompt file: [extractors/case-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md)
Counter-Example-Extractor: Warnings and Anti-Patterns
The counter-example-extractor hunts for warnings, failures, anti-patterns, and traps the author highlights.
This extractor surfaces negative knowledge: common mistakes, pitfalls to avoid, failed approaches, and cautionary tales. Often overlooked by other extractors, this "inverse" knowledge is critical for robust skill-building. Output: candidates/counter-examples.md.
Key prompt file: [extractors/counter-example-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/counter-example-extractor.md)
Glossary-Extractor: Key Concepts and Terminology
The glossary-extractor builds a dictionary of key concepts and specialized terminology.
This captures definitions, technical terms, recurring concepts, and the author's specific vocabulary—establishing the semantic foundation for later stages. Output: candidates/glossary.md.
Key prompt file: [extractors/glossary-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md)
Why Parallel Extraction? Three Design Rationale
The parallel architecture in Stage 1 serves three strategic purposes:
- Coverage through diversity — Different specialized lenses catch items others miss; a counter-example invisible to the framework-extractor gets flagged by its dedicated counterpart
- Speed through concurrency — Claude-code agents spawn simultaneously, reducing wall-clock runtime versus sequential processing
- Independence prevents contamination — Each extractor judges in isolation without seeing others' results; Stage 1.5 (V1 cross-domain verification) later merges overlapping findings intentionally
Running the Extractors: Example Commands
Each extractor follows an identical invocation pattern, differing only in prompt and output paths:
# Framework extractor
python run_extractor.py \
--overview BOOK_OVERVIEW.md \
--text book.txt \
--prompt extractors/framework-extractor.md \
--output candidates/frameworks.md
# Principle extractor (same pattern, different files)
python run_extractor.py \
--overview BOOK_OVERVIEW.md \
--text book.txt \
--prompt extractors/principle-extractor.md \
--output candidates/principles.md
All extractors produce YAML-structured candidates. Example output format:
id: f01
title: 逆向思维
type: framework
source_chapter: 第三讲
source_quote: |
"反过来想,总是反过来想..."
summary: |
The author proposes a reverse-thinking model...
tags: [decision, mental-model]
Summary
- Five parallel extractors in Cangjie-Skill Stage 1 each target distinct knowledge types: frameworks, principles, cases, counter-examples, and glossary terms
- All extractors share inputs (
BOOK_OVERVIEW.md, book text, specific prompts) but operate independently - Output files (
candidates/*.md) feed forward to Stage 1.5 verification and downstream skill generation - Parallel design maximizes coverage, speed, and judgment isolation as documented in [
methodology/02-stage1-parallel-extract.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md)
Frequently Asked Questions
What files do the five parallel extractors produce?
Each extractor writes to a dedicated file in candidates/: frameworks.md, principles.md, cases.md, counter-examples.md, and glossary.md. All use consistent YAML structure with fields for id, title, source_chapter, source_quote, summary, and tags.
Why five separate extractors instead of one comprehensive reader?
The design explicitly trades single-agent completeness for multi-agent specialization. Per the source methodology, different "perspectives catch items the others miss"—a framework-focused agent inherently overlooks anti-patterns that the counter-example-extractor targets. Later stages handle deduplication.
How do the extractors avoid duplicate findings across categories?
They don't—Stage 1 accepts intentional overlap. The V1 cross-domain verification step (Stage 1.5) merges and reconciles duplicates. Independence in Stage 1 ensures no extractor suppresses valid candidates due to perceived redundancy.
Can I run a single extractor without the others?
Yes. The architecture supports individual execution via run_extractor.py with --prompt and --output arguments swapped for each agent. This helps debug prompts or re-process specific knowledge types without repeating the full pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →