How the Five Parallel Extractors in cangjie-skill Function: A Complete Technical Guide
The five parallel extractors in cangjie-skill are autonomous LLM-driven modules that simultaneously process the same book text to isolate distinct knowledge types—principles, glossary terms, frameworks, cases, and counter-examples—producing structured YAML outputs that feed into a unified knowledge pipeline.
The kangarooking/cangjie-skill repository implements a sophisticated extraction pipeline designed to transform unstructured books into structured, actionable knowledge bases. At the heart of this system lies a parallel processing architecture where five specialized extractors operate concurrently, each applying distinct heuristics to mine specific categories of wisdom from the source text.
The Five Parallel Extractors: An Overview
During Stage 1 of the pipeline, cangjie-skill launches five extractor modules simultaneously. Each extractor receives identical inputs—the BOOK_OVERVIEW.md skeleton plus the full book text—but applies specialized recognition patterns to isolate a unique knowledge domain. This parallel design accelerates processing while ensuring that principles, terminology, mental models, historical cases, and anti-patterns are captured with domain-specific precision.
The extractors write their findings to separate candidate files (candidates/principles.md, candidates/glossary.md, etc.) before a subsequent Stage 1.5 – Triple Verify merges, deduplicates, and validates the combined output.
Stage 1: Parallel Extraction Deep Dive
Principle Extractor
File: extractors/principle-extractor.md
The Principle Extractor identifies concrete, actionable rules, checklists, maxims, and heuristics that readers can apply immediately. It scans for imperative language patterns such as "必须…" (must), "不要…" (do not), "要记住…" (remember), numbered lists, and repeated authoritative assertions.
Output Schema:
id: Unique identifiertitle: Principle nametype: principlesource_chapter: Origin locationsource_quote: Verbatim citationsummary: Actionable explanationtags: Categorization labels
# Principle Extractor example (principles.md)
- id: p01
title: Stop Doing List
type: principle
source_chapter: 第 2 部分 · 投资篇
source_quote: |
"不做什么比做什么更重要。我们的 stop doing list 比 to do list 长得多。"
summary: |
主动列出"绝对不做"的清单, 比列"要做"的清单更能防止重大错误。
适用于投资、战略、职业选择等"错一次就伤筋动骨"的场景。
tags: [principle, decision, negative-checklist]
Glossary Extractor
File: extractors/glossary-extractor.md
The Glossary Extractor constructs a core-concept dictionary of terms carrying specific meaning within the author's framework. It flags terms appearing three or more times, phrases followed by explicit definitions ("所谓 X, 是指…"), and words whose technical usage diverges from common parlance.
Output Schema:
id: Unique identifierterm: The concept nametype: termsource_chapter: Origin locationauthor_definition: Author's specific meaningkey_distinction: Differentiation from similar conceptswhy_it_matters: Contextual importancetags: Categorization labels
# Glossary Extractor example (glossary.md)
- id: g01
term: 能力圈
type: term
source_chapter: 第 2 讲
author_definition: |
"你真正能做出准确判断的知识边界。不是你知道什么, 而是你知道'你知道什么'和'你不知道什么'的边界。"
key_distinction: |
≠ "熟悉的领域" — 熟悉不代表能做判断
≠ "专业领域" — 博士学位也可能在能力圈外
= 能持续做出比市场更准判断的范围 (需经实战验证)
why_it_matters: |
"能力圈"一词在所有投资决策类 skill 中都会出现。
若沿用字典义, skill 会建议用户"评估一下是否熟悉该领域", 这是错的。
正确的用法是"评估自己过去在此领域的判断准确率"。
tags: [term, core-concept]
Framework Extractor
File: extractors/framework-extractor.md
The Framework Extractor detects transferable thinking models, decision frameworks, and reasoning methods. It recognizes named mental models, structural patterns like "面对 X 类问题时应该…", author-explicit declarations ("这是我的 mental model"), and logical constructs such as if-then or before-after sequences.
Output Schema:
id: Unique identifiertitle: Framework nametype: frameworksource_chapter: Origin locationsource_quote: Verbatim citationsummary: Explanation of the mental modeltags: Categorization labels
# Framework Extractor example (frameworks.md)
- id: f01
title: 逆向思维
type: framework
source_chapter: 第 3 讲
source_quote: |
"反过来想,总是反过来想。如果知道我会在哪里死去,那我就永远不去那里。"
summary: |
面对一个目标时, 不直接问"怎么达成", 而先问"什么会让我失败"。
列出失败因素后, 避免它们, 反向推出应做的事。
这比正向推理更有效, 因为人对"不想要什么"的判断通常比对"想要什么"更清晰。
tags: [decision, mental-model, inversion]
Case Extractor
File: extractors/case-extractor.md
The Case Extractor isolates concrete anecdotes, historical examples, or personal experiences cited by the author. It triggers on storytelling markers including "我曾经…" (I once…), "案例: …" (case study), explicit dates, organizational names, and first-person narratives.
Output Schema:
id: Unique identifiertitle: Case descriptiontype: casesource_chapter: Origin locationsource_quote: Verbatim citationsummary: Context and significancetags: Categorization labels
# Case Extractor example (cases.md)
- id: c01
title: 乔布斯的产品发布会
type: case
source_chapter: 第 5 讲
source_quote: |
"我记得第一次在大学里看到乔布斯的演讲,他把每一个细节都演练到极致。"
summary: |
该案例展示了“极致准备”对说服听众的效果,是作者在产品营销中反复提到的实践。
tags: [case, presentation, preparation]
Counter-Example Extractor
File: extractors/counter-example-extractor.md
The Counter-Example Extractor captures failure patterns, warnings, and anti-patterns illustrating what not to do. It identifies negative outcomes through phrases like "如果…会导致…" (if…then will result in…), "不要犯…错误" (don't make the mistake of…), and "一个常见的陷阱是…" (a common trap is…).
Output Schema:
id: Unique identifiertitle: Anti-pattern nametype: counter-examplesource_chapter: Origin locationsource_quote: Verbatim citationsummary: Explanation of the failure modetags: Categorization labels
# Counter‑Example Extractor example (counter-examples.md)
- id: ce01
title: 过度迭代的陷阱
type: counter-example
source_chapter: 第 7 讲
source_quote: |
"我曾在一个项目中每周都做一次大规模的功能迭代,结果团队疲惫、需求混乱。"
summary: |
频繁的大幅度迭代会导致技术债务累积、团队士气下降,提醒读者在迭代频率和规模上保持平衡。
tags: [failure, iteration, process]
Pipeline Architecture and Data Flow
The five parallel extractors operate within a strict multi-stage pipeline that ensures data integrity and prevents cross-contamination between knowledge types.
Pre-Stage 0 generates the BOOK_OVERVIEW.md, a high-level table of contents that provides structural context for all downstream extraction.
Stage 1 – Parallel Extraction launches the five extractors concurrently. Each module processes the complete book text independently, applying its specialized heuristics without interference from the others. The deliberately non-overlapping scopes—enforced through explicit "Not in scope" clauses in each extractor specification—ensure that ambiguous content routes to the correct module.
Stage 1.5 – Triple Verify performs cross-stream deduplication and validation. This stage resolves any edge cases where content might theoretically fit multiple extractors, merging the five YAML streams into a coherent, non-redundant knowledge base ready for downstream skill generation.
Summary
- Five autonomous extractors run simultaneously during Stage 1, each mining a distinct knowledge type from the same source text.
- Input standardization: All extractors consume
BOOK_OVERVIEW.mdplus the raw book content, ensuring consistent context. - Structured output: Every extractor produces YAML-compliant records with standardized fields (
id,title,type,source_chapter, etc.), enabling seamless downstream processing. - Zero-overlap design: Explicit scope boundaries and "Not in scope" rules prevent duplicate extraction across modules.
- Triple verification: Stage 1.5 deduplicates and validates the combined output before feeding into framework synthesis and skill generation.
Frequently Asked Questions
How do the five parallel extractors avoid extracting the same content?
Each extractor specification includes explicit "Not in scope" clauses that redirect ambiguous content to the appropriate module. For example, the Principle Extractor redirects undefined terminology to the Glossary Extractor, while the Framework Extractor passes pure historical anecdotes to the Case Extractor. Stage 1.5's triple-verification layer catches any residual overlap through cross-stream deduplication algorithms.
What file format do the extractors generate?
All five extractors produce YAML output with consistent schema conventions. Each record includes an id, title, type (principle, term, framework, case, or counter-example), source_chapter, source_quote, and tags field. This standardized contract allows the pipeline to merge the five candidate files (candidates/principles.md, candidates/glossary.md, etc.) without format conflicts.
Why use parallel extraction instead of a single unified model?
Parallel extraction allows domain-specific optimization of recognition heuristics for each knowledge type. A single model would struggle to simultaneously optimize for glossary term detection, failure pattern recognition, and mental model extraction. The concurrent design also speeds processing by allowing the five LLM-driven modules to execute simultaneously rather than sequentially.
Can I customize the recognition signals for a specific extractor?
Yes. Each extractor is defined in a standalone Markdown specification file within the extractors/ directory (e.g., extractors/principle-extractor.md). These documents contain the recognition signals, output schemas, and scope boundaries that guide the LLM. Modifying these specifications changes the extraction behavior without altering the core pipeline code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →