What Are the Differences Between the Five Parallel Extractors in Cangjie-Skill?
Cangjie-skill runs five extractor modules in parallel, each specializing in a distinct semantic class: principles, glossary terms, frameworks, cases, and counter-examples.
The cangjie-skill pipeline processes source books through concurrent extraction modules that share the same inputs—BOOK_OVERVIEW.md and raw book text—but produce structurally different outputs. Understanding these five parallel extractors reveals how the system transforms unstructured text into typed knowledge nodes for downstream skill generation.
How the Five Extractors Differ
Each extractor is defined in its own specification file under the extractors/ directory. They differ in recognition signals, output schemas, and downstream consumption patterns.
Principle Extractor
Primary goal: Capture actionable rules, checklists, and maxims that tell the reader what to do or not do.
- Recognition signals: Phrases like "必须…", "不要…", "要记住…", numbered or bullet lists, repeated assertions
- Key output fields:
type: principle,title,source_chapter,source_quote,summary,tags - Source: [
extractors/principle-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md)
# Principle Extractor output
- id: p01
title: Stop Doing List
type: principle
source_chapter: 第 2 部分 · 投资篇
source_quote: |
"不做什么比做什么更重要。我们的 stop doing list 比 to do list 长得多。"
summary: |
主动列出"绝对不做"的清单,比列"要做"的清单更能防止重大错误。
tags: [principle, decision, negative-checklist]
Glossary Extractor
Primary goal: Build a shared terminology dictionary for the entire skill pipeline.
- Recognition signals: Words appearing ≥3 times, explicitly defined terms ("所谓 X 是指…"), core-concept words unique to the book
- Key output fields:
term,author_definition,key_distinction,why_it_matters,tags - Source: [
extractors/glossary-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md)
# Glossary Extractor output
- id: g01
term: 能力圈
type: term
source_chapter: 第 2 讲
author_definition: |
"你真正能做出准确判断的知识边界。不是你知道什么, 而是你知道'你知道什么'和'你不知道什么'的边界。"
key_distinction: |
≠ "熟悉的领域" — 熟悉不代表能做判断
≠ "专业领域" — 博士学位也可能在能力圈外
why_it_matter: |
"能力圈"一词在所有投资决策类 skill 中都会出现……
tags: [term, core-concept]
Framework Extractor
Primary goal: Identify mental models or reasoning structures the author uses to think through problems.
- Recognition signals: Introductory sentences ("我们先…", "思考的框架是…"), repeated schematic language, diagrammatic cues
- Key output fields:
type: framework,title,components,example_usage,tags - Source: [
extractors/framework-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md)
# Framework Extractor output
- id: f01
title: 5‑Step Decision Framework
type: framework
source_chapter: 第 3 部分 · 决策篇
components:
- 确认问题
- 收集信息
- 构建模型
- 模拟结果
- 评估风险
example_usage: |
"在评估新项目时,先确认问题…"
tags: [framework, decision, process]
Case Extractor
Primary goal: Collect concrete examples from the author's personal experience.
- Recognition signals: First-person narration ("我…", "我们…"), story-like headings, explicit "案例" markers
- Key output fields:
type: case,title,source_chapter,quote,lesson,tags - Source: [
extractors/case-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md)
# Case Extractor output
- id: c01
title: "买下特斯拉"案例
type: case
source_chapter: 第 4 部分 · 案例篇
quote: |
"我在 2012 年买入特斯拉,当时市值只有 6 亿美元…"
lesson: |
早期进入高成长公司能获得指数级回报,但需做好长期持有准备。
tags: [case, investment, Tesla]
Counter-Example Extractor
Primary goal: Gather negative or failed patterns that warn readers about pitfalls.
- Recognition signals: Expressions like "不要…因为…会失败", "常见的错误是…", "如果…会出问题"
- Key output fields:
type: counter_example,title,source_chapter,description,preventive_tip,tags - Source: [
extractors/counter-example-extractor.md](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/counter-example-extractor.md)
# Counter‑Example Extractor output
- id: ce01
title: "盲目追高"失败模式
type: counter_example
source_chapter: 第 5 部分 · 风险篇
description: |
"很多投资者在股价快速上涨后追进,结果往往在回调时亏损。"
preventive_tip: |
只在回调后、已有明确基本面支撑时才考虑加仓。
tags: [counter_example, risk, overbuy]
Parallel Architecture Benefits
The five extractors operate concurrently as documented in methodology/02-stage1-parallel-extract.md. This design delivers three advantages:
- Separation of concerns — Each module handles one semantic class, preventing "over-extraction" (e.g., confusing a principle with a framework)
- Simplified downstream consumption — Skill generators subscribe only to node types they need
- Unified knowledge graph — All outputs merge into a single graph where the
typefield determines routing
Field Comparison Across Extractors
| Extractor | Core Identifier | Mandatory Narrative Field | Distinctive Field |
|---|---|---|---|
| Principle | source_quote |
summary |
Negative indicators (stop-doing signals) |
| Glossary | term |
author_definition |
key_distinction, why_it_matters |
| Framework | title |
components |
example_usage |
| Case | quote |
lesson |
First-person source attribution |
| Counter-Example | description |
preventive_tip |
Failure pattern documentation |
Summary
- Principles extract actionable dos and don'ts with emphasis on stop-doing lists
- Glossary builds shared vocabulary with author-defined boundaries
- Frameworks capture reusable mental models with component breakdowns
- Cases preserve author experiences with extracted lessons
- Counter-examples document failure modes with preventive guidance
All five run in parallel from shared inputs, producing typed nodes that feed into cangjie-skill's unified knowledge graph.
Frequently Asked Questions
Why does cangjie-skill use five separate extractors instead of one general extractor?
Five specialized extractors prevent semantic confusion and simplify downstream processing. A single extractor would need complex disambiguation logic to distinguish, for example, a numbered checklist (principle) from a step-by-step process (framework). By separating concerns at extraction time, each downstream skill generator receives pre-typed nodes matching its expected schema.
How does the Glossary Extractor decide which terms to include?
The Glossary Extractor selects terms based on frequency (≥3 occurrences), explicit definitions ("所谓 X 是指…"), or unique core-concept status within the book. This ensures the terminology dictionary covers concepts the author treats as foundational, not merely common words.
What happens if two extractors identify the same text passage?
The parallel extraction architecture assumes overlapping identification may occur. During the merge phase into the knowledge graph, deduplication occurs based on content hashing and source location. The type field determines which downstream modules receive the node, so a passage tagged as both principle and case would typically be resolved based on primary signal strength or manually curated rules.
Can I add custom extractors to the cangjie-skill pipeline?
The modular extractor design in extractors/ supports extension. A new extractor requires: a specification markdown file defining recognition signals and output schema, registration in the parallel stage configuration, and a downstream consumer that understands the new node type. The existing five extractors serve as templates for schema design and signal definition patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →