# What Are the Differences Between the Five Parallel Extractors in Cangjie-Skill?

> Explore the distinct semantic classes of Cangjie-skill's five parallel extractors: principles, glossary terms, frameworks, cases, and counter-examples. Understand their specializations.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: deep-dive
- Published: 2026-08-16

---

**Cangjie-skill runs five extractor modules in parallel, each specializing in a distinct semantic class: principles, glossary terms, frameworks, cases, and counter-examples.**

The cangjie-skill pipeline processes source books through concurrent extraction modules that share the same inputs—[`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) and raw book text—but produce structurally different outputs. Understanding these five parallel extractors reveals how the system transforms unstructured text into typed knowledge nodes for downstream skill generation.

## How the Five Extractors Differ

Each extractor is defined in its own specification file under the `extractors/` directory. They differ in **recognition signals**, **output schemas**, and **downstream consumption patterns**.

### Principle Extractor

**Primary goal:** Capture actionable rules, checklists, and maxims that tell the reader what to do or not do.

- **Recognition signals:** Phrases like "必须…", "不要…", "要记住…", numbered or bullet lists, repeated assertions
- **Key output fields:** `type: principle`, `title`, `source_chapter`, `source_quote`, `summary`, `tags`
- **Source:** [[`extractors/principle-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md)

```yaml

# Principle Extractor output

- id: p01
  title: Stop Doing List
  type: principle
  source_chapter: 第 2 部分 · 投资篇
  source_quote: |
    "不做什么比做什么更重要。我们的 stop doing list 比 to do list 长得多。"
  summary: |
    主动列出"绝对不做"的清单，比列"要做"的清单更能防止重大错误。
  tags: [principle, decision, negative-checklist]

```

### Glossary Extractor

**Primary goal:** Build a shared terminology dictionary for the entire skill pipeline.

- **Recognition signals:** Words appearing ≥3 times, explicitly defined terms ("所谓 X 是指…"), core-concept words unique to the book
- **Key output fields:** `term`, `author_definition`, `key_distinction`, `why_it_matters`, `tags`
- **Source:** [[`extractors/glossary-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md)

```yaml

# Glossary Extractor output

- id: g01
  term: 能力圈
  type: term
  source_chapter: 第 2 讲
  author_definition: |
    "你真正能做出准确判断的知识边界。不是你知道什么, 而是你知道'你知道什么'和'你不知道什么'的边界。"
  key_distinction: |
    ≠ "熟悉的领域" — 熟悉不代表能做判断
    ≠ "专业领域" — 博士学位也可能在能力圈外
  why_it_matter: |
    "能力圈"一词在所有投资决策类 skill 中都会出现……
  tags: [term, core-concept]

```

### Framework Extractor

**Primary goal:** Identify mental models or reasoning structures the author uses to think through problems.

- **Recognition signals:** Introductory sentences ("我们先…", "思考的框架是…"), repeated schematic language, diagrammatic cues
- **Key output fields:** `type: framework`, `title`, `components`, `example_usage`, `tags`
- **Source:** [[`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md)

```yaml

# Framework Extractor output

- id: f01
  title: 5‑Step Decision Framework
  type: framework
  source_chapter: 第 3 部分 · 决策篇
  components:
    - 确认问题
    - 收集信息
    - 构建模型
    - 模拟结果
    - 评估风险
  example_usage: |
    "在评估新项目时，先确认问题…"
  tags: [framework, decision, process]

```

### Case Extractor

**Primary goal:** Collect concrete examples from the author's personal experience.

- **Recognition signals:** First-person narration ("我…", "我们…"), story-like headings, explicit "案例" markers
- **Key output fields:** `type: case`, `title`, `source_chapter`, `quote`, `lesson`, `tags`
- **Source:** [[`extractors/case-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md)

```yaml

# Case Extractor output

- id: c01
  title: "买下特斯拉"案例
  type: case
  source_chapter: 第 4 部分 · 案例篇
  quote: |
    "我在 2012 年买入特斯拉，当时市值只有 6 亿美元…"
  lesson: |
    早期进入高成长公司能获得指数级回报，但需做好长期持有准备。
  tags: [case, investment, Tesla]

```

### Counter-Example Extractor

**Primary goal:** Gather negative or failed patterns that warn readers about pitfalls.

- **Recognition signals:** Expressions like "不要…因为…会失败", "常见的错误是…", "如果…会出问题"
- **Key output fields:** `type: counter_example`, `title`, `source_chapter`, `description`, `preventive_tip`, `tags`
- **Source:** [[`extractors/counter-example-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/counter-example-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/counter-example-extractor.md)

```yaml

# Counter‑Example Extractor output

- id: ce01
  title: "盲目追高"失败模式
  type: counter_example
  source_chapter: 第 5 部分 · 风险篇
  description: |
    "很多投资者在股价快速上涨后追进，结果往往在回调时亏损。"
  preventive_tip: |
    只在回调后、已有明确基本面支撑时才考虑加仓。
  tags: [counter_example, risk, overbuy]

```

## Parallel Architecture Benefits

The five extractors operate concurrently as documented in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md). This design delivers three advantages:

1. **Separation of concerns** — Each module handles one semantic class, preventing "over-extraction" (e.g., confusing a principle with a framework)
2. **Simplified downstream consumption** — Skill generators subscribe only to node types they need
3. **Unified knowledge graph** — All outputs merge into a single graph where the `type` field determines routing

## Field Comparison Across Extractors

| Extractor | Core Identifier | Mandatory Narrative Field | Distinctive Field |
|-----------|---------------|---------------------------|-------------------|
| Principle | `source_quote` | `summary` | Negative indicators (stop-doing signals) |
| Glossary | `term` | `author_definition` | `key_distinction`, `why_it_matters` |
| Framework | `title` | `components` | `example_usage` |
| Case | `quote` | `lesson` | First-person source attribution |
| Counter-Example | `description` | `preventive_tip` | Failure pattern documentation |

## Summary

- **Principles** extract actionable dos and don'ts with emphasis on stop-doing lists
- **Glossary** builds shared vocabulary with author-defined boundaries
- **Frameworks** capture reusable mental models with component breakdowns
- **Cases** preserve author experiences with extracted lessons
- **Counter-examples** document failure modes with preventive guidance

All five run in parallel from shared inputs, producing typed nodes that feed into cangjie-skill's unified knowledge graph.

## Frequently Asked Questions

### Why does cangjie-skill use five separate extractors instead of one general extractor?

Five specialized extractors prevent semantic confusion and simplify downstream processing. A single extractor would need complex disambiguation logic to distinguish, for example, a numbered checklist (principle) from a step-by-step process (framework). By separating concerns at extraction time, each downstream skill generator receives pre-typed nodes matching its expected schema.

### How does the Glossary Extractor decide which terms to include?

The Glossary Extractor selects terms based on frequency (≥3 occurrences), explicit definitions ("所谓 X 是指…"), or unique core-concept status within the book. This ensures the terminology dictionary covers concepts the author treats as foundational, not merely common words.

### What happens if two extractors identify the same text passage?

The parallel extraction architecture assumes overlapping identification may occur. During the merge phase into the knowledge graph, deduplication occurs based on content hashing and source location. The `type` field determines which downstream modules receive the node, so a passage tagged as both principle and case would typically be resolved based on primary signal strength or manually curated rules.

### Can I add custom extractors to the cangjie-skill pipeline?

The modular extractor design in `extractors/` supports extension. A new extractor requires: a specification markdown file defining recognition signals and output schema, registration in the parallel stage configuration, and a downstream consumer that understands the new node type. The existing five extractors serve as templates for schema design and signal definition patterns.