What Are the Differences Between the Five Parallel Extractors in Cangjie-Skill?

Cangjie-skill runs five extractor modules in parallel, each specializing in a distinct semantic class: principles, glossary terms, frameworks, cases, and counter-examples.

The cangjie-skill pipeline processes source books through concurrent extraction modules that share the same inputs—BOOK_OVERVIEW.md and raw book text—but produce structurally different outputs. Understanding these five parallel extractors reveals how the system transforms unstructured text into typed knowledge nodes for downstream skill generation.

How the Five Extractors Differ

Each extractor is defined in its own specification file under the extractors/ directory. They differ in recognition signals, output schemas, and downstream consumption patterns.

Principle Extractor

Primary goal: Capture actionable rules, checklists, and maxims that tell the reader what to do or not do.


# Principle Extractor output

- id: p01
  title: Stop Doing List
  type: principle
  source_chapter: 第 2 部分 · 投资篇
  source_quote: |
    "不做什么比做什么更重要。我们的 stop doing list 比 to do list 长得多。"
  summary: |
    主动列出"绝对不做"的清单,比列"要做"的清单更能防止重大错误。
  tags: [principle, decision, negative-checklist]

Glossary Extractor

Primary goal: Build a shared terminology dictionary for the entire skill pipeline.


# Glossary Extractor output

- id: g01
  term: 能力圈
  type: term
  source_chapter: 第 2 讲
  author_definition: |
    "你真正能做出准确判断的知识边界。不是你知道什么, 而是你知道'你知道什么'和'你不知道什么'的边界。"
  key_distinction: |
    ≠ "熟悉的领域" — 熟悉不代表能做判断
    ≠ "专业领域" — 博士学位也可能在能力圈外
  why_it_matter: |
    "能力圈"一词在所有投资决策类 skill 中都会出现……
  tags: [term, core-concept]

Framework Extractor

Primary goal: Identify mental models or reasoning structures the author uses to think through problems.


# Framework Extractor output

- id: f01
  title: 5‑Step Decision Framework
  type: framework
  source_chapter: 第 3 部分 · 决策篇
  components:
    - 确认问题
    - 收集信息
    - 构建模型
    - 模拟结果
    - 评估风险
  example_usage: |
    "在评估新项目时,先确认问题…"
  tags: [framework, decision, process]

Case Extractor

Primary goal: Collect concrete examples from the author's personal experience.


# Case Extractor output

- id: c01
  title: "买下特斯拉"案例
  type: case
  source_chapter: 第 4 部分 · 案例篇
  quote: |
    "我在 2012 年买入特斯拉,当时市值只有 6 亿美元…"
  lesson: |
    早期进入高成长公司能获得指数级回报,但需做好长期持有准备。
  tags: [case, investment, Tesla]

Counter-Example Extractor

Primary goal: Gather negative or failed patterns that warn readers about pitfalls.


# Counter‑Example Extractor output

- id: ce01
  title: "盲目追高"失败模式
  type: counter_example
  source_chapter: 第 5 部分 · 风险篇
  description: |
    "很多投资者在股价快速上涨后追进,结果往往在回调时亏损。"
  preventive_tip: |
    只在回调后、已有明确基本面支撑时才考虑加仓。
  tags: [counter_example, risk, overbuy]

Parallel Architecture Benefits

The five extractors operate concurrently as documented in methodology/02-stage1-parallel-extract.md. This design delivers three advantages:

  1. Separation of concerns — Each module handles one semantic class, preventing "over-extraction" (e.g., confusing a principle with a framework)
  2. Simplified downstream consumption — Skill generators subscribe only to node types they need
  3. Unified knowledge graph — All outputs merge into a single graph where the type field determines routing

Field Comparison Across Extractors

Extractor Core Identifier Mandatory Narrative Field Distinctive Field
Principle source_quote summary Negative indicators (stop-doing signals)
Glossary term author_definition key_distinction, why_it_matters
Framework title components example_usage
Case quote lesson First-person source attribution
Counter-Example description preventive_tip Failure pattern documentation

Summary

  • Principles extract actionable dos and don'ts with emphasis on stop-doing lists
  • Glossary builds shared vocabulary with author-defined boundaries
  • Frameworks capture reusable mental models with component breakdowns
  • Cases preserve author experiences with extracted lessons
  • Counter-examples document failure modes with preventive guidance

All five run in parallel from shared inputs, producing typed nodes that feed into cangjie-skill's unified knowledge graph.

Frequently Asked Questions

Why does cangjie-skill use five separate extractors instead of one general extractor?

Five specialized extractors prevent semantic confusion and simplify downstream processing. A single extractor would need complex disambiguation logic to distinguish, for example, a numbered checklist (principle) from a step-by-step process (framework). By separating concerns at extraction time, each downstream skill generator receives pre-typed nodes matching its expected schema.

How does the Glossary Extractor decide which terms to include?

The Glossary Extractor selects terms based on frequency (≥3 occurrences), explicit definitions ("所谓 X 是指…"), or unique core-concept status within the book. This ensures the terminology dictionary covers concepts the author treats as foundational, not merely common words.

What happens if two extractors identify the same text passage?

The parallel extraction architecture assumes overlapping identification may occur. During the merge phase into the knowledge graph, deduplication occurs based on content hashing and source location. The type field determines which downstream modules receive the node, so a passage tagged as both principle and case would typically be resolved based on primary signal strength or manually curated rules.

Can I add custom extractors to the cangjie-skill pipeline?

The modular extractor design in extractors/ supports extension. A new extractor requires: a specification markdown file defining recognition signals and output schema, registration in the parallel stage configuration, and a downstream consumer that understands the new node type. The existing five extractors serve as templates for schema design and signal definition patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →