The 5 Parallel Extractors in cangjie-skill: Architecture, Functions, and YAML Examples

The cangjie-skill pipeline uses five parallel extractors—framework, principle, case, counter-example, and glossary—to independently scan source texts and extract distinct knowledge types into structured YAML outputs.

The cangjie-skill repository implements a multi-stage document processing system that transforms raw texts (books, transcripts, lectures) into actionable skills. The first stage—parallel extraction—deploys five specialized sub-agents simultaneously. Each extractor has a narrowly defined scope, preventing cognitive overlap while maximizing coverage of the source material. This design is documented in methodology/02-stage1-parallel-extract.md and implemented through dedicated prompt files in the extractors/ directory.


What Are Parallel Extractors in cangjie-skill?

Parallel extractors are independent LLM agents that run concurrently during Stage 1 of the pipeline. Each agent receives the same inputs—a BOOK_OVERVIEW.md and the current book chunk—then processes it through its own specialized lens.

The architecture prioritizes isolation over integration at the extraction phase. By forcing each extractor to ignore content outside its domain (explicitly defined in "不属于你的" / "not yours" sections of each prompt), the system prevents single-view bias and ensures cleaner downstream merging in Stage 1.5 (three-way verification).


The 5 Parallel Extractors: Functions and Outputs

1. Framework Extractor

Function: Identifies thinking models, decision frameworks, and reasoning methods—systematic approaches authors use to solve problems.

What it captures: Mental models, analytical frameworks, structured thinking patterns, and repeatable decision procedures.

Output destination: candidates/frameworks.md

Source file: extractors/framework-extractor.md

This extractor distinguishes frameworks from principles by focusing on process rather than prescription. A framework provides a structured way to think through a problem; a principle tells you what to value or avoid.

- id: f01
  title: 逆向思维
  type: framework
  source_chapter: 第 3 讲
  source_quote: |
    "反过来想,总是反过来想。如果知道我会在哪里死去,那我就永远不去那里。"
  summary: |
    面对一个目标时, 不直接问"怎么达成", 而先问"什么会让我失败"。
    列出失败因素后, 避免它们, 反向推出应做的事。
  tags: [decision, mental-model, inversion]

2. Principle Extractor

Function: Extracts principles, checklists, rules, and maxims—compressed wisdom that guides action.

What it captures: Explicit rules, negative checklists ("stop doing" lists), heuristic guidelines, and author-stated "always/never" statements.

Output destination: candidates/principles.md

Source file: extractors/principle-extractor.md

The principle extractor targets actionable constraints. It specifically flags content where the author warns against common mistakes or prescribes invariant behaviors across situations.

- id: p01
  title: Stop Doing List
  type: principle
  source_chapter: 第 2 部分 · 投资篇
  source_quote: |
    "不做什么比做什么更重要。我们的 stop doing list 比 to do list 长得多。"
  summary: |
    主动列出"绝对不做"的清单, 比列"要做"的清单更能防止重大错误。
    适用于投资、战略、职业选择等"错一次就伤筋动骨"的场景。
  tags: [principle, decision, negative-checklist]

3. Case Extractor

Function: Pulls concrete cases where the author personally applied a method—first-hand implementation examples.

What it captures: Specific decisions the author made, with context, actions taken, and outcomes observed. Excludes generic illustrations; requires author involvement.

Output destination: candidates/cases.md

Source file: extractors/case-extractor.md

The case extractor enforces personal provenance. It rejects hypothetical scenarios or third-party anecdotes unless the author explicitly analyzes them. Valid cases include explicit bound_to fields linking the example to frameworks or principles it demonstrates.

- id: c01
  title: 投资 See's Candy
  type: case
  source_chapter: 第 5 讲
  source_quote: |
    "我们以 2500 万美元收购了 See's Candy...这是我们第一次为品牌溢价付费。"
  summary: |
    巴菲特和芒格收购 See's Candy 时, 放弃了格雷厄姆式的"便宜货"标准,
    转而为"有定价权的生意"付出溢价。这笔投资后来成了他们转向
    "优质企业+合理价格"策略的转折点。
  bound_to:
    - "能力圈 + 定价权"
    - "从便宜货到优质企业的转变"
  outcome: |
    该公司后续 30 年产生的现金流远超初始投资, 验证了新策略。
  tags: [case, investment, turning-point]

4. Counter-Example Extractor

Function: Captures author-warned failure modes, counter-examples, and traps—negative knowledge that prevents errors.

What it captures: Documented failures, cognitive biases with consequences, common pitfalls the author explicitly flags, and mechanisms by which good intentions go wrong.

Output destination: candidates/counter-examples.md

Source file: extractors/counter-example-extractor.md

This extractor builds defensive knowledge. Unlike the principle extractor's "what not to do," the counter-example extractor explains why things fail and how to recognize early warning signs. It includes structured fields for failure_mode, mechanism, and warning_signs.

- id: ce01
  title: 过度自信偏误
  type: counter-example
  source_chapter: 误判心理学 · 第 12 条
  source_quote: |
    "大多数人都认为自己比平均水平更聪明、更公正、更有能力。
     这种自我评价偏误在投资中尤其致命。"
  failure_mode: |
    在自己不懂的领域自认为懂, 导致做出超出能力圈的决策。
  mechanism: |
    人脑默认把"熟悉"等同于"理解", 把"喜欢"等同于"正确"。
    没有外部校正机制时, 过度自信会随成功次数累积而强化。
  warning_signs:
    - 决策时感到"这很简单"
    - 没有 plan B
    - 不愿意向人请教
  bound_to:
    - "能力圈判断"
    - "检查清单决策"
  tags: [counter-example, cognitive-bias, overconfidence]

5. Glossary Extractor

Function: Builds a shared glossary of key concepts—author-defined terms that acquire specialized meaning in context.

What it captures: Terms the author defines, distinguishes from common usage, or uses with technical precision. Excludes standard dictionary definitions.

Output destination: candidates/glossary.md

Source file: extractors/glossary-extractor.md

The glossary extractor prevents semantic drift. It captures author_definition, key_distinction (what the term is not), and why_it_matters (consequences of misunderstanding). This ensures downstream skills use terms as the author intended, not as generic concepts.

- id: g01
  term: 能力圈
  type: term
  source_chapter: 第 2 讲
  author_definition: |
    "你真正能做出准确判断的知识边界。不是你知道什么, 而是你知道'你知道什么'和'你不知道什么'的边界。"
  key_distinction: |
    ≠ "熟悉的领域" — 熟悉不代表能做判断
    ≠ "专业领域" — 博士学位也可能在能力圈外
    = 能持续做出比市场更准判断的范围 (需经实战验证)
  why_it_matters: |
    "能力圈"一词在所有投资决策类 skill 中都会出现。
    若沿用字典义, skill 会建议用户"评估一下是否熟悉该领域", 这是错的。
    正确的用法是"评估自己过去在此领域的判断准确率"。
  tags: [term, core-concept]

How the 5 Extractors Operate Together

Parallel Execution Model

All five extractors receive identical inputs and run simultaneously (with serial fallback). This design, documented in methodology/02-stage1-parallel-extract.md, provides three advantages:

  • Coverage breadth – No single extractor dominates the analysis
  • Conflict surfacing – Contradictory extractions become visible for Stage 1.5 resolution
  • Speed – Wall-clock time remains constant regardless of extractor count

Uniform Output Schema

Despite their specialized domains, all extractors emit structurally compatible YAML. Every candidate contains:

Field Purpose
id Unique identifier (prefix encodes extractor type: f-, p-, c-, ce-, g-)
type Discriminator for downstream routing
source_chapter Provenance for verification
source_quote Verbatim evidence
summary Condensed explanation
tags Cross-reference keywords

Isolation Boundaries

Each prompt file explicitly defines exclusion zones. The framework extractor is instructed to ignore principles; the case extractor skips counter-examples. This negative definition prevents duplicate extraction and forces clean categorical boundaries.


Summary

  • The 5 parallel extractors in cangjie-skill are: framework, principle, case, counter-example, and glossary extractors—each with a dedicated prompt file in extractors/.

  • Framework extractor captures thinking models and reasoning methods; principle extractor captures rules and checklists; case extractor captures author-personal implementations; counter-example extractor captures failure modes and traps; glossary extractor captures author-defined terms.

  • Parallel execution with isolated scopes maximizes coverage while minimizing overlap, producing five candidate files that feed into three-way verification (Stage 1.5) and eventual skill synthesis.

  • All extractors share a uniform YAML schema with id, type, source_chapter, source_quote, summary, and tags fields.


Frequently Asked Questions

How do the extractors avoid extracting the same content?

Each extractor prompt contains explicit exclusion instructions ("不属于你的" / "not yours" sections) that define what content belongs to the other four extractors. If content falls outside its defined scope, the extractor must ignore it. This negative framing, combined with distinct type labels in output, prevents cross-contamination.

Can I add a sixth parallel extractor to the pipeline?

The current architecture in methodology/02-stage1-parallel-extract.md supports extensibility, but requires: (1) a new prompt file in extractors/ with proper scope isolation, (2) a corresponding output file in candidates/, and (3) updates to Stage 1.5 verification logic to handle the new candidate type. The shared YAML schema should be respected for downstream compatibility.

Why does the case extractor require author personal involvement?

The case extractor specifically rejects generic illustrations and third-party anecdotes to ensure epistemic grounding. Author-personal cases carry implicit validity claims—"I did this and observed that"—which differ from "someone could do this." This distinction matters for skill construction because personally validated cases provide stronger evidence for method effectiveness.

What happens if two extractors disagree on categorizing the same passage?

Disagreements surface during Stage 1.5 three-way verification, where a separate arbitration agent reviews candidate files from all five extractors. The arbitrator evaluates source quotes, resolves duplicates, and assigns final categorizations based on the methodology's hierarchical disambiguation rules. The parallel extraction design intentionally preserves ambiguity rather than forcing premature resolution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →