# The 5 Parallel Extractors in cangjie-skill: Architecture, Functions, and YAML Examples

> Explore cangjie-skill's 5 parallel extractors framework principle case counter-example glossary. Understand their functions and get YAML examples for structured knowledge extraction.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: architecture
- Published: 2026-08-13

---

**The cangjie-skill pipeline uses five parallel extractors—framework, principle, case, counter-example, and glossary—to independently scan source texts and extract distinct knowledge types into structured YAML outputs.**

The **cangjie-skill** repository implements a multi-stage document processing system that transforms raw texts (books, transcripts, lectures) into actionable **skills**. The first stage—**parallel extraction**—deploys five specialized sub-agents simultaneously. Each extractor has a narrowly defined scope, preventing cognitive overlap while maximizing coverage of the source material. This design is documented in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) and implemented through dedicated prompt files in the `extractors/` directory.

---

## What Are Parallel Extractors in cangjie-skill?

**Parallel extractors** are **independent LLM agents** that run concurrently during *Stage 1* of the pipeline. Each agent receives the same inputs—a [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) and the current book chunk—then processes it through its own specialized lens.

The architecture prioritizes **isolation over integration** at the extraction phase. By forcing each extractor to ignore content outside its domain (explicitly defined in "不属于你的" / "not yours" sections of each prompt), the system prevents single-view bias and ensures cleaner downstream merging in Stage 1.5 (three-way verification).

---

## The 5 Parallel Extractors: Functions and Outputs

### 1. Framework Extractor

**Function:** Identifies **thinking models, decision frameworks, and reasoning methods**—systematic approaches authors use to solve problems.

**What it captures:** Mental models, analytical frameworks, structured thinking patterns, and repeatable decision procedures.

**Output destination:** [`candidates/frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/frameworks.md)

**Source file:** [`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md)

This extractor distinguishes frameworks from principles by focusing on *process* rather than *prescription*. A framework provides a structured way to think through a problem; a principle tells you what to value or avoid.

```yaml
- id: f01
  title: 逆向思维
  type: framework
  source_chapter: 第 3 讲
  source_quote: |
    "反过来想,总是反过来想。如果知道我会在哪里死去,那我就永远不去那里。"
  summary: |
    面对一个目标时, 不直接问"怎么达成", 而先问"什么会让我失败"。
    列出失败因素后, 避免它们, 反向推出应做的事。
  tags: [decision, mental-model, inversion]

```

---

### 2. Principle Extractor

**Function:** Extracts **principles, checklists, rules, and maxims**—compressed wisdom that guides action.

**What it captures:** Explicit rules, negative checklists ("stop doing" lists), heuristic guidelines, and author-stated "always/never" statements.

**Output destination:** [`candidates/principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/principles.md)

**Source file:** [`extractors/principle-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md)

The principle extractor targets *actionable constraints*. It specifically flags content where the author warns against common mistakes or prescribes invariant behaviors across situations.

```yaml
- id: p01
  title: Stop Doing List
  type: principle
  source_chapter: 第 2 部分 · 投资篇
  source_quote: |
    "不做什么比做什么更重要。我们的 stop doing list 比 to do list 长得多。"
  summary: |
    主动列出"绝对不做"的清单, 比列"要做"的清单更能防止重大错误。
    适用于投资、战略、职业选择等"错一次就伤筋动骨"的场景。
  tags: [principle, decision, negative-checklist]

```

---

### 3. Case Extractor

**Function:** Pulls **concrete cases where the author personally applied a method**—first-hand implementation examples.

**What it captures:** Specific decisions the author made, with context, actions taken, and outcomes observed. Excludes generic illustrations; requires author involvement.

**Output destination:** [`candidates/cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/cases.md)

**Source file:** [`extractors/case-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md)

The case extractor enforces **personal provenance**. It rejects hypothetical scenarios or third-party anecdotes unless the author explicitly analyzes them. Valid cases include explicit `bound_to` fields linking the example to frameworks or principles it demonstrates.

```yaml
- id: c01
  title: 投资 See's Candy
  type: case
  source_chapter: 第 5 讲
  source_quote: |
    "我们以 2500 万美元收购了 See's Candy...这是我们第一次为品牌溢价付费。"
  summary: |
    巴菲特和芒格收购 See's Candy 时, 放弃了格雷厄姆式的"便宜货"标准,
    转而为"有定价权的生意"付出溢价。这笔投资后来成了他们转向
    "优质企业+合理价格"策略的转折点。
  bound_to:
    - "能力圈 + 定价权"
    - "从便宜货到优质企业的转变"
  outcome: |
    该公司后续 30 年产生的现金流远超初始投资, 验证了新策略。
  tags: [case, investment, turning-point]

```

---

### 4. Counter-Example Extractor

**Function:** Captures **author-warned failure modes, counter-examples, and traps**—negative knowledge that prevents errors.

**What it captures:** Documented failures, cognitive biases with consequences, common pitfalls the author explicitly flags, and mechanisms by which good intentions go wrong.

**Output destination:** [`candidates/counter-examples.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/counter-examples.md)

**Source file:** [`extractors/counter-example-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/counter-example-extractor.md)

This extractor builds **defensive knowledge**. Unlike the principle extractor's "what not to do," the counter-example extractor explains *why* things fail and *how* to recognize early warning signs. It includes structured fields for `failure_mode`, `mechanism`, and `warning_signs`.

```yaml
- id: ce01
  title: 过度自信偏误
  type: counter-example
  source_chapter: 误判心理学 · 第 12 条
  source_quote: |
    "大多数人都认为自己比平均水平更聪明、更公正、更有能力。
     这种自我评价偏误在投资中尤其致命。"
  failure_mode: |
    在自己不懂的领域自认为懂, 导致做出超出能力圈的决策。
  mechanism: |
    人脑默认把"熟悉"等同于"理解", 把"喜欢"等同于"正确"。
    没有外部校正机制时, 过度自信会随成功次数累积而强化。
  warning_signs:
    - 决策时感到"这很简单"
    - 没有 plan B
    - 不愿意向人请教
  bound_to:
    - "能力圈判断"
    - "检查清单决策"
  tags: [counter-example, cognitive-bias, overconfidence]

```

---

### 5. Glossary Extractor

**Function:** Builds a **shared glossary of key concepts**—author-defined terms that acquire specialized meaning in context.

**What it captures:** Terms the author defines, distinguishes from common usage, or uses with technical precision. Excludes standard dictionary definitions.

**Output destination:** [`candidates/glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/glossary.md)

**Source file:** [`extractors/glossary-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md)

The glossary extractor prevents **semantic drift**. It captures `author_definition`, `key_distinction` (what the term is *not*), and `why_it_matters` (consequences of misunderstanding). This ensures downstream skills use terms as the author intended, not as generic concepts.

```yaml
- id: g01
  term: 能力圈
  type: term
  source_chapter: 第 2 讲
  author_definition: |
    "你真正能做出准确判断的知识边界。不是你知道什么, 而是你知道'你知道什么'和'你不知道什么'的边界。"
  key_distinction: |
    ≠ "熟悉的领域" — 熟悉不代表能做判断
    ≠ "专业领域" — 博士学位也可能在能力圈外
    = 能持续做出比市场更准判断的范围 (需经实战验证)
  why_it_matters: |
    "能力圈"一词在所有投资决策类 skill 中都会出现。
    若沿用字典义, skill 会建议用户"评估一下是否熟悉该领域", 这是错的。
    正确的用法是"评估自己过去在此领域的判断准确率"。
  tags: [term, core-concept]

```

---

## How the 5 Extractors Operate Together

### Parallel Execution Model

All five extractors receive **identical inputs** and run **simultaneously** (with serial fallback). This design, documented in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md), provides three advantages:

- **Coverage breadth** – No single extractor dominates the analysis
- **Conflict surfacing** – Contradictory extractions become visible for Stage 1.5 resolution
- **Speed** – Wall-clock time remains constant regardless of extractor count

### Uniform Output Schema

Despite their specialized domains, all extractors emit **structurally compatible YAML**. Every candidate contains:

| Field | Purpose |
|-------|---------|
| `id` | Unique identifier (prefix encodes extractor type: `f-`, `p-`, `c-`, `ce-`, `g-`) |
| `type` | Discriminator for downstream routing |
| `source_chapter` | Provenance for verification |
| `source_quote` | Verbatim evidence |
| `summary` | Condensed explanation |
| `tags` | Cross-reference keywords |

### Isolation Boundaries

Each prompt file explicitly defines exclusion zones. The framework extractor is instructed to ignore principles; the case extractor skips counter-examples. This **negative definition** prevents duplicate extraction and forces clean categorical boundaries.

---

## Summary

- **The 5 parallel extractors** in cangjie-skill are: **framework**, **principle**, **case**, **counter-example**, and **glossary** extractors—each with a dedicated prompt file in `extractors/`.

- **Framework extractor** captures thinking models and reasoning methods; **principle extractor** captures rules and checklists; **case extractor** captures author-personal implementations; **counter-example extractor** captures failure modes and traps; **glossary extractor** captures author-defined terms.

- **Parallel execution** with **isolated scopes** maximizes coverage while minimizing overlap, producing five candidate files that feed into three-way verification (Stage 1.5) and eventual skill synthesis.

- All extractors share a **uniform YAML schema** with `id`, `type`, `source_chapter`, `source_quote`, `summary`, and `tags` fields.

---

## Frequently Asked Questions

### How do the extractors avoid extracting the same content?

Each extractor prompt contains explicit **exclusion instructions** ("不属于你的" / "not yours" sections) that define what content belongs to the other four extractors. If content falls outside its defined scope, the extractor must ignore it. This negative framing, combined with distinct `type` labels in output, prevents cross-contamination.

### Can I add a sixth parallel extractor to the pipeline?

The current architecture in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) supports extensibility, but requires: (1) a new prompt file in `extractors/` with proper scope isolation, (2) a corresponding output file in `candidates/`, and (3) updates to Stage 1.5 verification logic to handle the new candidate type. The shared YAML schema should be respected for downstream compatibility.

### Why does the case extractor require author personal involvement?

The case extractor specifically rejects generic illustrations and third-party anecdotes to ensure **epistemic grounding**. Author-personal cases carry implicit validity claims—"I did this and observed that"—which differ from "someone could do this." This distinction matters for skill construction because personally validated cases provide stronger evidence for method effectiveness.

### What happens if two extractors disagree on categorizing the same passage?

Disagreements surface during **Stage 1.5 three-way verification**, where a separate arbitration agent reviews candidate files from all five extractors. The arbitrator evaluates source quotes, resolves duplicates, and assigns final categorizations based on the methodology's hierarchical disambiguation rules. The parallel extraction design intentionally preserves ambiguity rather than forcing premature resolution.