# Cangjie-Skill Stage 1 Parallel Extractors: What Are the Five Roles and How They Work

> Discover the five parallel extractors in Cangjie-Skill Stage 1: framework, principle, case, counter-example, and glossary. Learn how each independently scans for unique knowledge types to boost coverage.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: internals
- Published: 2026-08-14

---

The five parallel extractors in Cangjie-Skill Stage 1 are **framework-extractor**, **principle-extractor**, **case-extractor**, **counter-example-extractor**, and **glossary-extractor**—each independently scanning the same book for distinct knowledge types to maximize coverage.

The Cangjie-Skill pipeline transforms books into structured skills through a multi-stage process. In Stage 1, the challenge of "reading the book" is distributed across five specialized sub-agents that operate in parallel. According to the [methodology/02-stage1-parallel-extract.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) source file, all five extractors receive identical inputs—**[`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md)**, the raw book text, and their respective extractor prompts—but each applies a unique lens to surface different knowledge artifacts.

## The Five Parallel Extractor Roles

### Framework-Extractor: Thinking Models and Decision Methods

The **framework-extractor** identifies **mental models, decision frameworks, and reasoning methods** the author presents.

This extractor surfaces structured approaches to problem-solving: step-by-step processes, analytical frameworks, and thinking patterns that readers can apply broadly. Its output lands in [`candidates/frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/frameworks.md) as a YAML-structured list of frameworks.

**Key prompt file:** [[`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md)

### Principle-Extractor: Rules, Checklists, and Assertions

The **principle-extractor** captures **principles, checklists, rules, and explicit assertions**.

These are prescriptive statements—"always do X," "never do Y," numbered lists, and heuristic guidelines that the author treats as foundational truths. Output goes to [`candidates/principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/principles.md).

**Key prompt file:** [[`extractors/principle-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md)

### Case-Extractor: Concrete Examples from the Text

The **case-extractor** pulls **concrete examples the author actually uses** in the book.

These are specific scenarios, stories, applications, or worked examples that illustrate broader points. Unlike hypothetical illustrations, these are instances explicitly deployed by the author. Results saved to [`candidates/cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/cases.md).

**Key prompt file:** [[`extractors/case-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md)

### Counter-Example-Extractor: Warnings and Anti-Patterns

The **counter-example-extractor** hunts for **warnings, failures, anti-patterns, and traps** the author highlights.

This extractor surfaces negative knowledge: common mistakes, pitfalls to avoid, failed approaches, and cautionary tales. Often overlooked by other extractors, this "inverse" knowledge is critical for robust skill-building. Output: [`candidates/counter-examples.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/counter-examples.md).

**Key prompt file:** [[`extractors/counter-example-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/counter-example-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/counter-example-extractor.md)

### Glossary-Extractor: Key Concepts and Terminology

The **glossary-extractor** builds a **dictionary of key concepts and specialized terminology**.

This captures definitions, technical terms, recurring concepts, and the author's specific vocabulary—establishing the semantic foundation for later stages. Output: [`candidates/glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/glossary.md).

**Key prompt file:** [[`extractors/glossary-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md)](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md)

## Why Parallel Extraction? Three Design Rationale

The parallel architecture in Stage 1 serves three strategic purposes:

- **Coverage through diversity** — Different specialized lenses catch items others miss; a counter-example invisible to the framework-extractor gets flagged by its dedicated counterpart
- **Speed through concurrency** — Claude-code agents spawn simultaneously, reducing wall-clock runtime versus sequential processing
- **Independence prevents contamination** — Each extractor judges in isolation without seeing others' results; Stage 1.5 (V1 cross-domain verification) later merges overlapping findings intentionally

## Running the Extractors: Example Commands

Each extractor follows an identical invocation pattern, differing only in prompt and output paths:

```bash

# Framework extractor

python run_extractor.py \
  --overview BOOK_OVERVIEW.md \
  --text book.txt \
  --prompt extractors/framework-extractor.md \
  --output candidates/frameworks.md

# Principle extractor (same pattern, different files)

python run_extractor.py \
  --overview BOOK_OVERVIEW.md \
  --text book.txt \
  --prompt extractors/principle-extractor.md \
  --output candidates/principles.md

```

All extractors produce YAML-structured candidates. Example output format:

```yaml
id: f01
title: 逆向思维
type: framework
source_chapter: 第三讲
source_quote: |
  "反过来想,总是反过来想..."
summary: |
  The author proposes a reverse-thinking model...
tags: [decision, mental-model]

```

## Summary

- **Five parallel extractors** in Cangjie-Skill Stage 1 each target distinct knowledge types: frameworks, principles, cases, counter-examples, and glossary terms
- All extractors share inputs ([`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md), book text, specific prompts) but operate independently
- Output files (`candidates/*.md`) feed forward to Stage 1.5 verification and downstream skill generation
- Parallel design maximizes coverage, speed, and judgment isolation as documented in [[`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md)](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md)

## Frequently Asked Questions

### What files do the five parallel extractors produce?

Each extractor writes to a dedicated file in `candidates/`: [`frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/frameworks.md), [`principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/principles.md), [`cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/cases.md), [`counter-examples.md`](https://github.com/kangarooking/cangjie-skill/blob/main/counter-examples.md), and [`glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/glossary.md). All use consistent YAML structure with fields for `id`, `title`, `source_chapter`, `source_quote`, `summary`, and `tags`.

### Why five separate extractors instead of one comprehensive reader?

The design explicitly trades single-agent completeness for multi-agent specialization. Per the source methodology, different "perspectives catch items the others miss"—a framework-focused agent inherently overlooks anti-patterns that the counter-example-extractor targets. Later stages handle deduplication.

### How do the extractors avoid duplicate findings across categories?

They don't—Stage 1 accepts intentional overlap. The V1 cross-domain verification step (Stage 1.5) merges and reconciles duplicates. Independence in Stage 1 ensures no extractor suppresses valid candidates due to perceived redundancy.

### Can I run a single extractor without the others?

Yes. The architecture supports individual execution via [`run_extractor.py`](https://github.com/kangarooking/cangjie-skill/blob/main/run_extractor.py) with `--prompt` and `--output` arguments swapped for each agent. This helps debug prompts or re-process specific knowledge types without repeating the full pipeline.