# Roles of the Five Parallel Extractors in Stage 1 of the Cangjie-Skill Pipeline

> Discover the roles of five parallel extractors in the Cangjie-Skill pipeline. Learn how they maximize knowledge coverage by extracting frameworks, principles, cases, counter-examples, and glossary terms.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: internals
- Published: 2026-07-20

---

**The five parallel extractors in Stage 1 of the Cangjie-Skill pipeline are independent sub-agents that simultaneously analyze a book to extract distinct knowledge types—frameworks, principles, cases, counter-examples, and glossary terms—maximizing coverage while preventing cross-contamination.**

The Cangjie-Skill project (`kangarooking/cangjie-skill`) implements a multi-stage pipeline for converting books into executable skills. In Stage 1, the system employs five parallel extractors to perform the initial "reading" of the source material, ensuring comprehensive knowledge capture before downstream verification and linking stages process the extracted data.

## Overview of the Parallel Extraction Architecture

According to the methodology documentation in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md), Stage 1 splits the work of reading a book among five independent sub-agents. Each extractor receives identical inputs: the global [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md), the raw book text (or its path), and a specialized prompt file defining its extraction scope.

Despite sharing the same inputs, each agent operates in isolation, focusing exclusively on its designated knowledge domain. This design prevents cross-contamination while allowing Claude-code agents to spawn simultaneously, significantly reducing overall runtime.

## The Five Extractor Roles and Responsibilities

### 1. Framework Extractor

The **framework-extractor** identifies thinking models, decision frameworks, and reasoning methods embedded in the text. It outputs structured findings to [`candidates/frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/frameworks.md). This agent specifically targets mental models and systematic approaches the author uses for problem-solving, as defined in [`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md).

### 2. Principle Extractor

The **principle-extractor** captures principles, checklists, rules, and assertions. Its output feeds into [`candidates/principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/principles.md), creating a repository of actionable guidelines and heuristics extracted from the source material. The prompt logic resides in [`extractors/principle-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md).

### 3. Case Extractor

The **case-extractor** focuses on concrete examples the author actually uses within the book. Unlike generic illustration, this agent extracts specific instances and scenarios the author references, storing them in [`candidates/cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/cases.md) for later skill application. Configuration is handled via [`extractors/case-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md).

### 4. Counter-Example Extractor

The **counter-example-extractor** identifies warnings, failures, anti-patterns, and traps the author mentions. By cataloging these negative examples in [`candidates/counter-examples.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/counter-examples.md), the pipeline ensures that generated skills include boundary conditions and failure modes. The extraction rules are defined in [`extractors/counter-example-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/counter-example-extractor.md).

### 5. Glossary Extractor

The **glossary-extractor** builds a terminology dictionary by extracting key concepts and definitions. It populates [`candidates/glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/glossary.md) with domain-specific vocabulary, ensuring consistent terminology usage throughout the downstream skill generation process, guided by [`extractors/glossary-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md).

## Why Parallel Extraction Matters

The parallel architecture serves three critical functions in the Cangjie-Skill pipeline:

- **Coverage**: Different perspectives catch items others miss. For example, a counter-example that the framework extractor would skip gets captured by the counter-example extractor, ensuring no knowledge units slip through.

- **Speed**: Claude-code agents execute simultaneously rather than sequentially. This parallelization shortens the overall runtime of Stage 1 processing.

- **Independence**: Each extractor judges in isolation without influence from other agents. Later stages—specifically Stage 1.5 (V1 cross-domain verification)—handle the merging of overlapping results, maintaining clean separation of concerns during initial extraction.

## How to Run the Extractors

To execute a single extractor manually, use the helper script with the specific prompt and output configuration:

```bash
python run_extractor.py \
  --overview BOOK_OVERVIEW.md \
  --text book.txt \
  --prompt extractors/framework-extractor.md \
  --output candidates/frameworks.md

```

Swap the `--prompt` and `--output` arguments to run the other four extractors ([`principle-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/principle-extractor.md), [`case-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/case-extractor.md), [`counter-example-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/counter-example-extractor.md), [`glossary-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/glossary-extractor.md)).

Each extractor generates YAML-structured output blocks. A candidate entry must contain metadata fields including `id`, `title`, `type`, `source_chapter`, `source_quote`, `summary`, and `tags`:

```yaml
id: f01
title: 逆向思维
type: framework
source_chapter: 第三讲
source_quote: |
  "反过来想,总是反过来想..."
summary: |
  The author proposes a reverse-thinking model...
tags: [decision, mental-model]

```

## Summary

- The **five parallel extractors** in Stage 1 function as independent sub-agents that simultaneously process source books.
- Each extractor specializes in a distinct knowledge type: frameworks, principles, cases, counter-examples, or glossary terms.
- Parallel execution maximizes **coverage**, improves **speed** through simultaneous processing, and maintains **independence** to prevent cross-contamination.
- Extractors output to specific files in the `candidates/` directory, feeding downstream verification and skill-generation stages.
- The architecture is defined in [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) and implemented through specialized prompt files in the `extractors/` directory.

## Frequently Asked Questions

### What inputs do the five parallel extractors receive?

All five extractors receive the same three inputs: the global [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) file providing context about the book, the raw book text or its file path, and their specific extractor prompt file that defines their specialized extraction scope.

### Why does Cangjie-Skill use five separate extractors instead of one comprehensive agent?

Using five specialized agents maximizes coverage by ensuring different perspectives capture distinct knowledge types that a single agent might miss. It also enables parallel processing for speed and maintains isolation between extraction domains to prevent cross-contamination before the verification stage merges results.

### Where are the extractor outputs stored in the Cangjie-Skill repository?

Each extractor writes to a specific file in the `candidates/` directory: [`frameworks.md`](https://github.com/kangarooking/cangjie-skill/blob/main/frameworks.md), [`principles.md`](https://github.com/kangarooking/cangjie-skill/blob/main/principles.md), [`cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/cases.md), [`counter-examples.md`](https://github.com/kangarooking/cangjie-skill/blob/main/counter-examples.md), and [`glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/glossary.md). These files serve as inputs for Stage 1.5 (cross-domain verification) and subsequent pipeline stages.

### Can I run individual extractors independently of the full pipeline?

Yes. The repository provides a helper script (illustrated as [`run_extractor.py`](https://github.com/kangarooking/cangjie-skill/blob/main/run_extractor.py)) that allows you to execute single extractors by specifying the overview file, text source, prompt file, and output destination via command-line arguments.