Roles of the Five Parallel Extractors in Stage 1 of the Cangjie-Skill Pipeline

The five parallel extractors in Stage 1 of the Cangjie-Skill pipeline are independent sub-agents that simultaneously analyze a book to extract distinct knowledge types—frameworks, principles, cases, counter-examples, and glossary terms—maximizing coverage while preventing cross-contamination.

The Cangjie-Skill project (kangarooking/cangjie-skill) implements a multi-stage pipeline for converting books into executable skills. In Stage 1, the system employs five parallel extractors to perform the initial "reading" of the source material, ensuring comprehensive knowledge capture before downstream verification and linking stages process the extracted data.

Overview of the Parallel Extraction Architecture

According to the methodology documentation in methodology/02-stage1-parallel-extract.md, Stage 1 splits the work of reading a book among five independent sub-agents. Each extractor receives identical inputs: the global BOOK_OVERVIEW.md, the raw book text (or its path), and a specialized prompt file defining its extraction scope.

Despite sharing the same inputs, each agent operates in isolation, focusing exclusively on its designated knowledge domain. This design prevents cross-contamination while allowing Claude-code agents to spawn simultaneously, significantly reducing overall runtime.

The Five Extractor Roles and Responsibilities

1. Framework Extractor

The framework-extractor identifies thinking models, decision frameworks, and reasoning methods embedded in the text. It outputs structured findings to candidates/frameworks.md. This agent specifically targets mental models and systematic approaches the author uses for problem-solving, as defined in extractors/framework-extractor.md.

2. Principle Extractor

The principle-extractor captures principles, checklists, rules, and assertions. Its output feeds into candidates/principles.md, creating a repository of actionable guidelines and heuristics extracted from the source material. The prompt logic resides in extractors/principle-extractor.md.

3. Case Extractor

The case-extractor focuses on concrete examples the author actually uses within the book. Unlike generic illustration, this agent extracts specific instances and scenarios the author references, storing them in candidates/cases.md for later skill application. Configuration is handled via extractors/case-extractor.md.

4. Counter-Example Extractor

The counter-example-extractor identifies warnings, failures, anti-patterns, and traps the author mentions. By cataloging these negative examples in candidates/counter-examples.md, the pipeline ensures that generated skills include boundary conditions and failure modes. The extraction rules are defined in extractors/counter-example-extractor.md.

5. Glossary Extractor

The glossary-extractor builds a terminology dictionary by extracting key concepts and definitions. It populates candidates/glossary.md with domain-specific vocabulary, ensuring consistent terminology usage throughout the downstream skill generation process, guided by extractors/glossary-extractor.md.

Why Parallel Extraction Matters

The parallel architecture serves three critical functions in the Cangjie-Skill pipeline:

  • Coverage: Different perspectives catch items others miss. For example, a counter-example that the framework extractor would skip gets captured by the counter-example extractor, ensuring no knowledge units slip through.

  • Speed: Claude-code agents execute simultaneously rather than sequentially. This parallelization shortens the overall runtime of Stage 1 processing.

  • Independence: Each extractor judges in isolation without influence from other agents. Later stages—specifically Stage 1.5 (V1 cross-domain verification)—handle the merging of overlapping results, maintaining clean separation of concerns during initial extraction.

How to Run the Extractors

To execute a single extractor manually, use the helper script with the specific prompt and output configuration:

python run_extractor.py \
  --overview BOOK_OVERVIEW.md \
  --text book.txt \
  --prompt extractors/framework-extractor.md \
  --output candidates/frameworks.md

Swap the --prompt and --output arguments to run the other four extractors (principle-extractor.md, case-extractor.md, counter-example-extractor.md, glossary-extractor.md).

Each extractor generates YAML-structured output blocks. A candidate entry must contain metadata fields including id, title, type, source_chapter, source_quote, summary, and tags:

id: f01
title: 逆向思维
type: framework
source_chapter: 第三讲
source_quote: |
  "反过来想,总是反过来想..."
summary: |
  The author proposes a reverse-thinking model...
tags: [decision, mental-model]

Summary

  • The five parallel extractors in Stage 1 function as independent sub-agents that simultaneously process source books.
  • Each extractor specializes in a distinct knowledge type: frameworks, principles, cases, counter-examples, or glossary terms.
  • Parallel execution maximizes coverage, improves speed through simultaneous processing, and maintains independence to prevent cross-contamination.
  • Extractors output to specific files in the candidates/ directory, feeding downstream verification and skill-generation stages.
  • The architecture is defined in methodology/02-stage1-parallel-extract.md and implemented through specialized prompt files in the extractors/ directory.

Frequently Asked Questions

What inputs do the five parallel extractors receive?

All five extractors receive the same three inputs: the global BOOK_OVERVIEW.md file providing context about the book, the raw book text or its file path, and their specific extractor prompt file that defines their specialized extraction scope.

Why does Cangjie-Skill use five separate extractors instead of one comprehensive agent?

Using five specialized agents maximizes coverage by ensuring different perspectives capture distinct knowledge types that a single agent might miss. It also enables parallel processing for speed and maintains isolation between extraction domains to prevent cross-contamination before the verification stage merges results.

Where are the extractor outputs stored in the Cangjie-Skill repository?

Each extractor writes to a specific file in the candidates/ directory: frameworks.md, principles.md, cases.md, counter-examples.md, and glossary.md. These files serve as inputs for Stage 1.5 (cross-domain verification) and subsequent pipeline stages.

Can I run individual extractors independently of the full pipeline?

Yes. The repository provides a helper script (illustrated as run_extractor.py) that allows you to execute single extractors by specifying the overview file, text source, prompt file, and output destination via command-line arguments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →