How Cangjie-Skill Handles Parallel Extraction of Knowledge: A 5-Agent Pipeline Explained
Cangjie-skill uses a five-stage workflow where Stage 1 spawns five specialized extractor agents simultaneously, each targeting a distinct knowledge type (frameworks, principles, cases, counter-examples, and glossary terms), with automatic chunking for large texts and graceful fallback to serial execution when parallel runtime isn't available.
The kangarooking/cangjie-skill repository implements an agentic pipeline designed to extract methodological knowledge from long-form content like books or video transcripts. Parallel extraction is the cornerstone of its first stage, enabling comprehensive coverage without sacrificing processing speed. This article examines the architecture, implementation details, and fallback mechanisms that make this approach reliable.
How Parallel Extraction Works in Stage 1
The extraction pipeline begins with a single coordinated call to the Claude Agent tool. According to the source code in SKILL.md (lines 81-89), the main skill launches five independent Task agents in one operation—not sequentially—maximizing throughput and coverage.
Each agent receives identical global context (BOOK_OVERVIEW.md plus the source text or chunk) paired with its dedicated extractor prompt. They execute concurrently, writing results to separate candidate files without blocking one another.
The Five Specialized Extractor Agents
Each sub-agent corresponds to one extractor definition in the extractors/ folder. The following table maps agents to their responsibilities and output locations:
| Agent | Prompt File | Knowledge Type Extracted | Output Path |
|---|---|---|---|
| Framework extractor | extractors/framework-extractor.md |
Decision frameworks, mental models | books/<slug>/candidates/framework.md |
| Principle extractor | extractors/principle-extractor.md |
Principles, rules | books/<slug>/candidates/principle.md |
| Case extractor | extractors/case-extractor.md |
Concrete author examples | books/<slug>/candidates/case.md |
| Counter-example extractor | extractors/counter-example-extractor.md |
Failure patterns, traps | books/<slug>/candidates/counter-example.md |
| Glossary extractor | extractors/glossary-extractor.md |
Term definitions | books/<slug>/candidates/glossary.md |
As documented in extractors/framework-extractor.md (lines 33-40), each extractor operates independently: it reads the book overview, scans the (potentially chunked) text, and writes its YAML-formatted candidates to its designated file. This separation of concerns ensures no single agent is overwhelmed and each knowledge type receives focused attention.
Handling Large Texts Through Chunking
When source material exceeds a single sub-agent's context window, Cangjie-skill implements boundary-aware chunking. As specified in methodology/02-stage1-parallel-extract.md (lines 24-32), the system splits content on natural divisions:
- Chapter boundaries
- Volume breaks
- Video part transitions
Critical design invariant: every chunk includes the complete BOOK_OVERVIEW.md as a global anchor. This ensures extractors maintain holistic reasoning about the entire book while processing only a limited slice, preserving coherence across chunks.
Graceful Degradation to Serial Execution
The pipeline anticipates runtime constraints. Per methodology/02-stage1-parallel-extract.md (lines 22-23), when the execution environment cannot support parallel agents, Cangjie-skill degrades gracefully by running the same five extractor prompts sequentially. Output formats remain identical, ensuring pipeline compatibility regardless of infrastructure.
Post-Extraction Verification (Stage 1.5)
Parallelism introduces potential duplication. The pipeline addresses this through Stage 1.5 triple verification, described in SKILL.md (lines 98-105). After all sub-agents complete:
- Duplicate or overlapping candidates are identified and merged
- Sourcing confidence is validated
- Only well-attested knowledge units proceed to subsequent stages
This consolidation step guarantees that speed from parallelism never compromises quality.
Implementation Example
The following Python pseudocode illustrates the parallel spawn pattern as implemented in the actual agent tool:
# Define the five extractor configurations
agents = [
{"name": "framework_extractor", "prompt": read("extractors/framework-extractor.md")},
{"name": "principle_extractor", "prompt": read("extractors/principle-extractor.md")},
{"name": "case_extractor", "prompt": read("extractors/case-extractor.md")},
{"name": "counter_extractor", "prompt": read("extractors/counter-example-extractor.md")},
{"name": "glossary_extractor", "prompt": read("extractors/glossary-extractor.md")},
]
# Single call spawns all agents concurrently
results = agent_tool.spawn_many(
agents=agents,
inputs={
"book_overview": read("books/slug/BOOK_OVERVIEW.md"),
"source_text": large_text_or_chunks
}
)
# Each extractor writes to its dedicated candidate file
for r in results:
write(f"books/slug/candidates/{r['type']}.md", r["yaml_output"])
Key Source Files
| File | Role in Parallel Extraction |
|---|---|
methodology/02-stage1-parallel-extract.md |
Documents parallel strategy, chunking logic, and serial fallback |
SKILL.md |
Pipeline overview with parallel spawn step (lines 81-89, 98-105) |
extractors/framework-extractor.md |
Template for extractor prompts; defines output format and candidate file writing |
extractors/principle-extractor.md |
Dedicated prompt for principle/rule extraction |
extractors/case-extractor.md |
Dedicated prompt for concrete example extraction |
extractors/counter-example-extractor.md |
Dedicated prompt for failure pattern extraction |
extractors/glossary-extractor.md |
Dedicated prompt for term definition extraction |
books/<slug>/candidates/ |
Runtime directory storing raw parallel outputs before verification |
Summary
- Concurrency model: Five specialized agents launch simultaneously via single
agent_tool.spawn_many()call - Knowledge coverage: Each agent targets one knowledge type (frameworks, principles, cases, counter-examples, glossary)
- Scalability: Boundary-aware chunking with global context preservation handles arbitrary text lengths
- Resilience: Automatic fallback to serial execution maintains pipeline functionality on constrained runtimes
- Quality assurance: Stage 1.5 verification deduplicates and validates parallel outputs before downstream processing
Frequently Asked Questions
What happens if two extractors find the same knowledge unit?
The Stage 1.5 triple verification process identifies duplicates or overlaps and merges them. According to SKILL.md (lines 98-105), only candidates with confident sourcing survive this consolidation, ensuring the pipeline deduplicates without losing valid extractions.
Can I run Cangjie-skill on hardware that doesn't support parallel agents?
Yes. The pipeline automatically degrades to serial execution, running the five extractor prompts sequentially while preserving identical output formats and file locations. This fallback is documented in methodology/02-stage1-parallel-extract.md (lines 22-23).
How does chunking affect extraction quality?
Each chunk includes the complete BOOK_OVERVIEW.md as a global anchor. Per methodology/02-stage1-parallel-extract.md (lines 24-32), this design lets extractors reason about the whole book's context even when processing limited slices, maintaining coherence across boundaries.
Where are the parallel extraction results stored before verification?
Each extractor writes to its own candidate file in books/<slug>/candidates/, using the pattern <type>.md (e.g., framework.md, principle.md). These files serve as raw inputs to Stage 1.5 verification before qualified candidates advance to later pipeline stages.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →