What Is the Framework Extractor in cangjie-skill?
The framework-extractor is a Stage 1 pipeline component in kangarooking/cangjie-skill that scans book text and outlines to identify transferable thinking structures—such as mental models and decision frameworks—and outputs them as structured YAML records to books/<slug>/candidates/frameworks.md.
The framework-extractor operates as one of five parallel extractors in the cangjie-skill knowledge extraction pipeline. Located within the extractors/ directory of the kangarooking/cangjie-skill repository, this specialized component isolates reusable cognitive patterns from concrete examples and terminology. According to the specification in extractors/framework-extractor.md, it processes both the BOOK_OVERVIEW.md outline and raw chapter text during Stage 1 to populate the "framework" layer of the project's knowledge graph.
Core Responsibility and Scope
The framework-extractor's sole responsibility is identifying transferable thinking structures that authors employ to approach complex problems. The extractor scans for three distinct categories as defined in its specification:
- Thought models – Abstract cognitive patterns such as "能力圈" (circle of competence), "逆向思维" (inversion thinking), or "多元思维模型" (multiple mental models)
- Decision frameworks – Systematic evaluation approaches like "先问最坏情况再算期望值" (assessing worst-case scenarios before calculating expected value)
- Reasoning methods – Analytical techniques including "从第一性原理出发" (reasoning from first principles)
These elements are deliberately distinct from concrete cases, glossary definitions, or counter-examples. The pipeline routes those knowledge types to complementary components: the case-extractor, glossary-extractor, and counter-example-extractor.
Pipeline Architecture and Integration
The framework-extractor executes within a multi-stage orchestration defined in SKILL.md and the methodology documentation in methodology/01-stage0-adler.md:
| Pipeline Stage | Component | Role |
|---|---|---|
| Stage 0 | BOOK_OVERVIEW.md generation |
Creates the high-level book skeleton and chapter overview |
| Stage 1 | framework-extractor (parallel execution) | Extracts structured knowledge from raw text |
| Stage 1.5 | Triple-verification & deduplication | Consolidates overlapping results from all five extractors |
| Later Stages | Knowledge-base rendering | Converts curated candidates into published markdown articles |
During Stage 1, the framework-extractor runs in parallel with the principle-extractor, case-extractor, counter-example-extractor, and glossary-extractor. Each component writes to its own candidate file, leaving data consolidation to the Stage 1.5 deduplication step.
Input Processing and Output Specification
Expected Input Format
The extractor consumes two primary inputs documented in extractors/framework-extractor.md:
BOOK_OVERVIEW.md: The structured outline containing chapter summaries and thematic hierarchies- Book text: Either the complete manuscript or chunked segments for distributed processing
Example input structure:
# BOOK_OVERVIEW.md (excerpt)
## 第 3 讲 逆向思维
- 章节概述 ……
# Book text (excerpt)
> 反过来想, 总是反过来想。如果知道我会在哪里死去, 那我就永远不去那里。
> 这比正向推理更有效,因为人对“不想要什么” 的判断通常比对“想要什么” 更清晰。
YAML Output Schema
When the extractor identifies a candidate framework, it appends a YAML record to books/<slug>/candidates/frameworks.md following this exact schema:
- id: f01
title: 逆向思维
type: framework
source_chapter: 第 3 讲
source_quote: |
"反过来想,总是反过来想。如果知道我会在哪里死去,那我就永远不去那里。"
summary: |
面对一个目标时, 不直接问"怎么达成", 而先问"什么会让我失败"。
列出失败因素后, 避免它们, 反向推出应做的事。
这比正向推理更有效, 因为人对"不想要什么"的判断通常比对"想要什么"更清晰。
tags: [decision, mental-model, inversion]
Each record requires a unique identifier, precise source attribution via source_chapter, and descriptive tags for downstream classification. The one-record-per-line format facilitates parsing by the Stage 1.5 consolidation logic.
Implementation Pattern and Code Structure
While the actual orchestration is handled by the pipeline runner, the framework-extractor follows this operational contract illustrated in Python-style pseudo-code:
from framework_extractor import FrameworkExtractor
# Load inputs from repository structure
overview = load_markdown('BOOK_OVERVIEW.md')
text_chunks = split_book_into_chunks('book.txt')
# Execute extraction logic (parallel with other extractors)
frameworks = FrameworkExtractor().run(overview, text_chunks)
# Persist to candidate directory
write_yaml('books/my-book/candidates/frameworks.md', frameworks)
The implementation leverages the overview file to contextualize chapter boundaries while scanning for framework-specific signals: systematic thinking patterns, meta-cognitive strategies, and reusable decision heuristics.
Distinction from Parallel Extractors
The framework-extractor occupies a specific semantic niche within the five-extractor architecture:
- framework-extractor: Captures abstract thinking structures and mental models
- principle-extractor: Handles concrete rules, laws, and operational checklists
- case-extractor: Extracts specific author-used examples and narrative instances
- glossary-extractor: Processes term definitions and specialized vocabulary
- counter-example-extractor: Identifies cautionary tales and anti-patterns
This strict separation ensures that downstream consumers—such as ChatGPT prompts configured in the skill interface—can surface reusable mental models without conflating them with specific instances or terminology definitions.
Summary
- The framework-extractor is one of five parallel extractors operating during Stage 1 of the kangarooking/cangjie-skill pipeline
- It identifies transferable thinking structures including mental models, decision frameworks, and reasoning methods while excluding concrete cases and glossary terms
- Output is written to
books/<slug>/candidates/frameworks.mdas standardized YAML records containing ID, title, source attribution, and tags - It processes both the
BOOK_OVERVIEW.mdoutline and raw book text to contextualize extracted frameworks - Duplicate removal occurs during Stage 1.5, not within the extractor itself
- It functions alongside specialized extractors for principles, cases, glossary terms, and counter-examples
Frequently Asked Questions
What distinguishes the framework-extractor from other extractors in cangjie-skill?
The framework-extractor specifically targets abstract, transferable thinking patterns such as mental models and decision heuristics, while other extractors handle concrete knowledge types. The case-extractor captures specific stories and examples, the glossary-extractor processes terminology definitions, the principle-extractor handles rules and checklists, and the counter-example-extractor identifies failure patterns. This specialization allows the pipeline to build a layered knowledge graph where frameworks can be referenced independently of specific instances.
Where does the framework-extractor store its output?
The extractor writes candidate frameworks to books/<slug>/candidates/frameworks.md within the repository structure. Each extracted framework is appended as a YAML record containing fields for id, title, type, source_chapter, source_quote, summary, and tags. This location serves as the raw candidate pool before Stage 1.5 deduplication and final knowledge-base publishing.
How does the framework-extractor handle duplicate entries?
The extractor does not perform deduplication internally. According to the pipeline architecture documented in SKILL.md, duplicate removal occurs during Stage 1.5 ("triple-verification & deduplication") after all five parallel extractors have completed their work. The framework-extractor focuses solely on candidate identification and initial generation, leaving consolidation and conflict resolution to the subsequent pipeline stage.
What types of thinking structures does the framework-extractor identify?
As specified in extractors/framework-extractor.md, the extractor identifies three primary categories: thought models (cognitive patterns like "逆向思维" or inversion thinking), decision frameworks (systematic evaluation methods), and reasoning methods (analytical approaches such as first-principles thinking). It deliberately excludes concrete examples, glossary terms, and counter-examples, which are routed to their respective specialized extractors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →