# What Is the Framework Extractor in cangjie-skill?

> Discover the framework extractor in cangjie-skill. This Stage 1 pipeline component identifies transferable thinking structures like mental models and decision frameworks in book text, outputting them as structured YAML records.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: internals
- Published: 2026-07-18

---

**The framework-extractor is a Stage 1 pipeline component in kangarooking/cangjie-skill that scans book text and outlines to identify transferable thinking structures—such as mental models and decision frameworks—and outputs them as structured YAML records to `books/<slug>/candidates/frameworks.md`.**

The framework-extractor operates as one of five parallel extractors in the cangjie-skill knowledge extraction pipeline. Located within the `extractors/` directory of the kangarooking/cangjie-skill repository, this specialized component isolates reusable cognitive patterns from concrete examples and terminology. According to the specification in [`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md), it processes both the [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) outline and raw chapter text during Stage 1 to populate the "framework" layer of the project's knowledge graph.

## Core Responsibility and Scope

The framework-extractor's sole responsibility is identifying **transferable thinking structures** that authors employ to approach complex problems. The extractor scans for three distinct categories as defined in its specification:

- **Thought models** – Abstract cognitive patterns such as "能力圈" (circle of competence), "逆向思维" (inversion thinking), or "多元思维模型" (multiple mental models)
- **Decision frameworks** – Systematic evaluation approaches like "先问最坏情况再算期望值" (assessing worst-case scenarios before calculating expected value)
- **Reasoning methods** – Analytical techniques including "从第一性原理出发" (reasoning from first principles)

These elements are deliberately distinct from concrete cases, glossary definitions, or counter-examples. The pipeline routes those knowledge types to complementary components: the `case-extractor`, `glossary-extractor`, and `counter-example-extractor`.

## Pipeline Architecture and Integration

The framework-extractor executes within a multi-stage orchestration defined in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) and the methodology documentation in [`methodology/01-stage0-adler.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/01-stage0-adler.md):

| Pipeline Stage | Component | Role |
|----------------|-----------|------|
| **Stage 0** | [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) generation | Creates the high-level book skeleton and chapter overview |
| **Stage 1** | **framework-extractor** (parallel execution) | Extracts structured knowledge from raw text |
| **Stage 1.5** | Triple-verification & deduplication | Consolidates overlapping results from all five extractors |
| **Later Stages** | Knowledge-base rendering | Converts curated candidates into published markdown articles |

During Stage 1, the framework-extractor runs in parallel with the `principle-extractor`, `case-extractor`, `counter-example-extractor`, and `glossary-extractor`. Each component writes to its own candidate file, leaving data consolidation to the Stage 1.5 deduplication step.

## Input Processing and Output Specification

### Expected Input Format

The extractor consumes two primary inputs documented in [`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md):

1. **[`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md)**: The structured outline containing chapter summaries and thematic hierarchies
2. **Book text**: Either the complete manuscript or chunked segments for distributed processing

Example input structure:

```markdown

# BOOK_OVERVIEW.md (excerpt)

## 第 3 讲 逆向思维

- 章节概述 ……

# Book text (excerpt)

> 反过来想, 总是反过来想。如果知道我会在哪里死去, 那我就永远不去那里。
> 这比正向推理更有效，因为人对“不想要什么” 的判断通常比对“想要什么” 更清晰。

```

### YAML Output Schema

When the extractor identifies a candidate framework, it appends a YAML record to `books/<slug>/candidates/frameworks.md` following this exact schema:

```yaml
- id: f01
  title: 逆向思维
  type: framework
  source_chapter: 第 3 讲
  source_quote: |
    "反过来想,总是反过来想。如果知道我会在哪里死去,那我就永远不去那里。"
  summary: |
    面对一个目标时, 不直接问"怎么达成", 而先问"什么会让我失败"。
    列出失败因素后, 避免它们, 反向推出应做的事。
    这比正向推理更有效, 因为人对"不想要什么"的判断通常比对"想要什么"更清晰。
  tags: [decision, mental-model, inversion]

```

Each record requires a unique identifier, precise source attribution via `source_chapter`, and descriptive `tags` for downstream classification. The one-record-per-line format facilitates parsing by the Stage 1.5 consolidation logic.

## Implementation Pattern and Code Structure

While the actual orchestration is handled by the pipeline runner, the framework-extractor follows this operational contract illustrated in Python-style pseudo-code:

```python
from framework_extractor import FrameworkExtractor

# Load inputs from repository structure

overview = load_markdown('BOOK_OVERVIEW.md')
text_chunks = split_book_into_chunks('book.txt')

# Execute extraction logic (parallel with other extractors)

frameworks = FrameworkExtractor().run(overview, text_chunks)

# Persist to candidate directory

write_yaml('books/my-book/candidates/frameworks.md', frameworks)

```

The implementation leverages the overview file to contextualize chapter boundaries while scanning for framework-specific signals: systematic thinking patterns, meta-cognitive strategies, and reusable decision heuristics.

## Distinction from Parallel Extractors

The framework-extractor occupies a specific semantic niche within the five-extractor architecture:

- **framework-extractor**: Captures abstract thinking structures and mental models
- **principle-extractor**: Handles concrete rules, laws, and operational checklists  
- **case-extractor**: Extracts specific author-used examples and narrative instances
- **glossary-extractor**: Processes term definitions and specialized vocabulary
- **counter-example-extractor**: Identifies cautionary tales and anti-patterns

This strict separation ensures that downstream consumers—such as ChatGPT prompts configured in the skill interface—can surface reusable mental models without conflating them with specific instances or terminology definitions.

## Summary

- The **framework-extractor** is one of five parallel extractors operating during Stage 1 of the kangarooking/cangjie-skill pipeline
- It identifies **transferable thinking structures** including mental models, decision frameworks, and reasoning methods while excluding concrete cases and glossary terms
- Output is written to `books/<slug>/candidates/frameworks.md` as standardized YAML records containing ID, title, source attribution, and tags
- It processes both the [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) outline and raw book text to contextualize extracted frameworks
- Duplicate removal occurs during Stage 1.5, not within the extractor itself
- It functions alongside specialized extractors for principles, cases, glossary terms, and counter-examples

## Frequently Asked Questions

### What distinguishes the framework-extractor from other extractors in cangjie-skill?

The framework-extractor specifically targets abstract, transferable thinking patterns such as mental models and decision heuristics, while other extractors handle concrete knowledge types. The `case-extractor` captures specific stories and examples, the `glossary-extractor` processes terminology definitions, the `principle-extractor` handles rules and checklists, and the `counter-example-extractor` identifies failure patterns. This specialization allows the pipeline to build a layered knowledge graph where frameworks can be referenced independently of specific instances.

### Where does the framework-extractor store its output?

The extractor writes candidate frameworks to `books/<slug>/candidates/frameworks.md` within the repository structure. Each extracted framework is appended as a YAML record containing fields for `id`, `title`, `type`, `source_chapter`, `source_quote`, `summary`, and `tags`. This location serves as the raw candidate pool before Stage 1.5 deduplication and final knowledge-base publishing.

### How does the framework-extractor handle duplicate entries?

The extractor does not perform deduplication internally. According to the pipeline architecture documented in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md), duplicate removal occurs during Stage 1.5 ("triple-verification & deduplication") after all five parallel extractors have completed their work. The framework-extractor focuses solely on candidate identification and initial generation, leaving consolidation and conflict resolution to the subsequent pipeline stage.

### What types of thinking structures does the framework-extractor identify?

As specified in [`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md), the extractor identifies three primary categories: **thought models** (cognitive patterns like "逆向思维" or inversion thinking), **decision frameworks** (systematic evaluation methods), and **reasoning methods** (analytical approaches such as first-principles thinking). It deliberately excludes concrete examples, glossary terms, and counter-examples, which are routed to their respective specialized extractors.