# How Case Examples Are Extracted by cangjie-skill: The Case Extractor Pipeline Explained

> Discover how cangjie-skill extracts case examples using its Case Extractor pipeline to identify author-anchored narratives and output structured YAML.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: deep-dive
- Published: 2026-07-18

---

**cangjie-skill extracts case examples using a dedicated Case Extractor that scans source texts for author-anchored narratives during Stage 1 parallel extraction, outputting structured YAML bound to methodological topics.**

The `kangarooking/cangjie-skill` repository implements a sophisticated pipeline for distilling methodological knowledge from books. Understanding how case examples are extracted by cangjie-skill requires examining the **Case Extractor**, one of five parallel extractors that operate simultaneously during the initial processing stage. This component specifically targets instances where authors apply concepts personally or cite historical examples to illustrate methods.

## The Case Extractor Architecture

The Case Extractor functions as a specialized module within **Stage 1 – Parallel Extraction** of the cangjie-skill pipeline. According to [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md), this stage runs five extractors concurrently to parse different knowledge types from source material. The Case Extractor specifically focuses on locating concrete, author-anchored examples rather than abstract principles or theoretical frameworks.

The extractor operates as a deterministic pipeline that transforms raw book text into structured knowledge candidates. It runs alongside other specialized extractors (principles, frameworks, counter-examples, and analogies) to build a comprehensive knowledge base from a single source text.

## Input Sources and Data Requirements

Before extraction begins, the Case Extractor receives two mandatory inputs:

1. **[`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md)** – A structured outline containing chapter metadata and topical organization
2. **Raw book text** – The complete source content in plain text format

These inputs allow the extractor to contextualize findings within the book's structure while scanning for narrative patterns. The extractor references the overview file to accurately populate the `source_chapter` field in output records.

## Signal Detection and Pattern Matching

The core extraction logic relies on **linguistic cue detection** using regular-expression-style rules. The scanner identifies past-tense narratives that indicate concrete examples through specific Chinese and English patterns:

- **Personal experience markers**: `"1973 年，我曾…"` (In 1973, I once...)
- **Narrative openings**: `"有一次…"` (One time...)
- **Explicit case references**: `"某某公司的案例…"` (The case of X company...)
- **Attribution patterns**: `"巴菲特告诉我…"` (Buffett told me...)
- **Illustrative transitions**: `"比如…"` (For example...)

The detection algorithm looks for sequences combining temporal markers with reflective commentary or decision-making narratives. This distinguishes case examples from general exposition or hypothetical scenarios.

## Eligibility Validation and Scope Rules

Not all detected patterns qualify for extraction. The extractor applies strict **scope rules** to filter candidates:

- **Author-anchored requirement**: The event must represent a decision or experience directly undertaken by the author, or a historical case explicitly used to explain a specific method
- **Methodological binding**: Each case must link to at least one methodological topic via the mandatory `bound_to` field
- **Exclusion filters**: Pure background narration, fictional fables, and standalone principles without applied context are rejected

This validation ensures that only actionable, concept-illustrating examples proceed to the knowledge base, excluding anecdotal content that lacks methodological relevance.

## Output Schema and YAML Structure

Qualified cases are serialized into structured YAML blocks with specific mandatory fields. The schema defined in [`extractors/case-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md) requires:

```yaml
- id: <unique-id>
  title: <short description>
  type: case
  source_chapter: <chapter-name>
  source_quote: |
    "<original excerpt from the book>"
  summary: |
    <author’s reflective interpretation of the case>
  bound_to:                # at least one method-topic required

    - "<topic-1>"
    - "<topic-2>"
  outcome: |
    <result or impact described in the book, if any>
  tags: [case, <optional-tags>]

```

The **`bound_to`** field is particularly critical—it creates the semantic bridge between concrete examples and abstract methodology, enabling downstream consumption by **Stage 2 – A1 (Past Application)** processors.

## Integration with the Pipeline

Extracted cases persist to [`candidates/cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/cases.md), a generated Markdown file that aggregates all case YAML blocks from the extraction run. This file serves as the intermediate storage before Stage 2 processing combines cases with principles, frameworks, and counter-examples to construct the final skill knowledge base.

The central registry in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) maps the Case Extractor to its specification document ([`extractors/case-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md)) and tracks its output destination, maintaining the pipeline's provenance chain.

## Implementation Example

Developers can invoke the extractor programmatically as shown in the repository's reference implementation:

```python
from extractors.case_extractor import CaseExtractor

extractor = CaseExtractor(
    overview_path="BOOK_OVERVIEW.md",
    book_text_path="book.txt"
)

cases = extractor.run()
for case in cases:
    print(case["id"], case["title"], case["bound_to"])

```

Typical output follows this structure:

```yaml
- id: c01
  title: 投资 See's Candy
  type: case
  source_chapter: 第 5 讲
  source_quote: |
    "我们以 2500 万美元收购了 See's Candy...这是我们第一次为品牌溢价付费。"
  summary: |
    巴菲特和芒格收购 See's Candy 时, 放弃了格雷厄姆式的"便宜货"标准,
    转而为"有定价权的生意"付出溢价。这笔投资后来成了他们转向
    "优质企业+合理价格"策略的转折点。
  bound_to:
    - "能力圈 + 定价权"
    - "从便宜货到优质企业的转变"
  outcome: |
    该公司后续 30 年产生的现金流远超初始投资, 验证了新策略。
  tags: [case, investment, turning-point]

```

## Summary

- **cangjie-skill** uses a dedicated Case Extractor during Stage 1 parallel extraction to identify author-anchored examples from source texts.
- The extractor requires [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) and the raw book text as inputs to contextualize findings within chapter structures.
- Signal detection relies on regex patterns identifying past-tense narratives with reflective commentary, filtering for direct author experience or explicit method illustration.
- Mandatory **scope rules** require every case to include a `bound_to` field linking it to at least one methodological topic, excluding pure narration or fables.
- Valid cases output to [`candidates/cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/cases.md) as structured YAML, feeding into Stage 2 A1 (Past Application) processing for final knowledge base assembly.

## Frequently Asked Questions

### What input files does the Case Extractor require to function?

The Case Extractor requires two specific inputs: the [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) file containing the structured book outline, and the raw book text itself. These files enable the extractor to map discovered examples to their correct chapter contexts while scanning for linguistic patterns.

### How does cangjie-skill distinguish valid case examples from background narration?

The extractor applies strict eligibility checks based on **scope rules**. Valid cases must represent decisions or experiences directly made by the author, or historical cases explicitly used to explain a method. Pure background stories, fictional fables, and standalone principles without methodological application are automatically excluded.

### Why is the `bound_to` field mandatory in case extraction output?

The `bound_to` field creates the critical semantic link between concrete examples and abstract methodology. This mandatory linkage ensures that Stage 2 processors—specifically the A1 (Past Application) module—can consume case data effectively, connecting real-world applications to their underlying conceptual frameworks.

### Where are extracted case examples stored in the cangjie-skill pipeline?

Extracted cases are written to [`candidates/cases.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/cases.md) during Stage 1 processing. This generated file serves as intermediate storage containing all YAML-structured case blocks, which later aggregate with principles, frameworks, and counter-examples to construct the final skill knowledge base.