How Case Examples Are Extracted by cangjie-skill: The Case Extractor Pipeline Explained

cangjie-skill extracts case examples using a dedicated Case Extractor that scans source texts for author-anchored narratives during Stage 1 parallel extraction, outputting structured YAML bound to methodological topics.

The kangarooking/cangjie-skill repository implements a sophisticated pipeline for distilling methodological knowledge from books. Understanding how case examples are extracted by cangjie-skill requires examining the Case Extractor, one of five parallel extractors that operate simultaneously during the initial processing stage. This component specifically targets instances where authors apply concepts personally or cite historical examples to illustrate methods.

The Case Extractor Architecture

The Case Extractor functions as a specialized module within Stage 1 – Parallel Extraction of the cangjie-skill pipeline. According to methodology/02-stage1-parallel-extract.md, this stage runs five extractors concurrently to parse different knowledge types from source material. The Case Extractor specifically focuses on locating concrete, author-anchored examples rather than abstract principles or theoretical frameworks.

The extractor operates as a deterministic pipeline that transforms raw book text into structured knowledge candidates. It runs alongside other specialized extractors (principles, frameworks, counter-examples, and analogies) to build a comprehensive knowledge base from a single source text.

Input Sources and Data Requirements

Before extraction begins, the Case Extractor receives two mandatory inputs:

  1. BOOK_OVERVIEW.md – A structured outline containing chapter metadata and topical organization
  2. Raw book text – The complete source content in plain text format

These inputs allow the extractor to contextualize findings within the book's structure while scanning for narrative patterns. The extractor references the overview file to accurately populate the source_chapter field in output records.

Signal Detection and Pattern Matching

The core extraction logic relies on linguistic cue detection using regular-expression-style rules. The scanner identifies past-tense narratives that indicate concrete examples through specific Chinese and English patterns:

  • Personal experience markers: "1973 年,我曾…" (In 1973, I once...)
  • Narrative openings: "有一次…" (One time...)
  • Explicit case references: "某某公司的案例…" (The case of X company...)
  • Attribution patterns: "巴菲特告诉我…" (Buffett told me...)
  • Illustrative transitions: "比如…" (For example...)

The detection algorithm looks for sequences combining temporal markers with reflective commentary or decision-making narratives. This distinguishes case examples from general exposition or hypothetical scenarios.

Eligibility Validation and Scope Rules

Not all detected patterns qualify for extraction. The extractor applies strict scope rules to filter candidates:

  • Author-anchored requirement: The event must represent a decision or experience directly undertaken by the author, or a historical case explicitly used to explain a specific method
  • Methodological binding: Each case must link to at least one methodological topic via the mandatory bound_to field
  • Exclusion filters: Pure background narration, fictional fables, and standalone principles without applied context are rejected

This validation ensures that only actionable, concept-illustrating examples proceed to the knowledge base, excluding anecdotal content that lacks methodological relevance.

Output Schema and YAML Structure

Qualified cases are serialized into structured YAML blocks with specific mandatory fields. The schema defined in extractors/case-extractor.md requires:

- id: <unique-id>
  title: <short description>
  type: case
  source_chapter: <chapter-name>
  source_quote: |
    "<original excerpt from the book>"
  summary: |
    <author’s reflective interpretation of the case>
  bound_to:                # at least one method-topic required

    - "<topic-1>"
    - "<topic-2>"
  outcome: |
    <result or impact described in the book, if any>
  tags: [case, <optional-tags>]

The bound_to field is particularly critical—it creates the semantic bridge between concrete examples and abstract methodology, enabling downstream consumption by Stage 2 – A1 (Past Application) processors.

Integration with the Pipeline

Extracted cases persist to candidates/cases.md, a generated Markdown file that aggregates all case YAML blocks from the extraction run. This file serves as the intermediate storage before Stage 2 processing combines cases with principles, frameworks, and counter-examples to construct the final skill knowledge base.

The central registry in SKILL.md maps the Case Extractor to its specification document (extractors/case-extractor.md) and tracks its output destination, maintaining the pipeline's provenance chain.

Implementation Example

Developers can invoke the extractor programmatically as shown in the repository's reference implementation:

from extractors.case_extractor import CaseExtractor

extractor = CaseExtractor(
    overview_path="BOOK_OVERVIEW.md",
    book_text_path="book.txt"
)

cases = extractor.run()
for case in cases:
    print(case["id"], case["title"], case["bound_to"])

Typical output follows this structure:

- id: c01
  title: 投资 See's Candy
  type: case
  source_chapter: 第 5 讲
  source_quote: |
    "我们以 2500 万美元收购了 See's Candy...这是我们第一次为品牌溢价付费。"
  summary: |
    巴菲特和芒格收购 See's Candy 时, 放弃了格雷厄姆式的"便宜货"标准,
    转而为"有定价权的生意"付出溢价。这笔投资后来成了他们转向
    "优质企业+合理价格"策略的转折点。
  bound_to:
    - "能力圈 + 定价权"
    - "从便宜货到优质企业的转变"
  outcome: |
    该公司后续 30 年产生的现金流远超初始投资, 验证了新策略。
  tags: [case, investment, turning-point]

Summary

  • cangjie-skill uses a dedicated Case Extractor during Stage 1 parallel extraction to identify author-anchored examples from source texts.
  • The extractor requires BOOK_OVERVIEW.md and the raw book text as inputs to contextualize findings within chapter structures.
  • Signal detection relies on regex patterns identifying past-tense narratives with reflective commentary, filtering for direct author experience or explicit method illustration.
  • Mandatory scope rules require every case to include a bound_to field linking it to at least one methodological topic, excluding pure narration or fables.
  • Valid cases output to candidates/cases.md as structured YAML, feeding into Stage 2 A1 (Past Application) processing for final knowledge base assembly.

Frequently Asked Questions

What input files does the Case Extractor require to function?

The Case Extractor requires two specific inputs: the BOOK_OVERVIEW.md file containing the structured book outline, and the raw book text itself. These files enable the extractor to map discovered examples to their correct chapter contexts while scanning for linguistic patterns.

How does cangjie-skill distinguish valid case examples from background narration?

The extractor applies strict eligibility checks based on scope rules. Valid cases must represent decisions or experiences directly made by the author, or historical cases explicitly used to explain a method. Pure background stories, fictional fables, and standalone principles without methodological application are automatically excluded.

Why is the bound_to field mandatory in case extraction output?

The bound_to field creates the critical semantic link between concrete examples and abstract methodology. This mandatory linkage ensures that Stage 2 processors—specifically the A1 (Past Application) module—can consume case data effectively, connecting real-world applications to their underlying conceptual frameworks.

Where are extracted case examples stored in the cangjie-skill pipeline?

Extracted cases are written to candidates/cases.md during Stage 1 processing. This generated file serves as intermediate storage containing all YAML-structured case blocks, which later aggregate with principles, frameworks, and counter-examples to construct the final skill knowledge base.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →