How Glossary Extraction and Indexing Works in Cangjie-Skill: A Two-Stage Pipeline

The cangjie-skill repository implements glossary extraction and indexing through a coordinated two-stage pipeline that first extracts terminology candidates in parallel from book content, then promotes the curated glossary to a canonical location with Zettelkasten-style linking.

The cangjie-skill project automates the creation of structured knowledge artifacts from book content. Understanding the glossary extraction and indexing process reveals how the pipeline identifies critical terminology and makes it discoverable as a shared resource across downstream skills.

Stage 1: Parallel Glossary Extraction

The glossary extraction process runs as one of five concurrent extractors in Stage 1 of the pipeline, as defined in methodology/02-stage1-parallel-extract.md. This parallel execution ensures that terminology identification happens simultaneously with other content analysis tasks.

Input Sources and Selection Criteria

The Glossary Extractor receives two primary inputs: the book’s overview (BOOK_OVERVIEW.md) and the raw manuscript text. It scans the entire document and selects terms that satisfy any of the following four criteria:

  • The term appears at least three times throughout the book.
  • It is explicitly defined by the author using patterns like “所谓 X ,是指 …”.
  • It appears to be a common word but is used in a non-standard sense.
  • It forms part of the book’s core argument (such as antifragile in Antifragile).

This selection logic ensures the extractor captures only essential, author-specific terminology rather than general vocabulary.

YAML Output Schema and Temp File Storage

For each selected term, the extractor generates a structured YAML entry containing specific fields:

  • id: Unique identifier (e.g., g01)
  • term: The extracted terminology
  • type: Classification (typically "term")
  • source_chapter: Origin location in the manuscript
  • author_definition: Verbatim author explanation
  • key_distinction: Contrast between author usage and standard meaning
  • why_it_matters: Explanation of downstream relevance for skill generation
  • tags: Categorical markers (e.g., [term, core-concept])

The extractor writes these entries to a temporary candidate file at candidates/glossary.md. This temporary location allows for review before final promotion to the canonical glossary. The extraction logic is fully specified in extractors/glossary-extractor.md.

Stage 3: Zettelkasten-Style Indexing and Promotion

After all parallel extractors complete their work, the pipeline enters Stage 3, described in methodology/05-stage3-zettelkasten.md. This stage transforms temporary extraction outputs into permanent, linked artifacts.

Canonical Location Promotion

The temporary candidate file candidates/glossary.md is promoted to the public repository location books/<slug>/GLOSSARY.md, where <slug> represents the book identifier. This promotion establishes the glossary as the canonical terminology dictionary shared by every downstream skill derived from the book.

Master Index Integration

Simultaneously, the pipeline updates the master skill map (INDEX.md) to include a bidirectional link to the newly created glossary. The templates/INDEX.md.template automatically generates the entry:

术语词典: GLOSSARY.md

This linking follows Zettelkasten principles, ensuring that any skill in the repository can reference standardized definitions through the centralized glossary.

Key Implementation Files

The glossary extraction and indexing process relies on these specific repository files:

  • extractors/glossary-extractor.md: Contains the complete specification for selection criteria, extraction rules, and output format requirements.
  • methodology/02-stage1-parallel-extract.md: Documents the parallel execution architecture where the glossary extractor runs concurrently with four other extractors.
  • methodology/05-stage3-zettelkasten.md: Details the post-extraction linking logic and file promotion workflow.
  • templates/INDEX.md.template: Defines the template structure that inserts the glossary link into the master skill map.
  • SKILL.md: Provides the high-level pipeline overview and describes how Stage 1 and Stage 3 coordinate to produce the final glossary artifact.

Practical Implementation Examples

The following YAML structure represents a typical glossary entry produced by the extraction process:

- id: g01
  term: 能力圈
  type: term
  source_chapter: 第 2 讲
  author_definition: |
    "你真正能做出准确判断的知识边界。不是你知道什么, 而是你知道'你知道什么'和'你不知道什么'的边界。"
  key_distinction: |
    ≠ "熟悉的领域" — 熟悉不代表能做判断
    ≠ "专业领域" — 博士学位也可能在能力圈外
    = 能持续做出比市场更准判断的范围 (需经实战验证)
  why_it_matter: |
    "能力圈"一词在所有投资决策类 skill 中都会出现。若沿用字典义, skill 会建议用户"评估一下是否熟悉该领域", 这是错的。
  tags: [term, core-concept]

While the pipeline invokes this automatically, the conceptual command for running the glossary extractor aligns with this pattern:

python -m extractors.glossary_extractor \
  --overview BOOK_OVERVIEW.md \
  --source-text full_book.txt \
  --output candidates/glossary.md

During the Zettelkasten indexing stage, the promotion and index update operations follow this sequence:

mv candidates/glossary.md books/${BOOK_SLUG}/GLOSSARY.md

# Update the index with the glossary link (handled by template expansion)

sed -i "s|# INDEX|# INDEX\n- **术语词典**: [GLOSSARY.md](./GLOSSARY.md)|" INDEX.md

Summary

  • Glossary extraction operates as a parallel process in Stage 1, scanning book content against four specific criteria to identify critical terminology.
  • Selection criteria include frequency (≥3 occurrences), explicit author definition, non-standard usage of common words, and centrality to the book’s core argument.
  • YAML output captures structured metadata including definitions, key distinctions, and relevance explanations in candidates/glossary.md.
  • Indexing promotion moves the temporary file to books/<slug>/GLOSSARY.md during Stage 3, establishing the canonical terminology dictionary.
  • Zettelkasten linking automatically integrates the glossary into the master skill map via the template in templates/INDEX.md.template.

Frequently Asked Questions

What triggers the glossary extraction process in the pipeline?

The glossary extraction initiates automatically during Stage 1 when the pipeline receives a new book submission. The extractor receives BOOK_OVERVIEW.md and the raw manuscript text as inputs, then executes concurrently with four other specialized extractors. This parallel execution ensures comprehensive content analysis without sequential bottlenecks.

How does the pipeline distinguish between common words and technical terms?

The extractor applies the key_distinction field to differentiate author-specific usage from dictionary definitions. It flags terms that appear to be common words but carry non-standard meanings within the book’s context, and explicitly excludes general vocabulary unless the author redefines them or they appear at least three times with specific technical significance.

Where is the final glossary stored and how is it accessed by downstream skills?

After Stage 3 processing, the final glossary resides at books/<slug>/GLOSSARY.md in the repository root. Downstream skills access this canonical file through relative links from the master INDEX.md, which contains the entry "术语词典: GLOSSARY.md" generated by the index template.

Can the glossary extraction criteria be customized for specific domains?

Yes, the selection rules are defined in extractors/glossary-extractor.md and can be modified to adjust frequency thresholds or pattern matching for explicit definitions. However, the four core criteria (frequency, explicit definition, non-standard usage, and core argument relevance) provide the framework that ensures comprehensive coverage across diverse subject domains.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →