How Glossary Extraction and Indexing Works in Cangjie-Skill: A Two-Stage Pipeline
The cangjie-skill repository implements glossary extraction and indexing through a coordinated two-stage pipeline that first extracts terminology candidates in parallel from book content, then promotes the curated glossary to a canonical location with Zettelkasten-style linking.
The cangjie-skill project automates the creation of structured knowledge artifacts from book content. Understanding the glossary extraction and indexing process reveals how the pipeline identifies critical terminology and makes it discoverable as a shared resource across downstream skills.
Stage 1: Parallel Glossary Extraction
The glossary extraction process runs as one of five concurrent extractors in Stage 1 of the pipeline, as defined in methodology/02-stage1-parallel-extract.md. This parallel execution ensures that terminology identification happens simultaneously with other content analysis tasks.
Input Sources and Selection Criteria
The Glossary Extractor receives two primary inputs: the book’s overview (BOOK_OVERVIEW.md) and the raw manuscript text. It scans the entire document and selects terms that satisfy any of the following four criteria:
- The term appears at least three times throughout the book.
- It is explicitly defined by the author using patterns like “所谓 X ,是指 …”.
- It appears to be a common word but is used in a non-standard sense.
- It forms part of the book’s core argument (such as antifragile in Antifragile).
This selection logic ensures the extractor captures only essential, author-specific terminology rather than general vocabulary.
YAML Output Schema and Temp File Storage
For each selected term, the extractor generates a structured YAML entry containing specific fields:
id: Unique identifier (e.g.,g01)term: The extracted terminologytype: Classification (typically "term")source_chapter: Origin location in the manuscriptauthor_definition: Verbatim author explanationkey_distinction: Contrast between author usage and standard meaningwhy_it_matters: Explanation of downstream relevance for skill generationtags: Categorical markers (e.g.,[term, core-concept])
The extractor writes these entries to a temporary candidate file at candidates/glossary.md. This temporary location allows for review before final promotion to the canonical glossary. The extraction logic is fully specified in extractors/glossary-extractor.md.
Stage 3: Zettelkasten-Style Indexing and Promotion
After all parallel extractors complete their work, the pipeline enters Stage 3, described in methodology/05-stage3-zettelkasten.md. This stage transforms temporary extraction outputs into permanent, linked artifacts.
Canonical Location Promotion
The temporary candidate file candidates/glossary.md is promoted to the public repository location books/<slug>/GLOSSARY.md, where <slug> represents the book identifier. This promotion establishes the glossary as the canonical terminology dictionary shared by every downstream skill derived from the book.
Master Index Integration
Simultaneously, the pipeline updates the master skill map (INDEX.md) to include a bidirectional link to the newly created glossary. The templates/INDEX.md.template automatically generates the entry:
术语词典: GLOSSARY.md
This linking follows Zettelkasten principles, ensuring that any skill in the repository can reference standardized definitions through the centralized glossary.
Key Implementation Files
The glossary extraction and indexing process relies on these specific repository files:
extractors/glossary-extractor.md: Contains the complete specification for selection criteria, extraction rules, and output format requirements.methodology/02-stage1-parallel-extract.md: Documents the parallel execution architecture where the glossary extractor runs concurrently with four other extractors.methodology/05-stage3-zettelkasten.md: Details the post-extraction linking logic and file promotion workflow.templates/INDEX.md.template: Defines the template structure that inserts the glossary link into the master skill map.SKILL.md: Provides the high-level pipeline overview and describes how Stage 1 and Stage 3 coordinate to produce the final glossary artifact.
Practical Implementation Examples
The following YAML structure represents a typical glossary entry produced by the extraction process:
- id: g01
term: 能力圈
type: term
source_chapter: 第 2 讲
author_definition: |
"你真正能做出准确判断的知识边界。不是你知道什么, 而是你知道'你知道什么'和'你不知道什么'的边界。"
key_distinction: |
≠ "熟悉的领域" — 熟悉不代表能做判断
≠ "专业领域" — 博士学位也可能在能力圈外
= 能持续做出比市场更准判断的范围 (需经实战验证)
why_it_matter: |
"能力圈"一词在所有投资决策类 skill 中都会出现。若沿用字典义, skill 会建议用户"评估一下是否熟悉该领域", 这是错的。
tags: [term, core-concept]
While the pipeline invokes this automatically, the conceptual command for running the glossary extractor aligns with this pattern:
python -m extractors.glossary_extractor \
--overview BOOK_OVERVIEW.md \
--source-text full_book.txt \
--output candidates/glossary.md
During the Zettelkasten indexing stage, the promotion and index update operations follow this sequence:
mv candidates/glossary.md books/${BOOK_SLUG}/GLOSSARY.md
# Update the index with the glossary link (handled by template expansion)
sed -i "s|# INDEX|# INDEX\n- **术语词典**: [GLOSSARY.md](./GLOSSARY.md)|" INDEX.md
Summary
- Glossary extraction operates as a parallel process in Stage 1, scanning book content against four specific criteria to identify critical terminology.
- Selection criteria include frequency (≥3 occurrences), explicit author definition, non-standard usage of common words, and centrality to the book’s core argument.
- YAML output captures structured metadata including definitions, key distinctions, and relevance explanations in
candidates/glossary.md. - Indexing promotion moves the temporary file to
books/<slug>/GLOSSARY.mdduring Stage 3, establishing the canonical terminology dictionary. - Zettelkasten linking automatically integrates the glossary into the master skill map via the template in
templates/INDEX.md.template.
Frequently Asked Questions
What triggers the glossary extraction process in the pipeline?
The glossary extraction initiates automatically during Stage 1 when the pipeline receives a new book submission. The extractor receives BOOK_OVERVIEW.md and the raw manuscript text as inputs, then executes concurrently with four other specialized extractors. This parallel execution ensures comprehensive content analysis without sequential bottlenecks.
How does the pipeline distinguish between common words and technical terms?
The extractor applies the key_distinction field to differentiate author-specific usage from dictionary definitions. It flags terms that appear to be common words but carry non-standard meanings within the book’s context, and explicitly excludes general vocabulary unless the author redefines them or they appear at least three times with specific technical significance.
Where is the final glossary stored and how is it accessed by downstream skills?
After Stage 3 processing, the final glossary resides at books/<slug>/GLOSSARY.md in the repository root. Downstream skills access this canonical file through relative links from the master INDEX.md, which contains the entry "术语词典: GLOSSARY.md" generated by the index template.
Can the glossary extraction criteria be customized for specific domains?
Yes, the selection rules are defined in extractors/glossary-extractor.md and can be modified to adjust frequency thresholds or pattern matching for explicit definitions. However, the four core criteria (frequency, explicit definition, non-standard usage, and core argument relevance) provide the framework that ensures comprehensive coverage across diverse subject domains.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →