How GLOSSARY.md Functions as a Shared Terminology Dictionary in cangjie-skill

GLOSSARY.md serves as a centralized, shared terminology dictionary that guarantees consistent vocabulary across all AI skills generated by the cangjie-skill pipeline, preventing semantic drift when author-specific terms like "能力圈" (ability circle) risk being misinterpreted as generic concepts.

The kangarooking/cangjie-skill project implements a sophisticated book-to-skill conversion system where terminology precision directly impacts downstream AI behavior. The GLOSSARY.md file sits at the heart of this system, ensuring that specialized terms extracted from source material carry their original authorial meaning into every derived skill.

The Four-Phase Lifecycle of a Shared Terminology Dictionary

Phase 1: Intelligent Extraction by the Glossary Extractor

The Glossary Extractor (extractors/glossary-extractor.md) performs the initial term identification. According to the source code, this component scans the source book and selects terms meeting any of four strict criteria:

  • Appear at least three times in the book
  • Are explicitly defined by the author
  • Are common words used with non-standard meaning
  • Form part of the book's core arguments

The extractor outputs structured YAML containing term objects with fields for id, term, source_chapter, author_definition, and additional metadata. This structured format ensures machine-readable precision rather than loose textual definitions.

Phase 2: Promotion to Public-Facing Location

After parallel extraction, raw candidate files undergo a consolidation step. The methodology documentation in methodology/05-stage3-zettelkasten.md specifies that candidates/glossary.md is promoted to books/<slug>/GLOSSARY.md during Stage 3 (the Zettelkasten step). This placement at the root of each book's output directory establishes a well-known, stable location that downstream components can reliably reference.

Phase 3: Universal Reference Across All Skills

The shared terminology dictionary becomes the single source of truth for all downstream skill types. Whether generating framework skills, principle skills, case studies, or counter-example skills, each component imports terms from the same GLOSSARY.md. This architectural choice prevents a critical failure mode: independent interpretation of author-specific phrasing.

Consider the documented example of "能力圈" (ability circle). Without the shared dictionary, a skill might treat this as the generic phrase "ability range" and advise users to "evaluate whether the field is familiar." The corrected interpretation—"evaluate your historical judgment accuracy in this domain"—only propagates correctly because all skills reference the identical, author-curated definition from GLOSSARY.md.

Phase 4: Index Integration for Discoverability

The INDEX.md for each book explicitly links to its GLOSSARY.md, as specified in the Stage 3 instructions. This integration ensures that readers navigating the skill network can immediately locate authoritative term definitions without hunting through individual skill files.

YAML Structure: How Term Definitions Are Encoded

The shared terminology dictionary uses a precise schema that captures not just definitions but critical distinctions and usage guidance:


# Example entry produced by the Glossary Extractor

- id: g01
  term: 能力圈
  type: term
  source_chapter: 第 2 讲
  author_definition: |
    "你真正能做出准确判断的知识边界。不是你知道什么, 而是你知道'你知道什么'和'你不知道什么'的边界。"
  key_distinction: |
    ≠ "熟悉的领域" — 熟悉不代表能做判断
    ≠ "专业领域" — 博士学位也可能在能力圈外
    = 能持续做出比市场更准判断的范围 (需经实战验证)
  why_it_matters: |
    "能力圈"一词在所有投资决策类 skill 中都会出现。
    若沿用字典义, skill 会建议用户"评估一下是否熟悉该领域", 这是错的。
    正确的用法是"评估自己过去在此领域的判断准确率"。
  tags: [term, core-concept]

The key_distinction field is particularly significant—it explicitly documents negative definitions (what the term is not) alongside positive ones, preventing the most common form of semantic drift in AI systems.

Consolidation Workflow: From Candidate to Shared Dictionary

The pipeline enforces a strict promotion workflow:


# After extraction, promote the candidate file to the shared location

mv candidates/glossary.md books/my-book/GLOSSARY.md

# Verify that downstream skills can import it

grep -R "GLOSSARY.md" books/my-book/*  # should show references in SKILL.md, INDEX.md, etc.

This separation between candidates/ (temporary, parallel-extraction outputs) and books/<slug>/ (final, canonical outputs) ensures that only reviewed, consolidated terminology enters the shared namespace.

Architectural Significance in the cangjie-skill Pipeline

The GLOSSARY.md system addresses a fundamental challenge in AI skill generation: authorial terminology versus model prior knowledge. Large language models come with broad but shallow linguistic priors; specialized authors develop deep but narrow terminological precision. The shared terminology dictionary resolves this tension by:

  1. Overriding model defaults with author definitions for specific terms
  2. Propagating these overrides consistently across all skill variants
  3. Documenting the rationale for non-obvious interpretations

As implemented in kangarooking/cangjie-skill, this architecture makes the difference between a skill that generically "understands" business concepts and one that operationalizes a specific author's investment philosophy with fidelity.

Summary

  • GLOSSARY.md is generated by the Glossary Extractor (extractors/glossary-extractor.md) using four selection criteria including frequency thresholds and explicit author definition
  • It is consolidated during Stage 3 Zettelkasten processing into a public-facing location at books/<slug>/GLOSSARY.md
  • It serves as shared reference for all downstream skills—framework, principle, case, and counter-example—ensuring consistent vocabulary usage
  • It prevents semantic drift by encoding precise author definitions, negative distinctions, and usage guidance rather than relying on model priors
  • It integrates with INDEX.md to maintain discoverability within the skill network

Frequently Asked Questions

What makes a term qualify for inclusion in GLOSSARY.md?

A term must satisfy at least one of four criteria defined in extractors/glossary-extractor.md: appearing three or more times in the source book, receiving explicit author definition, being a common word with specialized usage, or functioning as a core argument component. This multi-criteria approach captures both high-frequency terminology and infrequent but critical conceptual vocabulary.

How does the shared dictionary handle terms with multiple interpretations?

Each entry includes a key_distinction field that explicitly documents what the term does not mean, alongside its positive definition. For "能力圈", the dictionary records that it ≠ "familiar territory" and ≠ "professional domain", preventing the most likely misinterpretations. The why_it_matters field further explains real-world consequences of definitional errors.

Can downstream skills modify or extend GLOSSARY.md definitions?

No—the architecture treats GLOSSARY.md as read-only once promoted to books/<slug>/. This immutability guarantee is essential for the "shared" property; if individual skills could override definitions, vocabulary consistency would collapse. Extensions or corrections require regenerating the source glossary and re-running the consolidation pipeline.

Where does GLOSSARY.md sit in the overall cangjie-skill methodology?

It emerges from parallel extraction (one of five extractors running simultaneously per SKILL.md), consolidates during Stage 3 Zettelkasten processing as documented in methodology/05-stage3-zettelkasten.md, and persists as a canonical reference through all subsequent skill generation stages. The file's position at the book directory root makes it accessible to any component without path dependencies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →