What Is the Purpose and Generation of GLOSSARY.md in Cangjie-Skill?
GLOSSARY.md serves as a centralized term dictionary shared across every skill produced by the Cangjie-Skill pipeline, automatically generated from extracted candidate terminology during the build process.
In the kangarooking/cangjie-skill repository, GLOSSARY.md plays a critical role in making technical skills accessible to readers. This article examines how this file is structured, why it matters for knowledge consumption, and the exact pipeline stages that create it.
Core Responsibilities of GLOSSARY.md
The GLOSSARY.md file fulfills three essential functions in the Cangjie-Skill ecosystem:
- Collecting key concepts — Extracts terminology from source materials such as research papers and domain articles.
- Providing a single reference point — Eliminates the need to search through individual skill sections for definitions.
- Ensuring repository visibility — Lives at the repository root (outside audit folders) and appears in the root-level
INDEX.mdfor immediate discoverability.
How GLOSSARY.md Is Generated
The generation of GLOSSARY.md follows a two-stage automated pipeline documented in the project's methodology files.
Stage 1: Parallel Extraction Creates the Candidate
During the Parallel Extraction stage, the glossary extractor scans raw source text and produces a candidate file. This intermediate output lives at candidates/glossary.md and contains raw term-definition pairs discovered in the source material.
As described in [methodology/02-stage1-parallel-extract.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md#L40), the extractor identifies domain-specific vocabulary and structures it for later processing.
Stage 2: Zettelkasten Promotes to Final Location
In the subsequent Zettelkasten stage, the pipeline transforms the candidate into the finished artifact. The system promotes candidates/glossary.md to books/<slug>/GLOSSARY.md.
According to [methodology/05-stage3-zettelkasten.md](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md#L33), this promotion step finalizes formatting and ensures the glossary integrates properly with the generated skill book. The SKILL.md overview confirms this transformation at line 129, describing the complete workflow as candidates/glossary.md → GLOSSARY.md.
Final Integration
Once generated, GLOSSARY.md receives automatic linkage from the root-level INDEX.md, making it part of the published skill output without manual intervention.
Structure and Format
The generated GLOSSARY.md follows simple Markdown conventions. Here is a representative example of its contents:
# Example entry inside the generated GLOSSARY.md
## 漢字 (Hanzi)
A logographic character used in Chinese writing, also adopted by Japanese (Kanji) and Korean (Hanja).
## Cangjie
A Chinese input method that maps characters to key sequences based on the shape of their components.
Each entry uses an H2 heading for the term followed by its plain-text definition. The file remains human-readable and can be viewed directly in the repository once a skill build completes.
Key Source Files
Understanding GLOSSARY.md generation requires familiarity with these repository locations:
methodology/02-stage1-parallel-extract.md— Defines the glossary extractor andcandidates/glossary.mdoutput.methodology/05-stage3-zettelkasten.md— Documents promotion tobooks/<slug>/GLOSSARY.md.SKILL.md— Summarizes the complete glossary generation workflow.extractors/glossary-extractor.md— Contains the extractor definition itself.books/<slug>/GLOSSARY.md— The final generated file in completed skill builds.
Summary
GLOSSARY.mdacts as a shared terminology reference across all skills in the Cangjie-Skill pipeline.- Generation is fully automated through the glossary extractor during Parallel Extraction and promotion during Zettelkasten.
- The file path transforms from
candidates/glossary.mdtobooks/<slug>/GLOSSARY.mdbefore publication. - Root-level visibility in
INDEX.mdensures readers can always access definitions without searching individual sections.
Frequently Asked Questions
What triggers the creation of GLOSSARY.md?
The glossary extractor runs automatically during the Parallel Extraction stage whenever source material contains extractable terminology. No manual invocation is required—the pipeline detects domain vocabulary and generates candidates/glossary.md as standard output.
Can I edit GLOSSARY.md after generation?
Direct editing is discouraged because subsequent pipeline runs will overwrite changes. To modify glossary content, adjust the source material or the extractor configuration in extractors/glossary-extractor.md, then rebuild the skill.
Why does GLOSSARY.md appear at the repository root rather than inside book folders?
The root-level INDEX.md references books/<slug>/GLOSSARY.md to create a unified entry point. The file technically resides within the book directory structure, but standardized linking makes it effectively available from the top-level navigation.
How does the glossary differ from inline definitions within skill sections?
Inline definitions explain terms in context for immediate comprehension. GLOSSARY.md consolidates all terminology into a single searchable reference, supporting readers who need quick lookups without reviewing entire sections.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →