# What Is the Purpose and Generation of GLOSSARY.md in Cangjie-Skill?

> Discover the purpose and generation of GLOSSARY.md in Cangjie-Skill. Learn how this centralized term dictionary is automatically created from candidate terminology during the build process.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: documentation
- Published: 2026-08-14

---

**[`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) serves as a centralized term dictionary shared across every skill produced by the Cangjie-Skill pipeline, automatically generated from extracted candidate terminology during the build process.**

In the `kangarooking/cangjie-skill` repository, [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) plays a critical role in making technical skills accessible to readers. This article examines how this file is structured, why it matters for knowledge consumption, and the exact pipeline stages that create it.

## Core Responsibilities of GLOSSARY.md

The [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) file fulfills three essential functions in the Cangjie-Skill ecosystem:

- **Collecting key concepts** — Extracts terminology from source materials such as research papers and domain articles.
- **Providing a single reference point** — Eliminates the need to search through individual skill sections for definitions.
- **Ensuring repository visibility** — Lives at the repository root (outside audit folders) and appears in the root-level [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md) for immediate discoverability.

## How GLOSSARY.md Is Generated

The generation of [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) follows a two-stage automated pipeline documented in the project's methodology files.

### Stage 1: Parallel Extraction Creates the Candidate

During the **Parallel Extraction** stage, the **glossary extractor** scans raw source text and produces a candidate file. This intermediate output lives at [`candidates/glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/glossary.md) and contains raw term-definition pairs discovered in the source material.

As described in [[`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md)](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md#L40), the extractor identifies domain-specific vocabulary and structures it for later processing.

### Stage 2: Zettelkasten Promotes to Final Location

In the subsequent **Zettelkasten** stage, the pipeline transforms the candidate into the finished artifact. The system promotes [`candidates/glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/glossary.md) to `books/<slug>/GLOSSARY.md`.

According to [[`methodology/05-stage3-zettelkasten.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md)](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md#L33), this promotion step finalizes formatting and ensures the glossary integrates properly with the generated skill book. The [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) overview confirms this transformation at [line 129](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md#L129), describing the complete workflow as `candidates/glossary.md → GLOSSARY.md`.

### Final Integration

Once generated, [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) receives automatic linkage from the root-level [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md), making it part of the published skill output without manual intervention.

## Structure and Format

The generated [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) follows simple Markdown conventions. Here is a representative example of its contents:

```markdown

# Example entry inside the generated GLOSSARY.md

## 漢字 (Hanzi)

A logographic character used in Chinese writing, also adopted by Japanese (Kanji) and Korean (Hanja).

## Cangjie

A Chinese input method that maps characters to key sequences based on the shape of their components.

```

Each entry uses an H2 heading for the term followed by its plain-text definition. The file remains human-readable and can be viewed directly in the repository once a skill build completes.

## Key Source Files

Understanding [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) generation requires familiarity with these repository locations:

- [`methodology/02-stage1-parallel-extract.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/02-stage1-parallel-extract.md) — Defines the glossary extractor and [`candidates/glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/glossary.md) output.
- [`methodology/05-stage3-zettelkasten.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md) — Documents promotion to `books/<slug>/GLOSSARY.md`.
- [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) — Summarizes the complete glossary generation workflow.
- [`extractors/glossary-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md) — Contains the extractor definition itself.
- `books/<slug>/GLOSSARY.md` — The final generated file in completed skill builds.

## Summary

- [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) acts as a **shared terminology reference** across all skills in the Cangjie-Skill pipeline.
- Generation is **fully automated** through the glossary extractor during Parallel Extraction and promotion during Zettelkasten.
- The file path transforms from [`candidates/glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/glossary.md) to `books/<slug>/GLOSSARY.md` before publication.
- Root-level visibility in [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md) ensures readers can always access definitions without searching individual sections.

## Frequently Asked Questions

### What triggers the creation of GLOSSARY.md?

The glossary extractor runs automatically during the **Parallel Extraction** stage whenever source material contains extractable terminology. No manual invocation is required—the pipeline detects domain vocabulary and generates [`candidates/glossary.md`](https://github.com/kangarooking/cangjie-skill/blob/main/candidates/glossary.md) as standard output.

### Can I edit GLOSSARY.md after generation?

Direct editing is discouraged because subsequent pipeline runs will overwrite changes. To modify glossary content, adjust the source material or the extractor configuration in [`extractors/glossary-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md), then rebuild the skill.

### Why does GLOSSARY.md appear at the repository root rather than inside book folders?

The root-level [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md) references `books/<slug>/GLOSSARY.md` to create a unified entry point. The file technically resides within the book directory structure, but standardized linking makes it effectively available from the top-level navigation.

### How does the glossary differ from inline definitions within skill sections?

Inline definitions explain terms in context for immediate comprehension. [`GLOSSARY.md`](https://github.com/kangarooking/cangjie-skill/blob/main/GLOSSARY.md) consolidates all terminology into a single searchable reference, supporting readers who need quick lookups without reviewing entire sections.