How the Ubiquitous Language Skill Extracts DDD Glossary Terms: A Complete Pipeline Guide
The ubiquitous language skill parses development conversations to identify domain concepts, flags linguistic conflicts like ambiguity and synonyms, and compiles a canonical Domain-Driven Design glossary into UBIQUITOUS_LANGUAGE.md through a deterministic seven-step pipeline.
The ubiquitous language skill is a rule-based tool in the mattpocock/skills repository that automates the extraction of Domain-Driven Design (DDD) terminology from unstructured dialogue. According to the source code in ubiquitous-language/SKILL.md, the skill implements a deterministic pipeline that transforms conversational text into a structured, version-controlled glossary suitable for bounded contexts and strategic design.
The Seven-Step DDD Term Extraction Pipeline
The skill processes conversation transcripts through distinct phases defined in the SKILL.md specification.
1. Conversation Scanning for Domain Concepts
The skill walks through the entire dialogue looking for domain-relevant nouns, verbs, and concepts (line 12). This initial pass identifies candidate terms that represent core domain entities, value objects, or processes without requiring explicit markup or annotations from participants.
2. Linguistic Problem Identification
While scanning, the skill flags three specific categories of linguistic issues that threaten DDD consistency:
- Ambiguity: The same word used for different concepts (lines 14-15)
- Synonyms: Different words referring to the same concept (lines 15-16)
- Vagueness: Terms that are overloaded or loosely defined (lines 16-17)
3. Canonical Glossary Construction
After identification, the skill builds one or more markdown tables. Each row contains a Term, a concise definition capped at one sentence, and an Aliases-to-avoid column listing synonyms or conflicting names (lines 30-34). This creates the authoritative dictionary for the domain.
4. Relationship Extraction
When conversations describe connections between concepts—such as "an Invoice belongs to exactly one Customer"—the skill captures these as bullet-point statements in a dedicated Relationships section (lines 44-45). This maps domain associations critical for understanding aggregates and boundaries.
5. Ambiguity Reporting
All flagged ambiguities are collected under a Flagged ambiguities heading, each accompanied by a clear recommendation to maintain terminological consistency (lines 56-57). This forces explicit resolution of overloaded language before it contaminates the codebase.
6. Example Dialogue Generation
To illustrate proper usage, the skill inserts a short mock conversation containing three to five exchanges that demonstrate the new terms in context (lines 68-71). This serves as executable documentation for domain newcomers.
7. File Persistence and Output
The assembled markdown is written to UBIQUITOUS_LANGUAGE.md in the working directory (lines 18-19), and a concise summary is echoed back to the user (lines 19-20). The output follows strict formatting rules ensuring compatibility with version control and static site generators.
Incremental Updates and State Management
When the skill is re-invoked within the same session, it first loads the existing glossary from the previous run. It then merges newly discovered terms, updates existing definitions, re-flags any new ambiguities, and rewrites the example dialogue to keep the document current (lines 88-92). This incremental approach prevents glossary drift and maintains historical accuracy of domain evolution.
Practical Implementation Example
Below is a minimal illustration of how a client triggers the skill and the resulting UBIQUITOUS_LANGUAGE.md structure.
# Pseudo-code for invoking the ubiquitous language skill
conversation = [
"Dev: When a Customer places an Order, do we create the Invoice immediately?",
"Domain: No — an Invoice is only generated after Fulfillment is confirmed.",
"Dev: If a Shipment is cancelled before dispatch, no Invoice exists for it?",
"Domain: Exactly."
]
# The skill is called with the current transcript
ubiquitous_language = run_skill("ubiquitous-language", conversation)
print(ubiquitous_language.summary) # Inline summary shown to the user
# The skill writes UBIQUITOUS_LANGUAGE.md to the cwd
Resulting UBIQUITOUS_LANGUAGE.md (truncated):
# Ubiquitous Language
## Order lifecycle
| Term | Definition | Aliases to avoid |
| ----------- | ----------------------------------------------------- | --------------------- |
| **Order** | A customer's request to purchase one or more items | Purchase, transaction |
| **Invoice** | A request for payment sent after fulfillment is done | Bill, payment request |
## People
| Term | Definition | Aliases to avoid |
| ------------ | ---------------------------------------------- | ---------------- |
| **Customer** | A person or organisation that places orders | Client, buyer |
| **User** | An authentication identity in the system | Login, account |
## Relationships
- An **Invoice** belongs to exactly one **Customer**
- An **Order** produces one or more **Invoices**
## Flagged ambiguities
- "account" was used to mean both **Customer** and **User** …
Summary
- The ubiquitous language skill implements a deterministic seven-step pipeline defined in
ubiquitous-language/SKILL.mdto extract DDD terms from conversation. - It specifically targets ambiguity, synonyms, and vagueness that violate ubiquitous language principles.
- Output is standardized as markdown tables in
UBIQUITOUS_LANGUAGE.mdwith Terms, Definitions, and Aliases-to-avoid columns. - The skill extracts relationships between domain concepts and documents them separately.
- Incremental updates allow the glossary to evolve across multiple sessions without losing historical context.
- Each run generates example dialogues demonstrating correct term usage in context.
Frequently Asked Questions
What triggers the ubiquitous language skill to extract terms?
The skill activates when provided a conversation transcript containing domain discussion. It scans every utterance for nouns, verbs, and concepts relevant to the business domain, then applies deterministic rules defined in ubiquitous-language/SKILL.md to filter and categorize these terms into a structured glossary.
How does the skill handle ambiguous terminology?
When the same word appears with different meanings across the conversation, the skill flags this as an ambiguity under the Flagged ambiguities section (lines 56-57). Each flag includes a recommendation to resolve the inconsistency, ensuring that the final glossary maintains the one-term-one-meaning principle required by Domain-Driven Design.
Can the ubiquitous language skill update an existing glossary?
Yes. When invoked in a session where UBIQUITOUS_LANGUAGE.md already exists, the skill loads the previous document, merges new terms, updates definitions, and rewrites the example dialogue (lines 88-92). This incremental approach prevents duplicate entries while allowing the domain model to evolve naturally over time.
What is the structure of the generated UBIQUITOUS_LANGUAGE.md file?
The output file contains markdown tables organizing terms by domain area, each with three columns: Term, Definition (maximum one sentence), and Aliases to avoid (lines 30-34). It also includes a Relationships section for domain associations and a Flagged ambiguities section for unresolved linguistic conflicts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →