# How the Ubiquitous Language Skill Extracts DDD Glossary Terms: A Complete Pipeline Guide

> Learn how the ubiquitous language skill extracts DDD glossary terms by parsing development conversations. Discover its deterministic seven-step pipeline for conflict resolution and glossary compilation.

- Repository: [Matt Pocock/skills](https://github.com/mattpocock/skills)
- Tags: how-to-guide
- Published: 2026-04-04

---

**The ubiquitous language skill parses development conversations to identify domain concepts, flags linguistic conflicts like ambiguity and synonyms, and compiles a canonical Domain-Driven Design glossary into [`UBIQUITOUS_LANGUAGE.md`](https://github.com/mattpocock/skills/blob/main/UBIQUITOUS_LANGUAGE.md) through a deterministic seven-step pipeline.**

The **ubiquitous language skill** is a rule-based tool in the `mattpocock/skills` repository that automates the extraction of **Domain-Driven Design (DDD)** terminology from unstructured dialogue. According to the source code in [`ubiquitous-language/SKILL.md`](https://github.com/mattpocock/skills/blob/main/ubiquitous-language/SKILL.md), the skill implements a deterministic pipeline that transforms conversational text into a structured, version-controlled glossary suitable for bounded contexts and strategic design.

## The Seven-Step DDD Term Extraction Pipeline

The skill processes conversation transcripts through distinct phases defined in the SKILL.md specification.

### 1. Conversation Scanning for Domain Concepts

The skill walks through the entire dialogue looking for **domain-relevant nouns, verbs, and concepts** (line 12). This initial pass identifies candidate terms that represent core domain entities, value objects, or processes without requiring explicit markup or annotations from participants.

### 2. Linguistic Problem Identification

While scanning, the skill flags three specific categories of linguistic issues that threaten DDD consistency:

- **Ambiguity**: The same word used for different concepts (lines 14-15)
- **Synonyms**: Different words referring to the same concept (lines 15-16)
- **Vagueness**: Terms that are overloaded or loosely defined (lines 16-17)

### 3. Canonical Glossary Construction

After identification, the skill builds one or more markdown tables. Each row contains a **Term**, a concise definition capped at one sentence, and an **Aliases-to-avoid** column listing synonyms or conflicting names (lines 30-34). This creates the authoritative dictionary for the domain.

### 4. Relationship Extraction

When conversations describe connections between concepts—such as "an **Invoice** belongs to exactly one **Customer**"—the skill captures these as bullet-point statements in a dedicated **Relationships** section (lines 44-45). This maps domain associations critical for understanding aggregates and boundaries.

### 5. Ambiguity Reporting

All flagged ambiguities are collected under a **Flagged ambiguities** heading, each accompanied by a clear recommendation to maintain terminological consistency (lines 56-57). This forces explicit resolution of overloaded language before it contaminates the codebase.

### 6. Example Dialogue Generation

To illustrate proper usage, the skill inserts a short mock conversation containing three to five exchanges that demonstrate the new terms in context (lines 68-71). This serves as executable documentation for domain newcomers.

### 7. File Persistence and Output

The assembled markdown is written to [`UBIQUITOUS_LANGUAGE.md`](https://github.com/mattpocock/skills/blob/main/UBIQUITOUS_LANGUAGE.md) in the working directory (lines 18-19), and a concise summary is echoed back to the user (lines 19-20). The output follows strict formatting rules ensuring compatibility with version control and static site generators.

## Incremental Updates and State Management

When the skill is re-invoked within the same session, it first **loads the existing glossary** from the previous run. It then merges newly discovered terms, updates existing definitions, re-flags any new ambiguities, and rewrites the example dialogue to keep the document current (lines 88-92). This incremental approach prevents glossary drift and maintains historical accuracy of domain evolution.

## Practical Implementation Example

Below is a minimal illustration of how a client triggers the skill and the resulting [`UBIQUITOUS_LANGUAGE.md`](https://github.com/mattpocock/skills/blob/main/UBIQUITOUS_LANGUAGE.md) structure.

```python

# Pseudo-code for invoking the ubiquitous language skill

conversation = [
    "Dev: When a Customer places an Order, do we create the Invoice immediately?",
    "Domain: No — an Invoice is only generated after Fulfillment is confirmed.",
    "Dev: If a Shipment is cancelled before dispatch, no Invoice exists for it?",
    "Domain: Exactly."
]

# The skill is called with the current transcript

ubiquitous_language = run_skill("ubiquitous-language", conversation)

print(ubiquitous_language.summary)   # Inline summary shown to the user

# The skill writes UBIQUITOUS_LANGUAGE.md to the cwd

```

Resulting **UBIQUITOUS_LANGUAGE.md** (truncated):

```markdown

# Ubiquitous Language

## Order lifecycle

| Term        | Definition                                            | Aliases to avoid      |
| ----------- | ----------------------------------------------------- | --------------------- |
| **Order**   | A customer's request to purchase one or more items     | Purchase, transaction |
| **Invoice** | A request for payment sent after fulfillment is done  | Bill, payment request |

## People

| Term         | Definition                                     | Aliases to avoid |
| ------------ | ---------------------------------------------- | ---------------- |
| **Customer** | A person or organisation that places orders    | Client, buyer    |
| **User**     | An authentication identity in the system       | Login, account   |

## Relationships

- An **Invoice** belongs to exactly one **Customer**
- An **Order** produces one or more **Invoices**

## Flagged ambiguities

- "account" was used to mean both **Customer** and **User** …

```

## Summary

- The ubiquitous language skill implements a **deterministic seven-step pipeline** defined in [`ubiquitous-language/SKILL.md`](https://github.com/mattpocock/skills/blob/main/ubiquitous-language/SKILL.md) to extract DDD terms from conversation.
- It specifically targets **ambiguity, synonyms, and vagueness** that violate ubiquitous language principles.
- Output is standardized as **markdown tables** in [`UBIQUITOUS_LANGUAGE.md`](https://github.com/mattpocock/skills/blob/main/UBIQUITOUS_LANGUAGE.md) with Terms, Definitions, and Aliases-to-avoid columns.
- The skill extracts **relationships** between domain concepts and documents them separately.
- **Incremental updates** allow the glossary to evolve across multiple sessions without losing historical context.
- Each run generates **example dialogues** demonstrating correct term usage in context.

## Frequently Asked Questions

### What triggers the ubiquitous language skill to extract terms?

The skill activates when provided a conversation transcript containing domain discussion. It scans every utterance for nouns, verbs, and concepts relevant to the business domain, then applies deterministic rules defined in [`ubiquitous-language/SKILL.md`](https://github.com/mattpocock/skills/blob/main/ubiquitous-language/SKILL.md) to filter and categorize these terms into a structured glossary.

### How does the skill handle ambiguous terminology?

When the same word appears with different meanings across the conversation, the skill flags this as an ambiguity under the **Flagged ambiguities** section (lines 56-57). Each flag includes a recommendation to resolve the inconsistency, ensuring that the final glossary maintains the one-term-one-meaning principle required by Domain-Driven Design.

### Can the ubiquitous language skill update an existing glossary?

Yes. When invoked in a session where [`UBIQUITOUS_LANGUAGE.md`](https://github.com/mattpocock/skills/blob/main/UBIQUITOUS_LANGUAGE.md) already exists, the skill loads the previous document, merges new terms, updates definitions, and rewrites the example dialogue (lines 88-92). This incremental approach prevents duplicate entries while allowing the domain model to evolve naturally over time.

### What is the structure of the generated UBIQUITOUS_LANGUAGE.md file?

The output file contains markdown tables organizing terms by domain area, each with three columns: **Term**, **Definition** (maximum one sentence), and **Aliases to avoid** (lines 30-34). It also includes a **Relationships** section for domain associations and a **Flagged ambiguities** section for unresolved linguistic conflicts.