# How the `/understand-domain` Command Extracts Business Domains, Flows, and Process Steps

> Discover how the /understand-domain command extracts business domains, flows, and process steps using a two-stage pipeline with Python scripts and LLM agents. Unlock deep insights into your code.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-11

---

**The `/understand-domain` command runs a two-stage pipeline that first extracts code context via [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py), then uses an LLM-backed agent to generate a structured domain graph with business domains, flows, and steps.**

The **Egonex-AI/Understand-Anything** repository provides a domain analysis tool that transforms raw source code into a visual representation of business logic. This command operates through a lightweight pipeline that bridges static code analysis with generative AI to extract business domains, flows, and process steps. The resulting **domain graph** enables developers to visualize cross-domain relationships and process flows directly in the dashboard.

## Stage 1: Source Code Context Extraction

The first stage runs the bundled script **[`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)** located at [`understand-anything-plugin/skills/understand-domain/extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand-domain/extract-domain-context.py). This script performs static analysis to build a structured context object that feeds into the LLM.

### Project Traversal and Filtering

The script traverses the source tree while respecting **`.gitignore`** patterns and a built-in skip list (`SKIP_DIRS`). It only processes files matching extensions defined in `SOURCE_EXTENSIONS`, applying configurable limits such as `MAX_FILE_TREE_DEPTH` and `MAX_FILES_TOTAL` to prevent overwhelming the downstream processor.

The output is a `fileTree` array containing relative paths of all qualifying source files.

### Entry Point Detection

The script iterates over each file and applies regular-expression patterns (`ENTRY_POINT_PATTERNS`) to identify critical integration points:

- **HTTP routes** (Express, FastAPI, NestJS, Next.js, GraphQL, gRPC)
- **CLI commands** and argparse sub-parsers
- **Event listeners** and cron-style schedules

For each match, the script records the file path, line number, entry type, description, exact match string, and a surrounding code snippet. These records populate the `entryPoints` array.

### File Signature Harvesting

A secondary pass (`extract_file_signatures`) targets a prioritized subset of files (capped at `MAX_SAMPLED_FILES`). For each file, it extracts:

- **Exports**: JavaScript/TypeScript `export` statements or Python `def`/`class` definitions
- **Imports**: First 20 import statements
- **Line count** and a truncated preview of the file content

This data structures the `fileSignatures` array, providing the LLM with function signatures and dependency graphs without full file contents.

### Metadata Collection

The script harvests project metadata by reading standard configuration files like [`package.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/package.json), [`pyproject.toml`](https://github.com/Egonex-AI/Understand-Anything/blob/main/pyproject.toml), and `README.*`. This `metadata` section provides context about project dependencies, scripts, and human-readable descriptions.

### Context Truncation

Because the downstream LLM agent operates within token budgets, the **`_truncate_to_fit`** function progressively compresses the payload. It limits the file tree depth, truncates code previews, and reduces signature counts until the JSON output is ≤ `MAX_OUTPUT_BYTES` (approximately 512 KB).

The final output is written to [`.understand-anything/intermediate/domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/domain-context.json).

## Stage 2: LLM-Based Domain Graph Generation

The second stage consumes [`domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-context.json) and uses the **domain-analyzer** agent to infer higher-level business concepts.

### Schema Definition

The domain graph structure is defined in **[`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts)**, which declares three node types and three edge types:

```typescript
// Node types
type Node = {
  id: string;                 // e.g. "domain:order-management"
  type: "domain" | "flow" | "step";
  label: string;
  domainMeta?: DomainMeta;    // optional metadata for domain/flow/step
};

// Edge types
type Edge = {
  source: string;
  target: string;
  type: "contains_flow" | "flow_step" | "cross_domain";
};

```

### Domain Analyzer Agent

The domain-analyzer agent (part of the core analyzer suite) reads the context JSON and prompts a language model to:

1. Identify **business domains** (e.g., "Order Management", "Payment Processing")
2. Map **flows** within each domain (e.g., "Order Creation")
3. Extract **process steps** within each flow (e.g., "Validate Order")
4. Establish relationships via edges: `contains_flow` (domain to flow), `flow_step` (flow to step), and `cross_domain` (domain to domain)

The agent validates the graph using `validateGraph` and persists the result to **[`domain-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-graph.json)** via [`packages/core/src/persistence/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/persistence/index.ts) (where `DOMAIN_GRAPH_FILE = "domain-graph.json"`).

## Visualization and Dashboard Integration

When the dashboard loads, it checks for [`domain-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-graph.json). If present, the UI switches to **Domain view** (`viewMode: "domain"`), rendering a horizontal flow graph of domains, flows, and steps. Cross-domain edges visualize dependencies between business areas.

The visualization logic resides in **[`packages/dashboard/src/components/DomainGraphView.tsx`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/dashboard/src/components/DomainGraphView.tsx)**, while the view state is managed in **[`packages/dashboard/src/store.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/dashboard/src/store.ts)**.

## Running the Command

You can execute the extraction pipeline manually or via the Claude skill interface.

### Manual Execution

Run the extraction script directly from the repository root:

```bash
understand-anything-plugin/skills/understand-domain/extract-domain-context.py /path/to/project

```

This creates [`.understand-anything/intermediate/domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/domain-context.json).

### Skill Invocation

Inside a Claude or Claude-Code session, invoke the command:

```bash
/understand-domain

```

According to **[`skills/understand-domain/SKILL.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/skills/understand-domain/SKILL.md)**, this executes:

```bash
python ./extract-domain-context.py "$PROJECT_ROOT"

```

### Sample Output Structure

The intermediate [`domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-context.json) contains structured data like:

```json
{
  "projectRoot": "/path/to/project",
  "fileCount": 124,
  "fileTree": ["src/api/orders.ts", "src/services/payment.py"],
  "entryPoints": [
    {
      "file": "src/api/orders.ts",
      "line": 12,
      "type": "http",
      "description": "Express/Koa route",
      "match": "app.get('/orders')",
      "snippet": "app.get('/orders', (req, res) => { /* … */"
    }
  ],
  "fileSignatures": [
    {
      "file": "src/services/payment.py",
      "exports": ["process_payment", "PaymentGateway"],
      "imports": ["stripe", "logging"],
      "lines": 210,
      "preview": "def process_payment(order):\n    ..."
    }
  ]
}

```

The resulting [`domain-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-graph.json) follows this structure:

```json
{
  "nodes": [
    { "id": "domain:order-management", "type": "domain", "label": "Order Management" },
    { "id": "flow:order-creation", "type": "flow", "label": "Order Creation" },
    { "id": "step:validate-order", "type": "step", "label": "Validate Order" }
  ],
  "edges": [
    { "source": "domain:order-management", "target": "flow:order-creation", "type": "contains_flow" },
    { "source": "flow:order-creation", "target": "step:validate-order", "type": "flow_step" }
  ]
}

```

## Summary

- **[`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)** performs static analysis to build `fileTree`, `entryPoints`, `fileSignatures`, and `metadata` while respecting `.gitignore` and size limits.
- The **`_truncate_to_fit`** function ensures the JSON payload stays within the `MAX_OUTPUT_BYTES` budget (≈512 KB).
- The **domain-analyzer** agent consumes the context JSON and generates a validated domain graph with nodes (`domain`, `flow`, `step`) and edges (`contains_flow`, `flow_step`, `cross_domain`).
- The graph schema is defined in **[`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts)** and persisted to **[`domain-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-graph.json)** via **[`packages/core/src/persistence/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/persistence/index.ts)**.
- The dashboard renders the graph in **Domain view** using **[`DomainGraphView.tsx`](https://github.com/Egonex-AI/Understand-Anything/blob/main/DomainGraphView.tsx)**, displaying business domains as horizontal flows with cross-domain relationships.

## Frequently Asked Questions

### What file types does the `/understand-domain` command analyze?

The command processes files matching extensions defined in the `SOURCE_EXTENSIONS` constant within [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py). It typically includes JavaScript, TypeScript, Python, and other source files while respecting `.gitignore` patterns and the internal `SKIP_DIRS` list to exclude directories like `node_modules` and `__pycache__`.

### How does the command identify business entry points?

The script applies `ENTRY_POINT_PATTERNS` regular expressions to detect HTTP routes (Express, FastAPI, NestJS), CLI commands, event listeners, and cron schedules. Each match records the file path, line number, entry type, and surrounding code snippet, enabling the LLM to understand the application's external interfaces.

### Where is the domain graph stored and how is it validated?

The generated domain graph is written to [`domain-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-graph.json) in the project root, as defined by `DOMAIN_GRAPH_FILE` in [`packages/core/src/persistence/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/persistence/index.ts). The **domain-analyzer** agent validates the graph structure against the TypeScript schema in [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts) before persistence, ensuring all nodes have valid types (`domain`, `flow`, `step`) and edges reference existing node IDs.

### Can I run the extraction without the full dashboard?

Yes. You can execute [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) manually against any project directory to generate [`.understand-anything/intermediate/domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/domain-context.json). This JSON file contains all extracted context and can be consumed independently or used to debug the LLM generation step without loading the dashboard interface.