# How the Domain-Analyzer Extracts Business Domains, Flows, and Process Steps from Codebases

> Discover how the Domain-Analyzer extracts business domains, flows, and process steps from codebases using LLM-powered semantic code analysis for deeper insights into your software architecture.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-18

---

**The Domain-Analyzer in Understand Anything uses an LLM to enrich structural code graphs with business semantics, identifying high-level domains, end-to-end flows, and granular process steps through schema-aware node extraction and normalization.**

The **Domain-Analyzer** is a specialized agent in the **Egonex-AI/Understand-Anything** multi-agent pipeline that transforms syntactic code structures into actionable business-domain knowledge. When you invoke the `/understand-domain` command, this agent consumes the structural knowledge graph produced by the file-analyzer and applies large language model (LLM) reasoning to surface architectural concepts hidden in your implementation.

## Input: Consuming the Structural Graph

The Domain-Analyzer begins its work after the initial scanning phase completes. The scanner discovers all source files, and the file-analyzer extracts syntactic facts—functions, classes, imports, and inheritance relationships—storing them in [`.understand-anything/knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json). This **structural graph** contains only technical edges such as `reads_from`, `calls`, and `inherits`, lacking any business context.

According to the README's "Multi-Agent Pipeline" section (lines 311-321), the Domain-Analyzer receives this graph as its primary input, preparing to enrich it with semantic meaning.

## LLM-Driven Semantic Extraction

The core intelligence of the Domain-Analyzer relies on **prompt engineering** and LLM inference. The analyzer feeds the model two critical inputs:

1. **Raw source code** from each relevant file or snippet
2. **Structural context**, including import maps, call graphs, and file hierarchy

In [`understand-anything-plugin/src/onboard-builder.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/src/onboard-builder.ts) (lines 51-52), the prompt explicitly instructs the LLM to surface "important architectural and domain concepts." The model identifies:

- **Business domains** – High-level product areas (e.g., *order-management*, *payment-processing*)
- **Flows** – End-to-end processes spanning multiple files (e.g., *order-creation flow*)
- **Steps** – Granular actions within flows (e.g., *validate-order*, *charge-card*)

## Domain-Specific Schema Types

Once the LLM identifies business concepts, the analyzer creates strongly-typed nodes and edges according to a strict schema defined in [`understand-anything-plugin/packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/schema.ts).

The schema introduces three new node types (lines 12-13):

- `domain` – Represents high-level business areas
- `flow` – Represents end-to-end business processes  
- `step` – Represents individual actions within flows

And three relationship edge types (lines 54-55):

- `contains_flow` – Links domains to their constituent flows
- `flow_step` – Connects flows to their sequential steps
- `cross_domain` – Identifies dependencies between different business domains

Each node stores additional metadata in the `domainMeta` field, capturing human-readable descriptions, complexity metrics, and classification tags.

## Normalization and Validation

Before persisting results, the **graph-builder** normalizes identifiers to ensure consistency across the knowledge graph. In [`understand-anything-plugin/packages/core/src/analyzer/normalize-graph.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/normalize-graph.ts) (lines 5-22 and 181-183), the system:

- Generates canonical IDs such as `domain:order-management`, `flow:create-order`, and `step:create-order:validate`
- Rewrites edge IDs to correctly wire `flow_step` relationships
- Ensures parent-child relationships between domains, flows, and steps maintain referential integrity

The `validateGraph` function in [`schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/schema.ts) then verifies that the domain graph respects type constraints and that `domainMeta` fields are preserved. Unit tests in [`understand-anything-plugin/packages/core/src/__tests__/domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/__tests__/domain-types.test.ts) (lines 68-79) enforce these validation rules, ensuring that domain-specific node types and edge aliases function correctly.

## Persistence and Visualization

The final **domain graph** is written to [`domain-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-graph.json), defined by the constant `DOMAIN_GRAPH_FILE` in [`understand-anything-plugin/packages/core/src/persistence/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/persistence/index.ts) (lines 150-176). This file separates business-domain knowledge from the structural graph, allowing independent analysis and versioning.

The Understand Anything dashboard reads this file when switching to *Domain* view. In [`understand-anything-plugin/packages/dashboard/src/store.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/dashboard/src/store.ts) (lines 14-18 and 34-38), the application registers the domain node and edge types, rendering domains, flows, and steps as a horizontal graph visualization with node type "domain" and edge category "domain."

## Practical Usage Examples

### CLI Invocation

Run the full analysis pipeline including domain extraction:

```bash
/understand

```

Or execute only the domain analysis step:

```bash
/understand-domain

```

### Programmatic Access (Node.js)

```typescript
import { analyzeProject } from '@understand-anything/core';
import { DOMAIN_GRAPH_FILE } from '@understand-anything/core/persistence';
import { readFile } from 'node:fs/promises';

// Execute the full pipeline including domain-analyzer
await analyzeProject({ root: '/path/to/project' });

// Load the generated domain graph
const domainGraph = JSON.parse(
  await readFile(`${process.cwd()}/.understand-anything/${DOMAIN_GRAPH_FILE}`, 'utf-8')
);

// List all detected business domains
const domains = domainGraph.nodes.filter(n => n.type === 'domain');
console.log('Business domains:', domains.map(d => d.name));

```

### Traversing Flows and Steps

```typescript
// Locate a specific flow
const flow = domainGraph.nodes.find(
  n => n.type === 'flow' && n.name === 'Create Order'
);

// Extract steps via flow_step edges
const stepEdges = domainGraph.edges.filter(
  e => e.type === 'flow_step' && e.source === flow.id
);

const steps = stepEdges.map(e => 
  domainGraph.nodes.find(n => n.id === e.target)
);

console.log('Process steps:', steps.map(s => s.name));

```

## Summary

- The Domain-Analyzer operates as a pipeline agent in **Understand Anything**, processing structural graphs into business-domain knowledge.
- It uses **LLM prompt engineering** (defined in [`onboard-builder.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/onboard-builder.ts)) to identify domains, flows, and steps from raw source code.
- The system employs **schema-aware types** (`domain`, `flow`, `step`) with specific edge relationships (`contains_flow`, `flow_step`, `cross_domain`).
- **ID normalization** in [`normalize-graph.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/normalize-graph.ts) ensures consistent entity naming (e.g., `domain:order-management`).
- Results persist to [`domain-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-graph.json) and render in the dashboard's Domain view, providing horizontal visualization of business architecture.

## Frequently Asked Questions

### How does the Domain-Analyzer handle large codebases?

The Domain-Analyzer processes the structural graph incrementally, feeding the LLM relevant code snippets rather than entire repositories at once. The analyzer respects the existing file boundaries discovered by the scanner, ensuring that business flows spanning multiple files are correctly linked through the `cross_domain` edge type while maintaining manageable token limits for the LLM context window.

### Can I customize the business domains the analyzer identifies?

Yes, the domain extraction relies on prompt engineering in [`understand-anything-plugin/src/onboard-builder.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/src/onboard-builder.ts). You can modify the prompt instructions (lines 51-52) to guide the LLM toward specific domain vocabularies or industry terminology relevant to your codebase, though the core schema types (`domain`, `flow`, `step`) remain standardized across the platform.

### What is the relationship between the structural graph and the domain graph?

The structural graph ([`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json)) contains only syntactic relationships—functions calling functions, classes importing modules. The Domain-Analyzer consumes this as input and produces [`domain-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-graph.json), which contains semantic business relationships. The domain graph references structural nodes but operates at a higher abstraction level, making it possible to analyze business processes without losing the ability to trace them down to specific implementation files.

### How does the system validate the extracted domain concepts?

The `validateGraph` function in [`schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/schema.ts) enforces type constraints on all domain nodes, ensuring that `domainMeta` fields are present and that edge relationships (`contains_flow`, `flow_step`) connect valid, existing nodes. Unit tests in [`__tests__/domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/__tests__/domain-types.test.ts) verify that the normalization process correctly generates IDs and maintains referential integrity between domains, flows, and steps.