# How the Domain-Analyzer Extracts Business Domains, Flows, and Process Steps from Code

> Discover how domain-analyzer extracts business domains, flows, and process steps from code by transforming source code into a knowledge graph with a large language model.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-06

---

**The domain-analyzer transforms raw source code into a structured knowledge graph by prompting a large language model to identify business domains, their internal flows, and granular process steps, then normalizing the response into typed nodes and edges before persisting the validated graph.**

The domain-analyzer is a core stage of the Analyze pipeline in the Understand-Anything repository that bridges the gap between implementation details and business architecture. As implemented in `Lum1104/Understand-Anything`, it automatically surfaces domain boundaries and process workflows by analyzing source code semantics, enabling developers to explore system logic through business-level abstractions rather than file structures.

## LLM-Driven Extraction of Business Concepts

In [`packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/llm-analyzer.ts), the `buildProjectSummaryPrompt` function constructs a prompt that explicitly asks the LLM to summarize the project and extract three distinct **business-level concepts**: domains, flows, and steps. The prompt requests the model to identify high-level business areas (domains), the processes within each domain (flows), and the granular operations composing each flow (steps).

The LLM returns a structured JSON payload that follows this hierarchical pattern:

```json
{
  "domains": [
    {
      "name": "order-management",
      "description": "...",
      "flows": [
        {
          "name": "order-creation",
          "steps": ["validate-order", "persist-order", "notify-customer"]
        }
      ]
    }
  ]
}

```

## Parsing and Normalizing the LLM Response

The raw LLM output undergoes strict validation through `parseProjectSummaryResponse` in the same [`llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/llm-analyzer.ts) file. This function safely ignores malformed entries or missing values, guaranteeing that only well-typed structures proceed to graph construction.

To handle semantic variations in the LLM output—such as synonyms like `business_domain` instead of the canonical `domain`—the pipeline invokes `normaliseGraph` from [`packages/core/src/analyzer/normalize-graph.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/normalize-graph.ts). This utility maps field aliases and type variants to standardized node types, ensuring consistency before the graph is built.

## Constructing the Domain Knowledge Graph

The `buildDomainGraph` function in [`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts) iterates over the normalized data to instantiate three specialized **node types**:

- **domain** – High-level business areas (e.g., "order-management")
- **flow** – Business processes within a specific domain (e.g., "order-creation")
- **step** – Individual operations that make up a flow (e.g., "validate-order")

Each node receives a `domainMeta` field defined in [`packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/types.ts), which stores the LLM-provided description, owner information, and architectural annotations.

The builder then establishes relationships through three **edge categories**:

- `contains_flow` – Links Domain → Flow
- `flow_step` – Links Flow → Step
- `cross_domain` – Links Step → Step when a step invokes logic residing in a different domain

## Validation and Persistence

Before serialization, the composed graph undergoes strict validation against the JSON schema defined in [`packages/core/src/schema.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/schema.ts). Invalid graphs raise descriptive errors, preventing corrupted data from reaching the storage layer.

Once validated, the `validateAndWriteDomainGraph` function in [`packages/core/src/persistence/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/persistence/index.ts) serializes the structure to [`domain-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/domain-graph.json). This file serves as the single source of truth for the dashboard, which consumes the graph to render the domain view (`viewMode: "domain"`), treating domain, flow, and step nodes as first-class citizens for business-level navigation.

## Implementation Example

The following TypeScript workflow demonstrates the complete pipeline from prompt generation to persistence:

```typescript
// 1️⃣ Build the LLM prompt (inside llm-analyzer.ts)
const prompt = buildProjectSummaryPrompt(
  allFilePaths,
  sampleFiles // a few representative source files
);

// 2️⃣ Call the LLM (handled by the surrounding agent) and parse the reply
const llmResponse = await callLlm(prompt);
const projectSummary = parseProjectSummaryResponse(llmResponse);
if (!projectSummary) throw new Error('LLM gave malformed project summary');

// 3️⃣ Normalise any alias (inside normalize-graph.ts)
const normalised = normaliseGraph(projectSummary);

// 4️⃣ Build the domain graph (inside graph-builder.ts)
const domainGraph = buildDomainGraph(normalised);

// 5️⃣ Validate and persist (inside persistence/index.ts)
await validateAndWriteDomainGraph(domainGraph);

```

## Summary

- The domain-analyzer uses **LLM-driven extraction** via `buildProjectSummaryPrompt` to surface business concepts (domains, flows, steps) directly from source code analysis.
- **Response parsing** and **type normalization** in [`normalize-graph.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/normalize-graph.ts) ensure consistent node typing regardless of LLM output variations or synonym usage.
- The **knowledge graph** construction in [`graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-builder.ts) creates three node types connected by semantic edges (`contains_flow`, `flow_step`, `cross_domain`) that map business architecture.
- **Schema validation** in [`schema.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/schema.ts) and persistence logic in [`persistence/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/persistence/index.ts) guarantee data integrity before writing to [`domain-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/domain-graph.json).
- The resulting graph enables business-level navigation of codebases through the Understand-Anything dashboard without manual documentation.

## Frequently Asked Questions

### What happens if the LLM returns malformed JSON?

The `parseProjectSummaryResponse` function in [`packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/llm-analyzer.ts) validates the LLM response and safely ignores any malformed or missing values. This defensive parsing ensures that only well-typed structures proceed to the graph construction phase, preventing pipeline failures from invalid LLM outputs.

### How does the analyzer distinguish between flows and steps?

In the knowledge graph constructed by `buildDomainGraph`, **flows** represent high-level business processes belonging to a specific domain, while **steps** represent the granular operations that constitute a flow. The hierarchical relationship is enforced through the `contains_flow` (Domain → Flow) and `flow_step` (Flow → Step) edge types, creating a clear parent-child lineage from domain boundaries down to individual processing tasks.

### What are cross-domain edges and when are they created?

**Cross-domain** edges connect steps that invoke logic residing in different business domains. During graph construction, if a step references functionality outside its parent domain, the [`graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-builder.ts) logic creates a `cross_domain` edge from the calling step to the target step. This reveals hidden dependencies between seemingly independent business areas that traditional static code analysis might miss.

### Where is the domain graph stored and how is it consumed?

The validated domain graph is serialized to [`domain-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/domain-graph.json) by the persistence layer in [`packages/core/src/persistence/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/persistence/index.ts). The Understand-Anything dashboard reads this file through its store layer ([`packages/dashboard/src/store.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/dashboard/src/store.ts)) and renders the domain view using `viewMode: "domain"`, allowing users to explore business domains, drill into flows, and inspect individual process steps interactively.