How the Domain-Analyzer Extracts Business Domains, Flows, and Process Steps from Code
The domain-analyzer transforms raw source code into a structured knowledge graph by prompting a large language model to identify business domains, their internal flows, and granular process steps, then normalizing the response into typed nodes and edges before persisting the validated graph.
The domain-analyzer is a core stage of the Analyze pipeline in the Understand-Anything repository that bridges the gap between implementation details and business architecture. As implemented in Lum1104/Understand-Anything, it automatically surfaces domain boundaries and process workflows by analyzing source code semantics, enabling developers to explore system logic through business-level abstractions rather than file structures.
LLM-Driven Extraction of Business Concepts
In packages/core/src/analyzer/llm-analyzer.ts, the buildProjectSummaryPrompt function constructs a prompt that explicitly asks the LLM to summarize the project and extract three distinct business-level concepts: domains, flows, and steps. The prompt requests the model to identify high-level business areas (domains), the processes within each domain (flows), and the granular operations composing each flow (steps).
The LLM returns a structured JSON payload that follows this hierarchical pattern:
{
"domains": [
{
"name": "order-management",
"description": "...",
"flows": [
{
"name": "order-creation",
"steps": ["validate-order", "persist-order", "notify-customer"]
}
]
}
]
}
Parsing and Normalizing the LLM Response
The raw LLM output undergoes strict validation through parseProjectSummaryResponse in the same llm-analyzer.ts file. This function safely ignores malformed entries or missing values, guaranteeing that only well-typed structures proceed to graph construction.
To handle semantic variations in the LLM output—such as synonyms like business_domain instead of the canonical domain—the pipeline invokes normaliseGraph from packages/core/src/analyzer/normalize-graph.ts. This utility maps field aliases and type variants to standardized node types, ensuring consistency before the graph is built.
Constructing the Domain Knowledge Graph
The buildDomainGraph function in packages/core/src/analyzer/graph-builder.ts iterates over the normalized data to instantiate three specialized node types:
- domain – High-level business areas (e.g., "order-management")
- flow – Business processes within a specific domain (e.g., "order-creation")
- step – Individual operations that make up a flow (e.g., "validate-order")
Each node receives a domainMeta field defined in packages/core/src/types.ts, which stores the LLM-provided description, owner information, and architectural annotations.
The builder then establishes relationships through three edge categories:
contains_flow– Links Domain → Flowflow_step– Links Flow → Stepcross_domain– Links Step → Step when a step invokes logic residing in a different domain
Validation and Persistence
Before serialization, the composed graph undergoes strict validation against the JSON schema defined in packages/core/src/schema.ts. Invalid graphs raise descriptive errors, preventing corrupted data from reaching the storage layer.
Once validated, the validateAndWriteDomainGraph function in packages/core/src/persistence/index.ts serializes the structure to domain-graph.json. This file serves as the single source of truth for the dashboard, which consumes the graph to render the domain view (viewMode: "domain"), treating domain, flow, and step nodes as first-class citizens for business-level navigation.
Implementation Example
The following TypeScript workflow demonstrates the complete pipeline from prompt generation to persistence:
// 1️⃣ Build the LLM prompt (inside llm-analyzer.ts)
const prompt = buildProjectSummaryPrompt(
allFilePaths,
sampleFiles // a few representative source files
);
// 2️⃣ Call the LLM (handled by the surrounding agent) and parse the reply
const llmResponse = await callLlm(prompt);
const projectSummary = parseProjectSummaryResponse(llmResponse);
if (!projectSummary) throw new Error('LLM gave malformed project summary');
// 3️⃣ Normalise any alias (inside normalize-graph.ts)
const normalised = normaliseGraph(projectSummary);
// 4️⃣ Build the domain graph (inside graph-builder.ts)
const domainGraph = buildDomainGraph(normalised);
// 5️⃣ Validate and persist (inside persistence/index.ts)
await validateAndWriteDomainGraph(domainGraph);
Summary
- The domain-analyzer uses LLM-driven extraction via
buildProjectSummaryPromptto surface business concepts (domains, flows, steps) directly from source code analysis. - Response parsing and type normalization in
normalize-graph.tsensure consistent node typing regardless of LLM output variations or synonym usage. - The knowledge graph construction in
graph-builder.tscreates three node types connected by semantic edges (contains_flow,flow_step,cross_domain) that map business architecture. - Schema validation in
schema.tsand persistence logic inpersistence/index.tsguarantee data integrity before writing todomain-graph.json. - The resulting graph enables business-level navigation of codebases through the Understand-Anything dashboard without manual documentation.
Frequently Asked Questions
What happens if the LLM returns malformed JSON?
The parseProjectSummaryResponse function in packages/core/src/analyzer/llm-analyzer.ts validates the LLM response and safely ignores any malformed or missing values. This defensive parsing ensures that only well-typed structures proceed to the graph construction phase, preventing pipeline failures from invalid LLM outputs.
How does the analyzer distinguish between flows and steps?
In the knowledge graph constructed by buildDomainGraph, flows represent high-level business processes belonging to a specific domain, while steps represent the granular operations that constitute a flow. The hierarchical relationship is enforced through the contains_flow (Domain → Flow) and flow_step (Flow → Step) edge types, creating a clear parent-child lineage from domain boundaries down to individual processing tasks.
What are cross-domain edges and when are they created?
Cross-domain edges connect steps that invoke logic residing in different business domains. During graph construction, if a step references functionality outside its parent domain, the graph-builder.ts logic creates a cross_domain edge from the calling step to the target step. This reveals hidden dependencies between seemingly independent business areas that traditional static code analysis might miss.
Where is the domain graph stored and how is it consumed?
The validated domain graph is serialized to domain-graph.json by the persistence layer in packages/core/src/persistence/index.ts. The Understand-Anything dashboard reads this file through its store layer (packages/dashboard/src/store.ts) and renders the domain view using viewMode: "domain", allowing users to explore business domains, drill into flows, and inspect individual process steps interactively.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →