How the Domain-Analyzer Extracts Business Domains, Flows, and Process Steps from Codebases
The Domain-Analyzer in Understand Anything uses an LLM to enrich structural code graphs with business semantics, identifying high-level domains, end-to-end flows, and granular process steps through schema-aware node extraction and normalization.
The Domain-Analyzer is a specialized agent in the Egonex-AI/Understand-Anything multi-agent pipeline that transforms syntactic code structures into actionable business-domain knowledge. When you invoke the /understand-domain command, this agent consumes the structural knowledge graph produced by the file-analyzer and applies large language model (LLM) reasoning to surface architectural concepts hidden in your implementation.
Input: Consuming the Structural Graph
The Domain-Analyzer begins its work after the initial scanning phase completes. The scanner discovers all source files, and the file-analyzer extracts syntactic facts—functions, classes, imports, and inheritance relationships—storing them in .understand-anything/knowledge-graph.json. This structural graph contains only technical edges such as reads_from, calls, and inherits, lacking any business context.
According to the README's "Multi-Agent Pipeline" section (lines 311-321), the Domain-Analyzer receives this graph as its primary input, preparing to enrich it with semantic meaning.
LLM-Driven Semantic Extraction
The core intelligence of the Domain-Analyzer relies on prompt engineering and LLM inference. The analyzer feeds the model two critical inputs:
- Raw source code from each relevant file or snippet
- Structural context, including import maps, call graphs, and file hierarchy
In understand-anything-plugin/src/onboard-builder.ts (lines 51-52), the prompt explicitly instructs the LLM to surface "important architectural and domain concepts." The model identifies:
- Business domains – High-level product areas (e.g., order-management, payment-processing)
- Flows – End-to-end processes spanning multiple files (e.g., order-creation flow)
- Steps – Granular actions within flows (e.g., validate-order, charge-card)
Domain-Specific Schema Types
Once the LLM identifies business concepts, the analyzer creates strongly-typed nodes and edges according to a strict schema defined in understand-anything-plugin/packages/core/src/schema.ts.
The schema introduces three new node types (lines 12-13):
domain– Represents high-level business areasflow– Represents end-to-end business processesstep– Represents individual actions within flows
And three relationship edge types (lines 54-55):
contains_flow– Links domains to their constituent flowsflow_step– Connects flows to their sequential stepscross_domain– Identifies dependencies between different business domains
Each node stores additional metadata in the domainMeta field, capturing human-readable descriptions, complexity metrics, and classification tags.
Normalization and Validation
Before persisting results, the graph-builder normalizes identifiers to ensure consistency across the knowledge graph. In understand-anything-plugin/packages/core/src/analyzer/normalize-graph.ts (lines 5-22 and 181-183), the system:
- Generates canonical IDs such as
domain:order-management,flow:create-order, andstep:create-order:validate - Rewrites edge IDs to correctly wire
flow_steprelationships - Ensures parent-child relationships between domains, flows, and steps maintain referential integrity
The validateGraph function in schema.ts then verifies that the domain graph respects type constraints and that domainMeta fields are preserved. Unit tests in understand-anything-plugin/packages/core/src/__tests__/domain-types.test.ts (lines 68-79) enforce these validation rules, ensuring that domain-specific node types and edge aliases function correctly.
Persistence and Visualization
The final domain graph is written to domain-graph.json, defined by the constant DOMAIN_GRAPH_FILE in understand-anything-plugin/packages/core/src/persistence/index.ts (lines 150-176). This file separates business-domain knowledge from the structural graph, allowing independent analysis and versioning.
The Understand Anything dashboard reads this file when switching to Domain view. In understand-anything-plugin/packages/dashboard/src/store.ts (lines 14-18 and 34-38), the application registers the domain node and edge types, rendering domains, flows, and steps as a horizontal graph visualization with node type "domain" and edge category "domain."
Practical Usage Examples
CLI Invocation
Run the full analysis pipeline including domain extraction:
/understand
Or execute only the domain analysis step:
/understand-domain
Programmatic Access (Node.js)
import { analyzeProject } from '@understand-anything/core';
import { DOMAIN_GRAPH_FILE } from '@understand-anything/core/persistence';
import { readFile } from 'node:fs/promises';
// Execute the full pipeline including domain-analyzer
await analyzeProject({ root: '/path/to/project' });
// Load the generated domain graph
const domainGraph = JSON.parse(
await readFile(`${process.cwd()}/.understand-anything/${DOMAIN_GRAPH_FILE}`, 'utf-8')
);
// List all detected business domains
const domains = domainGraph.nodes.filter(n => n.type === 'domain');
console.log('Business domains:', domains.map(d => d.name));
Traversing Flows and Steps
// Locate a specific flow
const flow = domainGraph.nodes.find(
n => n.type === 'flow' && n.name === 'Create Order'
);
// Extract steps via flow_step edges
const stepEdges = domainGraph.edges.filter(
e => e.type === 'flow_step' && e.source === flow.id
);
const steps = stepEdges.map(e =>
domainGraph.nodes.find(n => n.id === e.target)
);
console.log('Process steps:', steps.map(s => s.name));
Summary
- The Domain-Analyzer operates as a pipeline agent in Understand Anything, processing structural graphs into business-domain knowledge.
- It uses LLM prompt engineering (defined in
onboard-builder.ts) to identify domains, flows, and steps from raw source code. - The system employs schema-aware types (
domain,flow,step) with specific edge relationships (contains_flow,flow_step,cross_domain). - ID normalization in
normalize-graph.tsensures consistent entity naming (e.g.,domain:order-management). - Results persist to
domain-graph.jsonand render in the dashboard's Domain view, providing horizontal visualization of business architecture.
Frequently Asked Questions
How does the Domain-Analyzer handle large codebases?
The Domain-Analyzer processes the structural graph incrementally, feeding the LLM relevant code snippets rather than entire repositories at once. The analyzer respects the existing file boundaries discovered by the scanner, ensuring that business flows spanning multiple files are correctly linked through the cross_domain edge type while maintaining manageable token limits for the LLM context window.
Can I customize the business domains the analyzer identifies?
Yes, the domain extraction relies on prompt engineering in understand-anything-plugin/src/onboard-builder.ts. You can modify the prompt instructions (lines 51-52) to guide the LLM toward specific domain vocabularies or industry terminology relevant to your codebase, though the core schema types (domain, flow, step) remain standardized across the platform.
What is the relationship between the structural graph and the domain graph?
The structural graph (knowledge-graph.json) contains only syntactic relationships—functions calling functions, classes importing modules. The Domain-Analyzer consumes this as input and produces domain-graph.json, which contains semantic business relationships. The domain graph references structural nodes but operates at a higher abstraction level, making it possible to analyze business processes without losing the ability to trace them down to specific implementation files.
How does the system validate the extracted domain concepts?
The validateGraph function in schema.ts enforces type constraints on all domain nodes, ensuring that domainMeta fields are present and that edge relationships (contains_flow, flow_step) connect valid, existing nodes. Unit tests in __tests__/domain-types.test.ts verify that the normalization process correctly generates IDs and maintains referential integrity between domains, flows, and steps.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →