How the Domain Analyzer Extracts Business Domains, Flows, and Process Steps in Understand-Anything
The Domain Analyzer in Egonex-AI/Understand-Anything uses a two-stage pipeline—first scanning source code with extract-domain-context.py to detect entry points and signatures, then constructing a hierarchical domain graph that maps business domains to flows and individual process steps.
The Domain Analyzer is a specialized agent within the Understand-Anything repository that transforms raw source code into a structured domain graph. This graph represents business domains, their associated flows, and the granular process steps that implement them, enabling developers to visualize and navigate complex codebases through a business-logic lens.
Two-Stage Architecture of the Domain Analyzer
The extraction process operates in two complementary stages: an initial lightweight context generation followed by intelligent graph construction.
Stage 1: Context Generation with extract-domain-context.py
The Python scanner (extract-domain-context.py) performs a lightweight walk of the project tree when no existing knowledge graph is present. This stage respects .gitignore patterns and skips common artifacts like node_modules and .git directories to focus on relevant source files.
The scanner enforces a 512 KB limit on collected data to stay within LLM context windows. It truncates file trees, previews, and code snippets intelligently to capture essential metadata without overwhelming downstream processors.
Entry-point detection uses targeted regular expressions to identify HTTP routes, CLI commands, event listeners, cron schedules, GraphQL resolvers, gRPC services, and generic exported handlers. Each match records the file path, line number, entry type (http, cli, event, cron, or manual), and a representative code snippet. This detection logic spans lines 70-115 in the source file.
Stage 2: Domain Graph Construction
The Domain Analyzer agent (defined in domain-analyzer.md) receives either the generated domain-context.json or an existing knowledge-graph.json. It traverses file signatures—exports and imports—to discover logical boundaries such as controllers, services, handlers, and use-cases. These boundaries map to high-level business domains.
Each detected entry point becomes a flow node. The entry-point type populates the entryType field, while the specific route or command string becomes the entryPoint identifier.
How Entry Points Map to Business Flows
The analyzer treats entry points as the genesis of business flows. When the agent processes the context, it:
- Identifies flow boundaries by examining imports and exports that suggest architectural layers (e.g., a controller importing a service).
- Creates flow nodes linked to their parent domain via
contains_flowedges. - Populates metadata including the triggering mechanism (HTTP endpoint, scheduled job, etc.) and the implementing file location.
This mapping allows the dashboard to present a domain-first view where users can explore how external triggers cascade through the business logic.
Extracting Process Steps and Domain Metadata
The analyzer examines the code surrounding each entry point to infer process steps—individual actions such as validation, database writes, or external API calls. Step nodes inherit precise file paths and line ranges from the source, enabling accurate traceability from the graph back to the implementation.
For each node, the analyzer populates a domain-meta object containing:
- Entities: Business objects discovered through naming conventions
- Business rules: Logic inferred from code structure and comments
- Cross-domain interactions: Dependencies identified through import relationships or explicit cross-module calls
The agent emits three distinct edge types:
contains_flow: Links domain nodes to their constituent flowsflow_step: Orders steps within a flow using monotonically increasingweightvalues that encode execution sequencecross_domain: Describes interactions between separate business domains
Output Schema and Validation
The resulting domain graph conforms to the TypeScript schema defined in types.ts, where nodes are typed as "domain", "flow", or "step", and edges are strictly typed as "contains_flow", "flow_step", or "cross_domain".
Validation tests in domain-types.test.ts ensure that generated graphs comply with the schema and that edge aliases are properly normalized, preventing structural inconsistencies that could break the visualization layer.
Practical Usage Workflow
Integrate the Domain Analyzer into your development workflow using the following commands:
# Stage 1: Generate the lightweight context (if no knowledge graph exists)
python extract-domain-context.py /path/to/project
# Stage 2: Run the Domain Analyzer agent
understand --skill domain-analyzer /path/to/project
The analyzer writes intermediate outputs to:
<project-root>/.understand-anything/intermediate/domain-context.json
<project-root>/.understand-anything/intermediate/domain-analysis.json
Loading Results Programmatically
Access the generated domain graph in TypeScript applications:
import type { KnowledgeGraph } from '@understand-anything/core';
import { readFile } from 'node:fs/promises';
async function loadDomainGraph(root: string): Promise<KnowledgeGraph> {
const path = `${root}/.understand-anything/intermediate/domain-analysis.json`;
const raw = await readFile(path, 'utf-8');
return JSON.parse(raw) as KnowledgeGraph;
}
Or invoke the Python scanner programmatically:
from pathlib import Path
import subprocess, json
project = Path("/my/project")
subprocess.run(["python", "extract-domain-context.py", str(project)], check=True)
# Load the context for custom processing
ctx_path = project / ".understand-anything" / "intermediate" / "domain-context.json"
with ctx_path.open() as f:
context = json.load(f)
print(f"Detected {len(context['entryPoints'])} entry points")
Summary
- The Domain Analyzer operates in two stages: Python-based context generation (
extract-domain-context.py) and agent-based graph construction (domain-analyzer.md). - Entry-point detection uses regex patterns to identify HTTP routes, CLI commands, events, cron jobs, and other triggers across lines 70-115 of the scanner.
- The output is a hierarchical domain graph with three node types (domain, flow, step) and three edge types (
contains_flow,flow_step,cross_domain). - Process steps inherit file paths and line ranges from source code, while
weightvalues on edges encode execution order. - All outputs follow the strict schema in
types.tsand are validated bydomain-types.test.ts.
Frequently Asked Questions
What file types does the Domain Analyzer support?
The analyzer language-agnostic regex patterns in extract-domain-context.py detect entry points across common backend languages including JavaScript, TypeScript, Python, Go, and Java. The scanner respects .gitignore automatically, skipping binary files and dependency directories like node_modules to focus on implementation code.
How does the analyzer handle large codebases?
The extract-domain-context.py scanner limits data collection to 512 KB per project to respect LLM context windows. It intelligently truncates file trees and code previews while preserving critical metadata such as entry-point locations and signature exports, ensuring the Domain Analyzer can process monorepos and large legacy systems without overwhelming the graph construction phase.
What is the difference between domain-context.json and domain-analysis.json?
The domain-context.json file is the raw output of the Python scanner containing file trees, entry points, and code snippets. The domain-analysis.json file is the processed output of the Domain Analyzer agent containing the complete hierarchical graph with business domains, flows, process steps, and relationship edges. The former is input; the latter is the final navigable structure.
Can I extend the entry-point detection patterns?
Yes. The entry-point detection logic in extract-domain-context.py (lines 70-115) uses configurable regular expressions. You can modify these patterns to recognize custom framework conventions, internal APIs, or domain-specific triggers before running the scanner. The agent will then map these custom entry points to flows using the same graph construction logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →