How the Domain Analyzer Extracts Business Domains and Flows in Understand-Anything
The domain analyzer extracts business domains and flows through a two-stage pipeline that first scans source code to build a structured domain-context.json file, then uses a specialized LLM agent to synthesize a hierarchical domain graph mapping business concepts to specific code locations.
The Understand-Anything repository by Egonex-AI automates architectural discovery through a domain analyzer that transforms raw source code into actionable business knowledge. Unlike static documentation, this system programmatically identifies how technical implementations map to business capabilities by combining static analysis with large language model reasoning. The analyzer outputs a machine-readable graph connecting high-level domains to concrete code entry points and execution steps.
Stage 1: Lightweight Code Scanning and Context Generation
The extraction process begins with extract-domain-context.py, a Python scanner located at understand-anything-plugin/skills/understand-domain/extract-domain-context.py. This script performs a high-performance walk of the project tree while respecting .gitignore rules and predefined SKIP_DIRS to avoid irrelevant directories like node_modules or .git (lines 22-48).
Filtering Source Files with .gitignore and Skip Directories
The scanner initializes by filtering the file system against exclusion criteria. It only processes files matching extensions defined in SOURCE_EXTENSIONS and explicitly ignores paths listed in SKIP_DIRS (lines 22-48). This ensures the subsequent analysis focuses on relevant source code while maintaining performance on large repositories.
Detecting Entry Points Across Protocols
During traversal, the scanner identifies integration surfaces using regex patterns defined in ENTRY_POINT_PATTERNS (lines 68-115). These patterns detect:
- HTTP routes and API endpoints
- CLI commands and argument parsers
- Event listeners and message queue handlers
- Cron job definitions and scheduled tasks
- GraphQL resolvers and gRPC service definitions
Each match records the file path, line number, entry type, a code snippet, and a generated description (lines 68-115). These entry points serve as the anchors for later business flow discovery, revealing where external actors trigger system behavior.
Extracting File Signatures and Metadata
For files likely containing business logic—identified through a "priority keywords" filter—the scanner generates lightweight file signatures. Each signature includes exported functions, imported dependencies, line counts, and a short preview of the content (lines 48-56, 67-99). Additionally, the scanner extracts metadata from standard project files including package.json, pyproject.toml, and README.md to capture project context and dependencies (lines 59-69).
Packaging the domain-context.json Output
All collected data—the file tree, entry points, signatures, and metadata—is compiled into domain-context.json. The scanner applies size constraints via MAX_OUTPUT_BYTES to ensure the output remains within LLM context limits, truncating content if necessary while preserving structural integrity (lines 43-73).
Stage 2: LLM-Driven Domain Graph Synthesis
Once the context file is generated, the Domain Analyzer agent defined in understand-anything-plugin/agents/domain-analyzer.md processes the structured data. This agent can ingest either the freshly generated domain-context.json (Option A) or an existing knowledge-graph.json (Option B) for iterative refinement (lines 13-20).
Consuming Context with the Domain Analyzer Agent
The agent receives the JSON context containing the file tree, entry points, and signatures. Rather than reading raw source code directly, it reasons over the pre-structured metadata to identify architectural boundaries. This approach reduces token consumption while maintaining precision in business concept extraction.
Inferring Business Domains, Flows, and Steps
Using the provided context, the LLM agent infers three hierarchical levels of business knowledge:
- Business Domains – High-level capability areas such as "Order Management" or "User Authentication"
- Business Flows – Concrete processes like "Create Order" or "Process Refund" derived from entry-point types (HTTP routes, CLI commands, etc.)
- Business Steps – Individual actions within each flow, mapped to specific source files and line ranges
The agent correlates entry points detected in the scanning phase with business processes, creating traceability between technical implementation and business capability.
Enforcing Schema Consistency and Validation
The agent outputs a JSON document following a strict schema defined in lines 27-92 of domain-analyzer.md. This schema mandates:
- Nodes for domains, flows, and steps with kebab-case IDs
- Edges linking flows to their parent domains and steps to their parent flows
- Weight ordering for steps to indicate execution sequence
- Validation rules ensuring every flow connects to at least one domain
Running the Domain Analyzer
Execute the domain extraction pipeline using the following commands:
# Stage 1: Generate domain-context.json from source code
python understand-anything-plugin/skills/understand-domain/extract-domain-context.py /path/to/my-app
# Output: .understand-anything/intermediate/domain-context.json
# Stage 2: Run the Domain Analyzer agent to produce domain-graph.json
understand --skill domain-analyzer \
--input .understand-anything/intermediate/domain-context.json \
--output .understand-anything/intermediate/domain-graph.json
The second command invokes the LLM agent through the Understand-Anything CLI wrapper. The agent reads the structured context and writes the final domain-graph.json containing the discovered business architecture.
Summary
- The domain analyzer operates in two distinct stages: a lightweight Python scanner followed by an LLM-driven synthesis agent.
extract-domain-context.pyfilters source files using.gitignoreandSKIP_DIRS(lines 22-48), then detects entry points viaENTRY_POINT_PATTERNS(lines 68-115) to identify HTTP routes, CLI commands, and event handlers.- File signatures and project metadata from files like
package.jsonare packaged intodomain-context.jsonwith size limits enforced byMAX_OUTPUT_BYTES(lines 43-73). - The Domain Analyzer agent processes this context to infer Business Domains, Business Flows, and Business Steps, outputting a structured
domain-graph.jsonfollowing the schema defined in lines 27-92. - The system maintains traceability by mapping each business step to specific source file locations and line numbers.
Frequently Asked Questions
What file types does the domain analyzer scan?
The scanner processes only files matching extensions defined in the SOURCE_EXTENSIONS constant while respecting .gitignore patterns and skipping directories listed in SKIP_DIRS such as node_modules, .git, and build artifacts. This filtering occurs in lines 22-48 of extract-domain-context.py.
How does the analyzer identify business entry points?
The analyzer uses a comprehensive regex collection stored in ENTRY_POINT_PATTERNS (lines 68-115) to detect HTTP routes, CLI commands, event listeners, cron jobs, GraphQL resolvers, and gRPC services. Each match captures the file path, line number, entry type, and surrounding code snippet to establish where business processes originate.
Can the domain analyzer process existing knowledge graphs?
Yes. The Domain Analyzer agent accepts two input options: Option A consumes the domain-context.json generated by the scanner, while Option B can process an existing knowledge-graph.json to refine or extend previously extracted business domains without rescanning the entire codebase (lines 13-20).
Where are the output files stored?
By default, the scanner writes domain-context.json to .understand-anything/intermediate/domain-context.json. The Domain Analyzer agent then outputs domain-graph.json to the configured intermediate directory. Both files follow strict JSON schemas that preserve mappings between business concepts and source code locations (lines 27-92).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →