How the Domain Analyzer in Egonex-AI Understand Anything Extracts Business Domains, Flows, and Process Steps
The Domain Analyzer employs a two-stage pipeline—first scanning codebases with extract-domain-context.py to detect entry points and file signatures, then constructing a hierarchical domain graph via the Domain Analyzer agent that maps business domains to flows and individual process steps.
The Egonex-AI Understand-Anything repository provides a sophisticated Domain Analyzer that transforms raw source code into structured business logic representations. This tool extracts business domains, flows, and process steps by combining static analysis with intelligent graph construction. The analyzer outputs a validated JSON graph that visualizes how business logic flows through your codebase.
Two-Stage Extraction Pipeline
The Domain Analyzer operates through two complementary stages that progressively refine raw code into a structured domain graph.
Stage 1: Context Generation with extract-domain-context.py
Located at understand-anything-plugin/skills/understand-domain/extract-domain-context.py, this Python script performs a lightweight scan when no existing knowledge graph is present. The utility walks the project tree while respecting .gitignore patterns and skipping common artifacts like node_modules and .git directories, as implemented at lines 49-57.
The scanner implements strict memory management to preserve LLM context windows, limiting input data to a maximum of 512 KB by truncating file trees, previews, and code snippets at lines 43-71. This ensures the analyzer can process large repositories without exceeding token limits.
Entry-point detection relies on a comprehensive set of regular expressions defined at lines 70-115 that identify HTTP routes, CLI commands, event listeners, cron schedules, GraphQL resolvers, gRPC services, and generic exported handlers. Each match records the file path, line number, invocation type (http/cli/event/cron/manual), and a representative code snippet.
Stage 2: Domain Graph Construction via the Domain Analyzer Agent
Defined in understand-anything-plugin/agents/domain-analyzer.md, the Domain Analyzer agent processes either the generated domain-context.json or an existing knowledge-graph.json. It traverses file signatures—including exports and imports—to discover logical boundaries such as controllers, services, handlers, and use-cases that constitute distinct business domains.
For each detected entry point, the agent creates a flow node populated with entryType (http, cli, event, etc.) and entryPoint (the specific route or command string), as defined at lines 29-33. The agent then analyzes surrounding code using the captured snippets to infer process steps—individual actions like validation, database writes, or external API calls. Each step node inherits the precise file path and line range from the source code, as specified at lines 75-81.
Detecting Entry Points and Business Flows
Business flows originate from entry points detected during the initial scan. The analyzer categorizes these entry points by their invocation mechanism, creating a clear boundary between external triggers and internal processing.
The agent maps detected entry points to flow nodes in the domain graph. These nodes contain metadata describing the trigger type and specific entry point identifier, enabling the visualization of how external requests or scheduled jobs initiate business processes through the entryType and entryPoint fields.
Mapping Process Steps and Domain Relationships
Process steps represent the granular operations within a business flow. The Domain Analyzer identifies these steps by examining the code structure surrounding each entry point, capturing validation logic, data transformations, persistence operations, and external service calls.
The analyzer constructs three distinct edge types to represent relationships:
contains_flow: Links domain nodes to their associated flow nodesflow_step: Connects steps within a flow using monotonically increasingweightvalues that encode execution order, as implemented at lines 99-104cross_domain: Describes interactions between business domains derived from import relationships or explicit cross-module calls, as defined at lines 88-92
Each node includes a domain-meta object containing entities, business rules, and cross-domain interactions inferred from naming conventions and code comments discovered in scanned signatures, as specified in types.ts at lines 29-36.
Output Schema and Validation
The domain graph adheres to the TypeScript schema defined in understand-anything-plugin/packages/core/src/types.ts, where node types are restricted to "domain", "flow", or "step" at lines 6-7, and edge types belong to "contains_flow", "flow_step", or "cross_domain" categories at lines 18-19.
Validation tests in domain-types.test.ts at lines 5-63 ensure the generated graph complies with the expected schema, verifying that edge aliases are normalized and all required fields are present. The final output writes to .understand-anything/intermediate/domain-analysis.json for consumption by the dashboard visualization.
Practical Implementation Examples
Execute the domain analysis workflow using the following commands:
# Generate the lightweight context (if no knowledge graph exists)
python extract-domain-context.py /path/to/project
# Run the Domain Analyzer skill
understand --skill domain-analyzer /path/to/project
Load the generated domain graph in TypeScript applications:
import type { KnowledgeGraph } from '@understand-anything/core';
import { readFile } from 'node:fs/promises';
async function loadDomainGraph(root: string): Promise<KnowledgeGraph> {
const path = `${root}/.understand-anything/intermediate/domain-analysis.json`;
const raw = await readFile(path, 'utf-8');
return JSON.parse(raw) as KnowledgeGraph;
}
Invoke the Python scanner programmatically for custom processing:
from pathlib import Path
import subprocess, json
project = Path("/my/project")
subprocess.run(["python", "extract-domain-context.py", str(project)], check=True)
# Load the context for custom processing
ctx_path = project / ".understand-anything" / "intermediate" / "domain-context.json"
with ctx_path.open() as f:
context = json.load(f)
print(f"Detected {len(context['entryPoints'])} entry points")
Summary
- The Domain Analyzer uses a two-stage pipeline: context generation via
extract-domain-context.pyfollowed by graph construction through the Domain Analyzer agent defined indomain-analyzer.md. - Entry points are detected using regex patterns for HTTP routes, CLI commands, events, cron jobs, and other invocation types at lines 70-115, with metadata including file paths and line numbers.
- Process steps are inferred from code analysis around entry points at lines 75-81, capturing validation, database operations, and external calls with precise source locations.
- The output follows a strict schema defined in
types.tswith three node types (domain, flow, step) and three edge types (contains_flow,flow_step,cross_domain). - Results are validated against
domain-types.test.tsand written to.understand-anything/intermediate/domain-analysis.json.
Frequently Asked Questions
What file types does the domain analyzer detect as entry points?
The analyzer recognizes HTTP routes, CLI commands, event listeners, cron schedules, GraphQL resolvers, gRPC services, and generic exported handlers through regex patterns defined in extract-domain-context.py at lines 70-115. Each detected entry point records the file path, line number, invocation type, and a code snippet for downstream analysis.
How does the analyzer handle large codebases that exceed LLM context limits?
The context generator limits data to 512 KB by truncating file trees and snippets, as implemented in extract-domain-context.py at lines 43-71. This ensures the Domain Analyzer agent receives manageable inputs regardless of repository size while preserving essential structural information.
What is the relationship between domains, flows, and steps in the graph?
Domains contain flows via contains_flow edges, while flows comprise ordered steps connected by flow_step edges with monotonically increasing weight values encoding execution sequence, as defined in domain-analyzer.md at lines 99-104. Cross-domain interactions are mapped via cross_domain edges at lines 88-92.
How is the domain graph schema validated?
The generated graph is validated against the TypeScript definitions in types.ts where node types are restricted to "domain", "flow", or "step" at lines 6-7, and edge types are enumerated at lines 18-19. The test suite in domain-types.test.ts at lines 5-63 ensures compliance with these schema requirements and verifies that edge aliases are properly normalized.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →