How the Domain Analyzer Extracts Business Domains and Flows in Understand-Anything

The domain analyzer extracts business domains and flows through a two-stage pipeline that first scans source code to build a structured domain-context.json file, then uses a specialized LLM agent to synthesize a hierarchical domain graph mapping business concepts to specific code locations.

The Understand-Anything repository by Egonex-AI automates architectural discovery through a domain analyzer that transforms raw source code into actionable business knowledge. Unlike static documentation, this system programmatically identifies how technical implementations map to business capabilities by combining static analysis with large language model reasoning. The analyzer outputs a machine-readable graph connecting high-level domains to concrete code entry points and execution steps.

Stage 1: Lightweight Code Scanning and Context Generation

The extraction process begins with extract-domain-context.py, a Python scanner located at understand-anything-plugin/skills/understand-domain/extract-domain-context.py. This script performs a high-performance walk of the project tree while respecting .gitignore rules and predefined SKIP_DIRS to avoid irrelevant directories like node_modules or .git (lines 22-48).

Filtering Source Files with .gitignore and Skip Directories

The scanner initializes by filtering the file system against exclusion criteria. It only processes files matching extensions defined in SOURCE_EXTENSIONS and explicitly ignores paths listed in SKIP_DIRS (lines 22-48). This ensures the subsequent analysis focuses on relevant source code while maintaining performance on large repositories.

Detecting Entry Points Across Protocols

During traversal, the scanner identifies integration surfaces using regex patterns defined in ENTRY_POINT_PATTERNS (lines 68-115). These patterns detect:

  • HTTP routes and API endpoints
  • CLI commands and argument parsers
  • Event listeners and message queue handlers
  • Cron job definitions and scheduled tasks
  • GraphQL resolvers and gRPC service definitions

Each match records the file path, line number, entry type, a code snippet, and a generated description (lines 68-115). These entry points serve as the anchors for later business flow discovery, revealing where external actors trigger system behavior.

Extracting File Signatures and Metadata

For files likely containing business logic—identified through a "priority keywords" filter—the scanner generates lightweight file signatures. Each signature includes exported functions, imported dependencies, line counts, and a short preview of the content (lines 48-56, 67-99). Additionally, the scanner extracts metadata from standard project files including package.json, pyproject.toml, and README.md to capture project context and dependencies (lines 59-69).

Packaging the domain-context.json Output

All collected data—the file tree, entry points, signatures, and metadata—is compiled into domain-context.json. The scanner applies size constraints via MAX_OUTPUT_BYTES to ensure the output remains within LLM context limits, truncating content if necessary while preserving structural integrity (lines 43-73).

Stage 2: LLM-Driven Domain Graph Synthesis

Once the context file is generated, the Domain Analyzer agent defined in understand-anything-plugin/agents/domain-analyzer.md processes the structured data. This agent can ingest either the freshly generated domain-context.json (Option A) or an existing knowledge-graph.json (Option B) for iterative refinement (lines 13-20).

Consuming Context with the Domain Analyzer Agent

The agent receives the JSON context containing the file tree, entry points, and signatures. Rather than reading raw source code directly, it reasons over the pre-structured metadata to identify architectural boundaries. This approach reduces token consumption while maintaining precision in business concept extraction.

Inferring Business Domains, Flows, and Steps

Using the provided context, the LLM agent infers three hierarchical levels of business knowledge:

  1. Business Domains – High-level capability areas such as "Order Management" or "User Authentication"
  2. Business Flows – Concrete processes like "Create Order" or "Process Refund" derived from entry-point types (HTTP routes, CLI commands, etc.)
  3. Business Steps – Individual actions within each flow, mapped to specific source files and line ranges

The agent correlates entry points detected in the scanning phase with business processes, creating traceability between technical implementation and business capability.

Enforcing Schema Consistency and Validation

The agent outputs a JSON document following a strict schema defined in lines 27-92 of domain-analyzer.md. This schema mandates:

  • Nodes for domains, flows, and steps with kebab-case IDs
  • Edges linking flows to their parent domains and steps to their parent flows
  • Weight ordering for steps to indicate execution sequence
  • Validation rules ensuring every flow connects to at least one domain

Running the Domain Analyzer

Execute the domain extraction pipeline using the following commands:


# Stage 1: Generate domain-context.json from source code

python understand-anything-plugin/skills/understand-domain/extract-domain-context.py /path/to/my-app

# Output: .understand-anything/intermediate/domain-context.json

# Stage 2: Run the Domain Analyzer agent to produce domain-graph.json

understand --skill domain-analyzer \
  --input .understand-anything/intermediate/domain-context.json \
  --output .understand-anything/intermediate/domain-graph.json

The second command invokes the LLM agent through the Understand-Anything CLI wrapper. The agent reads the structured context and writes the final domain-graph.json containing the discovered business architecture.

Summary

  • The domain analyzer operates in two distinct stages: a lightweight Python scanner followed by an LLM-driven synthesis agent.
  • extract-domain-context.py filters source files using .gitignore and SKIP_DIRS (lines 22-48), then detects entry points via ENTRY_POINT_PATTERNS (lines 68-115) to identify HTTP routes, CLI commands, and event handlers.
  • File signatures and project metadata from files like package.json are packaged into domain-context.json with size limits enforced by MAX_OUTPUT_BYTES (lines 43-73).
  • The Domain Analyzer agent processes this context to infer Business Domains, Business Flows, and Business Steps, outputting a structured domain-graph.json following the schema defined in lines 27-92.
  • The system maintains traceability by mapping each business step to specific source file locations and line numbers.

Frequently Asked Questions

What file types does the domain analyzer scan?

The scanner processes only files matching extensions defined in the SOURCE_EXTENSIONS constant while respecting .gitignore patterns and skipping directories listed in SKIP_DIRS such as node_modules, .git, and build artifacts. This filtering occurs in lines 22-48 of extract-domain-context.py.

How does the analyzer identify business entry points?

The analyzer uses a comprehensive regex collection stored in ENTRY_POINT_PATTERNS (lines 68-115) to detect HTTP routes, CLI commands, event listeners, cron jobs, GraphQL resolvers, and gRPC services. Each match captures the file path, line number, entry type, and surrounding code snippet to establish where business processes originate.

Can the domain analyzer process existing knowledge graphs?

Yes. The Domain Analyzer agent accepts two input options: Option A consumes the domain-context.json generated by the scanner, while Option B can process an existing knowledge-graph.json to refine or extend previously extracted business domains without rescanning the entire codebase (lines 13-20).

Where are the output files stored?

By default, the scanner writes domain-context.json to .understand-anything/intermediate/domain-context.json. The Domain Analyzer agent then outputs domain-graph.json to the configured intermediate directory. Both files follow strict JSON schemas that preserve mappings between business concepts and source code locations (lines 27-92).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →