How the Domain Analyzer in Egonex-AI Understand-Anything Extracts Business Domains, Flows, and Process Steps

The Domain Analyzer transforms raw source code into a structured domain graph by first scanning for entry points and signatures via extract-domain-context.py, then constructing hierarchical relationships between business domains, flows, and process steps using the agent logic defined in domain-analyzer.md.

The Domain Analyzer in the Egonex-AI/Understand-Anything repository is a specialized agent that bridges the gap between raw code and business architecture. It operates through a two-stage pipeline that first extracts lightweight context from the codebase, then builds a comprehensive domain graph modeling how business logic actually flows through your application.

The Two-Stage Extraction Architecture

The domain analyzer operates through complementary stages: a lightweight context generation scan followed by intelligent graph construction.

Stage 1: Context Generation

If a full knowledge graph does not yet exist, the skill extract-domain-context.py performs an initial project tree walk. This script respects .gitignore patterns and skips common artifacts like node_modules and .git directories【extract-domain-context.py†L49-L57】.

To stay within LLM context window limits, the scanner truncates file trees, previews, and snippets to a maximum of 512 KB【extract-domain-context.py†L43-L71】. During this scan, it collects:

  • Entry-point signatures
  • File exports and imports
  • High-level metadata

Stage 2: Domain Graph Construction

The Domain Analyzer agent receives either the generated domain-context.json or an existing knowledge-graph.json. It then traverses file signatures to discover logical boundaries such as controllers, services, handlers, and use-cases. These boundaries map to domains (high-level business areas) and form the foundation of the hierarchical graph.

Detecting Entry Points and Flow Boundaries

Entry-point detection uses regular expressions that recognize HTTP routes, CLI commands, event listeners, cron schedules, GraphQL resolvers, gRPC services, and generic exported handlers【extract-domain-context.py†L70-L115】.

Each match records:

  • The file path and line number
  • The entry type (http, cli, event, cron, or manual)
  • A short code snippet for context

The agent maps each detected entry point to a flow node, where the entry-point type becomes the entryType field and the route or command string becomes entryPoint【domain-analyzer.md†L29-L33】.

Mapping Process Steps and Domain Relationships

Step Extraction

The analyzer examines surrounding code using snippets and adjacent lines to infer process steps—individual actions that the flow performs such as validation, database writes, or external API calls. Step nodes inherit the file path and line range from the source file【domain-analyzer.md†L75-L81】.

Domain Metadata and Edge Relationships

Each node receives a domain-meta object containing entities, business rules, and cross-domain interactions based on naming conventions and comments discovered in scanned signatures【types.ts†L29-L36】.

The resulting graph emits three layers of edges:

  • contains_flow links a domain node to each of its flow nodes
  • flow_step orders steps inside a flow using monotonically increasing weight values that encode step sequence【domain-analyzer.md†L99-L104】
  • cross_domain describes interactions between domains derived from import relationships or explicit cross-module calls【domain-analyzer.md†L88-L92】

The schema defines node type values as "domain", "flow", or "step"【types.ts†L6-L7】, while edges belong to the "contains_flow", "flow_step", or "cross_domain" types【types.ts†L18-L19】. Validation tests in domain-types.test.ts ensure graph compliance【domain-types.test.ts†L5-L63】.

Running the Domain Analyzer

Execute the workflow in two phases:


# 1️⃣ Generate the lightweight context (if no knowledge-graph exists)

python extract-domain-context.py /path/to/project

# 2️⃣ Run the Domain Analyzer skill

understand --skill domain-analyzer /path/to/project

The skill reads the produced domain-context.json, builds the hierarchical domain graph, and writes it to:


<project-root>/.understand-anything/intermediate/domain-analysis.json

The dashboard visualizes this graph in a domain-first view, allowing exploration of how business logic flows through the codebase.

Programmatic Usage Examples

Load the generated domain graph in TypeScript:

import type { KnowledgeGraph } from '@understand-anything/core';
import { readFile } from 'node:fs/promises';

async function loadDomainGraph(root: string): Promise<KnowledgeGraph> {
  const path = `${root}/.understand-anything/intermediate/domain-analysis.json`;
  const raw = await readFile(path, 'utf-8');
  return JSON.parse(raw) as KnowledgeGraph;
}

Invoke the Python scanner programmatically:

from pathlib import Path
import subprocess, json

project = Path("/my/project")
subprocess.run(["python", "extract-domain-context.py", str(project)], check=True)

# Load the context for custom processing

ctx_path = project / ".understand-anything" / "intermediate" / "domain-context.json"
with ctx_path.open() as f:
    context = json.load(f)
print(f"Detected {len(context['entryPoints'])} entry points")

Key Source Files

File Role
extract-domain-context.py Scans the repository, builds domain-context.json with file tree, entry points, signatures, and metadata.
domain-analyzer.md Agent definition that transforms the context into a hierarchical domain graph.
types.ts Type definitions for nodes, edges, and the overall KnowledgeGraph schema used by the analyzer.
domain-types.test.ts Test suite ensuring the domain graph conforms to the expected schema and edge aliases are normalized.
merge-subdomain-graphs.py Utility that combines multiple domain-analysis outputs into a single graph for multi-repository scenarios.

Summary

  • The Domain Analyzer uses a two-stage pipeline: context generation via extract-domain-context.py and graph construction via the domain analyzer agent.
  • Entry points are detected using regex patterns for HTTP routes, CLI commands, events, and scheduled jobs, then mapped to flow nodes.
  • Process steps are extracted from code snippets and linked via flow_step edges with weighted ordering.
  • The output follows the KnowledgeGraph schema defined in types.ts, with validation ensuring structural integrity.
  • Results are written to .understand-anything/intermediate/domain-analysis.json for dashboard visualization.

Frequently Asked Questions

What is the difference between a domain, a flow, and a step in the analyzer output?

A domain represents a high-level business area (such as "billing" or "authentication") discovered from logical code boundaries like controllers and services. A flow corresponds to an entry point (an HTTP endpoint or CLI command) that initiates business logic. A step is an individual action within that flow, such as validation logic or a database query. The three connect via contains_flow and flow_step relationships.

How does the scanner handle large repositories without exceeding LLM context limits?

The extract-domain-context.py script enforces a 512 KB ceiling on the data it extracts by truncating file trees, previews, and code snippets. It also respects .gitignore patterns to skip irrelevant directories like node_modules and .git, ensuring only pertinent source information enters the context window.

Can the Domain Analyzer work with existing knowledge graphs rather than raw source?

Yes. The agent accepts either a freshly generated domain-context.json or an existing knowledge-graph.json. If the knowledge graph already exists, the analyzer bypasses the initial scanning stage and proceeds directly to domain graph construction, enriching the existing structure with business domain hierarchies.

What entry point types does the domain analyzer recognize?

According to the source code in extract-domain-context.py【extract-domain-context.py†L70-L115】, the analyzer recognizes HTTP routes, CLI commands, event listeners, cron schedules, GraphQL resolvers, gRPC services, and manual exported handlers. Each type is stored in the entryType field of the flow node.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →