# How the Domain Analyzer in Egonex-AI Understand Anything Extracts Business Domains, Flows, and Process Steps

> Discover how Egonex-AI's Domain Analyzer extracts business domains, flows, and process steps using a two-stage pipeline of code scanning and hierarchical graph construction. Learn more today.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-13

---

**The Domain Analyzer employs a two-stage pipeline—first scanning codebases with [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) to detect entry points and file signatures, then constructing a hierarchical domain graph via the Domain Analyzer agent that maps business domains to flows and individual process steps.**

The Egonex-AI Understand-Anything repository provides a sophisticated Domain Analyzer that transforms raw source code into structured business logic representations. This tool extracts business domains, flows, and process steps by combining static analysis with intelligent graph construction. The analyzer outputs a validated JSON graph that visualizes how business logic flows through your codebase.

## Two-Stage Extraction Pipeline

The Domain Analyzer operates through two complementary stages that progressively refine raw code into a structured domain graph.

### Stage 1: Context Generation with extract-domain-context.py

Located at [`understand-anything-plugin/skills/understand-domain/extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand-domain/extract-domain-context.py), this Python script performs a lightweight scan when no existing knowledge graph is present. The utility walks the project tree while respecting `.gitignore` patterns and skipping common artifacts like `node_modules` and `.git` directories, as implemented at lines 49-57.

The scanner implements strict memory management to preserve LLM context windows, limiting input data to a maximum of 512 KB by truncating file trees, previews, and code snippets at lines 43-71. This ensures the analyzer can process large repositories without exceeding token limits.

Entry-point detection relies on a comprehensive set of regular expressions defined at lines 70-115 that identify HTTP routes, CLI commands, event listeners, cron schedules, GraphQL resolvers, gRPC services, and generic exported handlers. Each match records the file path, line number, invocation type (http/cli/event/cron/manual), and a representative code snippet.

### Stage 2: Domain Graph Construction via the Domain Analyzer Agent

Defined in [`understand-anything-plugin/agents/domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/agents/domain-analyzer.md), the Domain Analyzer agent processes either the generated [`domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-context.json) or an existing [`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json). It traverses file signatures—including exports and imports—to discover logical boundaries such as controllers, services, handlers, and use-cases that constitute distinct business domains.

For each detected entry point, the agent creates a flow node populated with `entryType` (http, cli, event, etc.) and `entryPoint` (the specific route or command string), as defined at lines 29-33. The agent then analyzes surrounding code using the captured snippets to infer process steps—individual actions like validation, database writes, or external API calls. Each step node inherits the precise file path and line range from the source code, as specified at lines 75-81.

## Detecting Entry Points and Business Flows

Business flows originate from entry points detected during the initial scan. The analyzer categorizes these entry points by their invocation mechanism, creating a clear boundary between external triggers and internal processing.

The agent maps detected entry points to flow nodes in the domain graph. These nodes contain metadata describing the trigger type and specific entry point identifier, enabling the visualization of how external requests or scheduled jobs initiate business processes through the `entryType` and `entryPoint` fields.

## Mapping Process Steps and Domain Relationships

Process steps represent the granular operations within a business flow. The Domain Analyzer identifies these steps by examining the code structure surrounding each entry point, capturing validation logic, data transformations, persistence operations, and external service calls.

The analyzer constructs three distinct edge types to represent relationships:

- `contains_flow`: Links domain nodes to their associated flow nodes
- `flow_step`: Connects steps within a flow using monotonically increasing `weight` values that encode execution order, as implemented at lines 99-104
- `cross_domain`: Describes interactions between business domains derived from import relationships or explicit cross-module calls, as defined at lines 88-92

Each node includes a `domain-meta` object containing entities, business rules, and cross-domain interactions inferred from naming conventions and code comments discovered in scanned signatures, as specified in [`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts) at lines 29-36.

## Output Schema and Validation

The domain graph adheres to the TypeScript schema defined in [`understand-anything-plugin/packages/core/src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/types.ts), where node types are restricted to `"domain"`, `"flow"`, or `"step"` at lines 6-7, and edge types belong to `"contains_flow"`, `"flow_step"`, or `"cross_domain"` categories at lines 18-19.

Validation tests in [`domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-types.test.ts) at lines 5-63 ensure the generated graph complies with the expected schema, verifying that edge aliases are normalized and all required fields are present. The final output writes to [`.understand-anything/intermediate/domain-analysis.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/domain-analysis.json) for consumption by the dashboard visualization.

## Practical Implementation Examples

Execute the domain analysis workflow using the following commands:

```bash

# Generate the lightweight context (if no knowledge graph exists)

python extract-domain-context.py /path/to/project

# Run the Domain Analyzer skill

understand --skill domain-analyzer /path/to/project

```

Load the generated domain graph in TypeScript applications:

```typescript
import type { KnowledgeGraph } from '@understand-anything/core';
import { readFile } from 'node:fs/promises';

async function loadDomainGraph(root: string): Promise<KnowledgeGraph> {
  const path = `${root}/.understand-anything/intermediate/domain-analysis.json`;
  const raw = await readFile(path, 'utf-8');
  return JSON.parse(raw) as KnowledgeGraph;
}

```

Invoke the Python scanner programmatically for custom processing:

```python
from pathlib import Path
import subprocess, json

project = Path("/my/project")
subprocess.run(["python", "extract-domain-context.py", str(project)], check=True)

# Load the context for custom processing

ctx_path = project / ".understand-anything" / "intermediate" / "domain-context.json"
with ctx_path.open() as f:
    context = json.load(f)
print(f"Detected {len(context['entryPoints'])} entry points")

```

## Summary

- The Domain Analyzer uses a two-stage pipeline: context generation via [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) followed by graph construction through the Domain Analyzer agent defined in [`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md).
- Entry points are detected using regex patterns for HTTP routes, CLI commands, events, cron jobs, and other invocation types at lines 70-115, with metadata including file paths and line numbers.
- Process steps are inferred from code analysis around entry points at lines 75-81, capturing validation, database operations, and external calls with precise source locations.
- The output follows a strict schema defined in [`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts) with three node types (domain, flow, step) and three edge types (`contains_flow`, `flow_step`, `cross_domain`).
- Results are validated against [`domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-types.test.ts) and written to [`.understand-anything/intermediate/domain-analysis.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/domain-analysis.json).

## Frequently Asked Questions

### What file types does the domain analyzer detect as entry points?

The analyzer recognizes HTTP routes, CLI commands, event listeners, cron schedules, GraphQL resolvers, gRPC services, and generic exported handlers through regex patterns defined in [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) at lines 70-115. Each detected entry point records the file path, line number, invocation type, and a code snippet for downstream analysis.

### How does the analyzer handle large codebases that exceed LLM context limits?

The context generator limits data to 512 KB by truncating file trees and snippets, as implemented in [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) at lines 43-71. This ensures the Domain Analyzer agent receives manageable inputs regardless of repository size while preserving essential structural information.

### What is the relationship between domains, flows, and steps in the graph?

Domains contain flows via `contains_flow` edges, while flows comprise ordered steps connected by `flow_step` edges with monotonically increasing `weight` values encoding execution sequence, as defined in [`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md) at lines 99-104. Cross-domain interactions are mapped via `cross_domain` edges at lines 88-92.

### How is the domain graph schema validated?

The generated graph is validated against the TypeScript definitions in [`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts) where node types are restricted to `"domain"`, `"flow"`, or `"step"` at lines 6-7, and edge types are enumerated at lines 18-19. The test suite in [`domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-types.test.ts) at lines 5-63 ensures compliance with these schema requirements and verifies that edge aliases are properly normalized.