# How the Domain Analyzer in Egonex-AI Understand-Anything Extracts Business Domains, Flows, and Process Steps

> Discover how Egonex-AI's Domain Analyzer extracts business domains, flows, and process steps from source code. Learn about its two-step transformation process for structured domain graphs.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-14

---

**The Domain Analyzer transforms raw source code into a structured domain graph by first scanning for entry points and signatures via [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py), then constructing hierarchical relationships between business domains, flows, and process steps using the agent logic defined in [`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md).**

The Domain Analyzer in the **Egonex-AI/Understand-Anything** repository is a specialized agent that bridges the gap between raw code and business architecture. It operates through a two-stage pipeline that first extracts lightweight context from the codebase, then builds a comprehensive domain graph modeling how business logic actually flows through your application.

## The Two-Stage Extraction Architecture

The domain analyzer operates through complementary stages: a lightweight context generation scan followed by intelligent graph construction.

### Stage 1: Context Generation

If a full knowledge graph does not yet exist, the skill **[`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)** performs an initial project tree walk. This script respects `.gitignore` patterns and skips common artifacts like `node_modules` and `.git` directories【[`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)†L49-L57】.

To stay within LLM context window limits, the scanner truncates file trees, previews, and snippets to a maximum of **512 KB**【[`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)†L43-L71】. During this scan, it collects:

- Entry-point signatures
- File exports and imports
- High-level metadata

### Stage 2: Domain Graph Construction

The **Domain Analyzer** agent receives either the generated [`domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-context.json) or an existing [`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json). It then traverses file signatures to discover logical boundaries such as controllers, services, handlers, and use-cases. These boundaries map to **domains** (high-level business areas) and form the foundation of the hierarchical graph.

## Detecting Entry Points and Flow Boundaries

Entry-point detection uses regular expressions that recognize HTTP routes, CLI commands, event listeners, cron schedules, GraphQL resolvers, gRPC services, and generic exported handlers【[`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)†L70-L115】.

Each match records:
- The file path and line number
- The **entry type** (`http`, `cli`, `event`, `cron`, or `manual`)
- A short code snippet for context

The agent maps each detected entry point to a **flow** node, where the entry-point type becomes the `entryType` field and the route or command string becomes `entryPoint`【[`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md)†L29-L33】.

## Mapping Process Steps and Domain Relationships

### Step Extraction

The analyzer examines surrounding code using snippets and adjacent lines to infer **process steps**—individual actions that the flow performs such as validation, database writes, or external API calls. Step nodes inherit the file path and line range from the source file【[`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md)†L75-L81】.

### Domain Metadata and Edge Relationships

Each node receives a **domain-meta** object containing entities, business rules, and cross-domain interactions based on naming conventions and comments discovered in scanned signatures【[`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts)†L29-L36】.

The resulting graph emits three layers of edges:

- **`contains_flow`** links a domain node to each of its flow nodes
- **`flow_step`** orders steps inside a flow using monotonically increasing `weight` values that encode step sequence【[`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md)†L99-L104】
- **`cross_domain`** describes interactions between domains derived from import relationships or explicit cross-module calls【[`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md)†L88-L92】

The schema defines node `type` values as `"domain"`, `"flow"`, or `"step"`【[`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts)†L6-L7】, while edges belong to the `"contains_flow"`, `"flow_step"`, or `"cross_domain"` types【[`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts)†L18-L19】. Validation tests in [`domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-types.test.ts) ensure graph compliance【[`domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-types.test.ts)†L5-L63】.

## Running the Domain Analyzer

Execute the workflow in two phases:

```bash

# 1️⃣ Generate the lightweight context (if no knowledge-graph exists)

python extract-domain-context.py /path/to/project

# 2️⃣ Run the Domain Analyzer skill

understand --skill domain-analyzer /path/to/project

```

The skill reads the produced [`domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-context.json), builds the hierarchical domain graph, and writes it to:

```

<project-root>/.understand-anything/intermediate/domain-analysis.json

```

The dashboard visualizes this graph in a domain-first view, allowing exploration of how business logic flows through the codebase.

### Programmatic Usage Examples

Load the generated domain graph in TypeScript:

```typescript
import type { KnowledgeGraph } from '@understand-anything/core';
import { readFile } from 'node:fs/promises';

async function loadDomainGraph(root: string): Promise<KnowledgeGraph> {
  const path = `${root}/.understand-anything/intermediate/domain-analysis.json`;
  const raw = await readFile(path, 'utf-8');
  return JSON.parse(raw) as KnowledgeGraph;
}

```

Invoke the Python scanner programmatically:

```python
from pathlib import Path
import subprocess, json

project = Path("/my/project")
subprocess.run(["python", "extract-domain-context.py", str(project)], check=True)

# Load the context for custom processing

ctx_path = project / ".understand-anything" / "intermediate" / "domain-context.json"
with ctx_path.open() as f:
    context = json.load(f)
print(f"Detected {len(context['entryPoints'])} entry points")

```

## Key Source Files

| File | Role |
|------|------|
| [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) | Scans the repository, builds [`domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-context.json) with file tree, entry points, signatures, and metadata. |
| [`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md) | Agent definition that transforms the context into a hierarchical domain graph. |
| [`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts) | Type definitions for nodes, edges, and the overall `KnowledgeGraph` schema used by the analyzer. |
| [`domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-types.test.ts) | Test suite ensuring the domain graph conforms to the expected schema and edge aliases are normalized. |
| [`merge-subdomain-graphs.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/merge-subdomain-graphs.py) | Utility that combines multiple domain-analysis outputs into a single graph for multi-repository scenarios. |

## Summary

- The Domain Analyzer uses a two-stage pipeline: **context generation** via [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) and **graph construction** via the domain analyzer agent.
- Entry points are detected using regex patterns for HTTP routes, CLI commands, events, and scheduled jobs, then mapped to flow nodes.
- Process steps are extracted from code snippets and linked via `flow_step` edges with weighted ordering.
- The output follows the `KnowledgeGraph` schema defined in [`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts), with validation ensuring structural integrity.
- Results are written to [`.understand-anything/intermediate/domain-analysis.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/domain-analysis.json) for dashboard visualization.

## Frequently Asked Questions

### What is the difference between a domain, a flow, and a step in the analyzer output?

**A domain** represents a high-level business area (such as "billing" or "authentication") discovered from logical code boundaries like controllers and services. **A flow** corresponds to an entry point (an HTTP endpoint or CLI command) that initiates business logic. **A step** is an individual action within that flow, such as validation logic or a database query. The three connect via `contains_flow` and `flow_step` relationships.

### How does the scanner handle large repositories without exceeding LLM context limits?

The [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) script enforces a **512 KB ceiling** on the data it extracts by truncating file trees, previews, and code snippets. It also respects `.gitignore` patterns to skip irrelevant directories like `node_modules` and `.git`, ensuring only pertinent source information enters the context window.

### Can the Domain Analyzer work with existing knowledge graphs rather than raw source?

Yes. The agent accepts either a freshly generated [`domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-context.json) or an existing [`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json). If the knowledge graph already exists, the analyzer bypasses the initial scanning stage and proceeds directly to domain graph construction, enriching the existing structure with business domain hierarchies.

### What entry point types does the domain analyzer recognize?

According to the source code in [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)【[`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)†L70-L115】, the analyzer recognizes **HTTP routes**, **CLI commands**, **event listeners**, **cron schedules**, **GraphQL resolvers**, **gRPC services**, and **manual exported handlers**. Each type is stored in the `entryType` field of the flow node.