# How the Domain Analyzer Extracts Business Domains, Flows, and Process Steps in Understand-Anything

> Learn how the Understand Anything Domain Analyzer extracts business domains, flows, and process steps using its two-stage code scanning and hierarchical graph construction pipeline. Understand your code's structure.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-09

---

**The Domain Analyzer in Egonex-AI/Understand-Anything uses a two-stage pipeline—first scanning source code with [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) to detect entry points and signatures, then constructing a hierarchical domain graph that maps business domains to flows and individual process steps.**

The **Domain Analyzer** is a specialized agent within the [Understand-Anything](https://github.com/Egonex-AI/Understand-Anything) repository that transforms raw source code into a structured domain graph. This graph represents business domains, their associated flows, and the granular process steps that implement them, enabling developers to visualize and navigate complex codebases through a business-logic lens.

## Two-Stage Architecture of the Domain Analyzer

The extraction process operates in two complementary stages: an initial lightweight context generation followed by intelligent graph construction.

### Stage 1: Context Generation with extract-domain-context.py

The **Python scanner** ([`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)) performs a lightweight walk of the project tree when no existing knowledge graph is present. This stage respects `.gitignore` patterns and skips common artifacts like `node_modules` and `.git` directories to focus on relevant source files.

The scanner enforces a **512 KB limit** on collected data to stay within LLM context windows. It truncates file trees, previews, and code snippets intelligently to capture essential metadata without overwhelming downstream processors.

**Entry-point detection** uses targeted regular expressions to identify HTTP routes, CLI commands, event listeners, cron schedules, GraphQL resolvers, gRPC services, and generic exported handlers. Each match records the file path, line number, entry type (`http`, `cli`, `event`, `cron`, or `manual`), and a representative code snippet. This detection logic spans lines 70-115 in the source file.

### Stage 2: Domain Graph Construction

The **Domain Analyzer agent** (defined in [`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md)) receives either the generated [`domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-context.json) or an existing [`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json). It traverses file signatures—exports and imports—to discover logical boundaries such as controllers, services, handlers, and use-cases. These boundaries map to high-level business **domains**.

Each detected entry point becomes a **flow** node. The entry-point type populates the `entryType` field, while the specific route or command string becomes the `entryPoint` identifier.

## How Entry Points Map to Business Flows

The analyzer treats entry points as the genesis of business flows. When the agent processes the context, it:

1. **Identifies flow boundaries** by examining imports and exports that suggest architectural layers (e.g., a controller importing a service).
2. **Creates flow nodes** linked to their parent domain via `contains_flow` edges.
3. **Populates metadata** including the triggering mechanism (HTTP endpoint, scheduled job, etc.) and the implementing file location.

This mapping allows the dashboard to present a *domain-first* view where users can explore how external triggers cascade through the business logic.

## Extracting Process Steps and Domain Metadata

The analyzer examines the code surrounding each entry point to infer **process steps**—individual actions such as validation, database writes, or external API calls. Step nodes inherit precise file paths and line ranges from the source, enabling accurate traceability from the graph back to the implementation.

For each node, the analyzer populates a **domain-meta** object containing:
- **Entities**: Business objects discovered through naming conventions
- **Business rules**: Logic inferred from code structure and comments
- **Cross-domain interactions**: Dependencies identified through import relationships or explicit cross-module calls

The agent emits three distinct edge types:
- **`contains_flow`**: Links domain nodes to their constituent flows
- **`flow_step`**: Orders steps within a flow using monotonically increasing `weight` values that encode execution sequence
- **`cross_domain`**: Describes interactions between separate business domains

## Output Schema and Validation

The resulting domain graph conforms to the TypeScript schema defined in **[`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts)**, where nodes are typed as `"domain"`, `"flow"`, or `"step"`, and edges are strictly typed as `"contains_flow"`, `"flow_step"`, or `"cross_domain"`.

Validation tests in **[`domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-types.test.ts)** ensure that generated graphs comply with the schema and that edge aliases are properly normalized, preventing structural inconsistencies that could break the visualization layer.

## Practical Usage Workflow

Integrate the Domain Analyzer into your development workflow using the following commands:

```bash

# Stage 1: Generate the lightweight context (if no knowledge graph exists)

python extract-domain-context.py /path/to/project

# Stage 2: Run the Domain Analyzer agent

understand --skill domain-analyzer /path/to/project

```

The analyzer writes intermediate outputs to:

```

<project-root>/.understand-anything/intermediate/domain-context.json
<project-root>/.understand-anything/intermediate/domain-analysis.json

```

### Loading Results Programmatically

Access the generated domain graph in TypeScript applications:

```typescript
import type { KnowledgeGraph } from '@understand-anything/core';
import { readFile } from 'node:fs/promises';

async function loadDomainGraph(root: string): Promise<KnowledgeGraph> {
  const path = `${root}/.understand-anything/intermediate/domain-analysis.json`;
  const raw = await readFile(path, 'utf-8');
  return JSON.parse(raw) as KnowledgeGraph;
}

```

Or invoke the Python scanner programmatically:

```python
from pathlib import Path
import subprocess, json

project = Path("/my/project")
subprocess.run(["python", "extract-domain-context.py", str(project)], check=True)

# Load the context for custom processing

ctx_path = project / ".understand-anything" / "intermediate" / "domain-context.json"
with ctx_path.open() as f:
    context = json.load(f)
print(f"Detected {len(context['entryPoints'])} entry points")

```

## Summary

- The **Domain Analyzer** operates in two stages: Python-based context generation ([`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py)) and agent-based graph construction ([`domain-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analyzer.md)).
- **Entry-point detection** uses regex patterns to identify HTTP routes, CLI commands, events, cron jobs, and other triggers across lines 70-115 of the scanner.
- The output is a **hierarchical domain graph** with three node types (domain, flow, step) and three edge types (`contains_flow`, `flow_step`, `cross_domain`).
- **Process steps** inherit file paths and line ranges from source code, while `weight` values on edges encode execution order.
- All outputs follow the strict schema in **[`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts)** and are validated by **[`domain-types.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-types.test.ts)**.

## Frequently Asked Questions

### What file types does the Domain Analyzer support?

The analyzer language-agnostic regex patterns in [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) detect entry points across common backend languages including JavaScript, TypeScript, Python, Go, and Java. The scanner respects `.gitignore` automatically, skipping binary files and dependency directories like `node_modules` to focus on implementation code.

### How does the analyzer handle large codebases?

The [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) scanner limits data collection to 512 KB per project to respect LLM context windows. It intelligently truncates file trees and code previews while preserving critical metadata such as entry-point locations and signature exports, ensuring the Domain Analyzer can process monorepos and large legacy systems without overwhelming the graph construction phase.

### What is the difference between domain-context.json and domain-analysis.json?

The [`domain-context.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-context.json) file is the raw output of the Python scanner containing file trees, entry points, and code snippets. The [`domain-analysis.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/domain-analysis.json) file is the processed output of the Domain Analyzer agent containing the complete hierarchical graph with business domains, flows, process steps, and relationship edges. The former is input; the latter is the final navigable structure.

### Can I extend the entry-point detection patterns?

Yes. The entry-point detection logic in [`extract-domain-context.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extract-domain-context.py) (lines 70-115) uses configurable regular expressions. You can modify these patterns to recognize custom framework conventions, internal APIs, or domain-specific triggers before running the scanner. The agent will then map these custom entry points to flows using the same graph construction logic.