How the `/understand-domain` Command Extracts Business Domains, Flows, and Process Steps
The /understand-domain command runs a two-stage pipeline that first extracts code context via extract-domain-context.py, then uses an LLM-backed agent to generate a structured domain graph with business domains, flows, and steps.
The Egonex-AI/Understand-Anything repository provides a domain analysis tool that transforms raw source code into a visual representation of business logic. This command operates through a lightweight pipeline that bridges static code analysis with generative AI to extract business domains, flows, and process steps. The resulting domain graph enables developers to visualize cross-domain relationships and process flows directly in the dashboard.
Stage 1: Source Code Context Extraction
The first stage runs the bundled script extract-domain-context.py located at understand-anything-plugin/skills/understand-domain/extract-domain-context.py. This script performs static analysis to build a structured context object that feeds into the LLM.
Project Traversal and Filtering
The script traverses the source tree while respecting .gitignore patterns and a built-in skip list (SKIP_DIRS). It only processes files matching extensions defined in SOURCE_EXTENSIONS, applying configurable limits such as MAX_FILE_TREE_DEPTH and MAX_FILES_TOTAL to prevent overwhelming the downstream processor.
The output is a fileTree array containing relative paths of all qualifying source files.
Entry Point Detection
The script iterates over each file and applies regular-expression patterns (ENTRY_POINT_PATTERNS) to identify critical integration points:
- HTTP routes (Express, FastAPI, NestJS, Next.js, GraphQL, gRPC)
- CLI commands and argparse sub-parsers
- Event listeners and cron-style schedules
For each match, the script records the file path, line number, entry type, description, exact match string, and a surrounding code snippet. These records populate the entryPoints array.
File Signature Harvesting
A secondary pass (extract_file_signatures) targets a prioritized subset of files (capped at MAX_SAMPLED_FILES). For each file, it extracts:
- Exports: JavaScript/TypeScript
exportstatements or Pythondef/classdefinitions - Imports: First 20 import statements
- Line count and a truncated preview of the file content
This data structures the fileSignatures array, providing the LLM with function signatures and dependency graphs without full file contents.
Metadata Collection
The script harvests project metadata by reading standard configuration files like package.json, pyproject.toml, and README.*. This metadata section provides context about project dependencies, scripts, and human-readable descriptions.
Context Truncation
Because the downstream LLM agent operates within token budgets, the _truncate_to_fit function progressively compresses the payload. It limits the file tree depth, truncates code previews, and reduces signature counts until the JSON output is ≤ MAX_OUTPUT_BYTES (approximately 512 KB).
The final output is written to .understand-anything/intermediate/domain-context.json.
Stage 2: LLM-Based Domain Graph Generation
The second stage consumes domain-context.json and uses the domain-analyzer agent to infer higher-level business concepts.
Schema Definition
The domain graph structure is defined in packages/core/src/schema.ts, which declares three node types and three edge types:
// Node types
type Node = {
id: string; // e.g. "domain:order-management"
type: "domain" | "flow" | "step";
label: string;
domainMeta?: DomainMeta; // optional metadata for domain/flow/step
};
// Edge types
type Edge = {
source: string;
target: string;
type: "contains_flow" | "flow_step" | "cross_domain";
};
Domain Analyzer Agent
The domain-analyzer agent (part of the core analyzer suite) reads the context JSON and prompts a language model to:
- Identify business domains (e.g., "Order Management", "Payment Processing")
- Map flows within each domain (e.g., "Order Creation")
- Extract process steps within each flow (e.g., "Validate Order")
- Establish relationships via edges:
contains_flow(domain to flow),flow_step(flow to step), andcross_domain(domain to domain)
The agent validates the graph using validateGraph and persists the result to domain-graph.json via packages/core/src/persistence/index.ts (where DOMAIN_GRAPH_FILE = "domain-graph.json").
Visualization and Dashboard Integration
When the dashboard loads, it checks for domain-graph.json. If present, the UI switches to Domain view (viewMode: "domain"), rendering a horizontal flow graph of domains, flows, and steps. Cross-domain edges visualize dependencies between business areas.
The visualization logic resides in packages/dashboard/src/components/DomainGraphView.tsx, while the view state is managed in packages/dashboard/src/store.ts.
Running the Command
You can execute the extraction pipeline manually or via the Claude skill interface.
Manual Execution
Run the extraction script directly from the repository root:
understand-anything-plugin/skills/understand-domain/extract-domain-context.py /path/to/project
This creates .understand-anything/intermediate/domain-context.json.
Skill Invocation
Inside a Claude or Claude-Code session, invoke the command:
/understand-domain
According to skills/understand-domain/SKILL.md, this executes:
python ./extract-domain-context.py "$PROJECT_ROOT"
Sample Output Structure
The intermediate domain-context.json contains structured data like:
{
"projectRoot": "/path/to/project",
"fileCount": 124,
"fileTree": ["src/api/orders.ts", "src/services/payment.py"],
"entryPoints": [
{
"file": "src/api/orders.ts",
"line": 12,
"type": "http",
"description": "Express/Koa route",
"match": "app.get('/orders')",
"snippet": "app.get('/orders', (req, res) => { /* … */"
}
],
"fileSignatures": [
{
"file": "src/services/payment.py",
"exports": ["process_payment", "PaymentGateway"],
"imports": ["stripe", "logging"],
"lines": 210,
"preview": "def process_payment(order):\n ..."
}
]
}
The resulting domain-graph.json follows this structure:
{
"nodes": [
{ "id": "domain:order-management", "type": "domain", "label": "Order Management" },
{ "id": "flow:order-creation", "type": "flow", "label": "Order Creation" },
{ "id": "step:validate-order", "type": "step", "label": "Validate Order" }
],
"edges": [
{ "source": "domain:order-management", "target": "flow:order-creation", "type": "contains_flow" },
{ "source": "flow:order-creation", "target": "step:validate-order", "type": "flow_step" }
]
}
Summary
extract-domain-context.pyperforms static analysis to buildfileTree,entryPoints,fileSignatures, andmetadatawhile respecting.gitignoreand size limits.- The
_truncate_to_fitfunction ensures the JSON payload stays within theMAX_OUTPUT_BYTESbudget (≈512 KB). - The domain-analyzer agent consumes the context JSON and generates a validated domain graph with nodes (
domain,flow,step) and edges (contains_flow,flow_step,cross_domain). - The graph schema is defined in
packages/core/src/schema.tsand persisted todomain-graph.jsonviapackages/core/src/persistence/index.ts. - The dashboard renders the graph in Domain view using
DomainGraphView.tsx, displaying business domains as horizontal flows with cross-domain relationships.
Frequently Asked Questions
What file types does the /understand-domain command analyze?
The command processes files matching extensions defined in the SOURCE_EXTENSIONS constant within extract-domain-context.py. It typically includes JavaScript, TypeScript, Python, and other source files while respecting .gitignore patterns and the internal SKIP_DIRS list to exclude directories like node_modules and __pycache__.
How does the command identify business entry points?
The script applies ENTRY_POINT_PATTERNS regular expressions to detect HTTP routes (Express, FastAPI, NestJS), CLI commands, event listeners, and cron schedules. Each match records the file path, line number, entry type, and surrounding code snippet, enabling the LLM to understand the application's external interfaces.
Where is the domain graph stored and how is it validated?
The generated domain graph is written to domain-graph.json in the project root, as defined by DOMAIN_GRAPH_FILE in packages/core/src/persistence/index.ts. The domain-analyzer agent validates the graph structure against the TypeScript schema in packages/core/src/schema.ts before persistence, ensuring all nodes have valid types (domain, flow, step) and edges reference existing node IDs.
Can I run the extraction without the full dashboard?
Yes. You can execute extract-domain-context.py manually against any project directory to generate .understand-anything/intermediate/domain-context.json. This JSON file contains all extracted context and can be consumed independently or used to debug the LLM generation step without loading the dashboard interface.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →