# How the Multi-Agent Pipeline in Understand-Anything Orchestrates Codebase Analysis

> Discover how the multi-agent pipeline in Understand Anything orchestrates codebase analysis. Learn how raw code transforms into an interactive knowledge graph step by step.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: architecture
- Published: 2026-06-01

---

**The multi-agent pipeline in Understand-Anything sequentially executes seven specialized agents—starting with project discovery and ending with graph validation—to transform raw code into an interactive knowledge graph via the `/understand` command.**

The open-source tool [Lum1104/Understand-Anything](https://github.com/Lum1104/Understand-Anything) automates codebase comprehension through a deterministic, stage-gated workflow. By combining deterministic static analysis scripts with LLM-powered semantic enrichment, the multi-agent pipeline breaks repository analysis into discrete, scalable phases that generate a queryable knowledge graph.

## The Seven Specialized Agents

The pipeline delegates work to focused agents, each defined in dedicated markdown specifications under `understand-anything-plugin/agents/`:

- **project-scanner** – Discovers file structure, detects languages, and builds import maps using deterministic scripts.
- **file-analyzer** – Extracts structural nodes (functions, classes) and edges via Tree-sitter, then enriches them with LLM summaries.
- **architecture-analyzer** – Groups files into logical layers (API, service, infrastructure) based on import topology and semantic analysis.
- **tour-builder** – Generates a pedagogical walkthrough of the codebase by ranking entry points and critical paths.
- **graph-reviewer** – Validates graph integrity, removes dangling edges, and optionally triggers a full LLM consistency review.
- **domain-analyzer** *(optional)* – Extracts business-domain concepts and workflows when invoked via `/understand-domain`.
- **article-analyzer** *(optional)* – Parses wiki-style knowledge bases into entities and relationships for `/understand-knowledge`.

The high-level architecture is documented in [README.md lines 83-95](https://github.com/Lum1104/Understand-Anything/blob/main/README.md#L83-L95).

## Step-by-Step Orchestration Flow

When you run `/understand`, the orchestrator in [`understand-chat.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-chat.ts) triggers a five-phase execution sequence. File analysis runs in parallel batches, while architectural and tour phases run sequentially to ensure data dependencies are met.

### Phase 1: Project Discovery (project-scanner)

The **project-scanner** agent executes three deterministic sub-steps:

1. **Narrative Extraction** – Reads [`README.md`](https://github.com/Lum1104/Understand-Anything/blob/main/README.md), [`package.json`](https://github.com/Lum1104/Understand-Anything/blob/main/package.json), and configuration files to extract project context via LLM.
2. **File Enumeration** – Executes [`scan-project.mjs`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/scan-project.mjs) to produce `files[]` with language detection, categorization, and line counts.
3. **Import Mapping** – Runs [`extract-import-map.mjs`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/extract-import-map.mjs) to generate a deterministic cross-language import graph.

Outputs are written to [`.understand-anything/tmp/ua-scan-files.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/tmp/ua-scan-files.json) and [`ua-import-map-output.json`](https://github.com/Lum1104/Understand-Anything/blob/main/ua-import-map-output.json) as specified in [project-scanner.md lines 51-78](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/project-scanner.md#L51-L78).

### Phase 2: Structural Analysis (file-analyzer)

The **file-analyzer** processes the discovered files in parallel batches (default: ~20-30 files per batch, up to 5 concurrent). For each batch:

- It invokes [`extract-structure.mjs`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/extract-structure.mjs) to parse ASTs and extract raw structural data.
- It enriches each node with LLM-generated summaries, complexity metrics, and semantic tags.
- It emits `batch-<idx>.json` fragments containing `nodes[]` and `edges[]` arrays.

Incremental runs cache results, re-processing only modified files to maintain performance on large monorepos.

### Phase 3: Architectural Layering (architecture-analyzer)

The **architecture-analyzer** consumes the complete node and edge set to compute logical layers. As implemented in [architecture-analyzer.md lines 14-66](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/architecture-analyzer.md#L14-L66), the agent:

- Groups files by directory topology and cross-reference density.
- Calculates adjacency matrices and import-directionality signals.
- Assigns every file node to exactly one layer (e.g., `layer:api`, `layer:service`, `layer:data`).

The result is persisted to [`layers.json`](https://github.com/Lum1104/Understand-Anything/blob/main/layers.json) for downstream consumption.

### Phase 4: Guided Tour Generation (tour-builder)

The **tour-builder** synthesizes a pedagogical entry point into the codebase. Using topology metrics (fan-in/fan-out, BFS traversal order, coupling clusters) detailed in [tour-builder.md lines 12-38](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/tour-builder.md#L12-L38), it constructs a 5-15 step tour that progresses from high-level entry points (like [`main.py`](https://github.com/Lum1104/Understand-Anything/blob/main/main.py) or [`index.js`](https://github.com/Lum1104/Understand-Anything/blob/main/index.js)) through critical business logic to infrastructure details.

### Phase 5: Graph Review and Final Assembly

The **graph-reviewer** performs integrity checks—detecting orphaned nodes or dangling edges—before optionally invoking an LLM for semantic consistency validation. Finally, the orchestrator merges all intermediate artefacts:

- [`ua-scan-files.json`](https://github.com/Lum1104/Understand-Anything/blob/main/ua-scan-files.json) (project metadata)
- `batch-*.json` (structural nodes/edges)
- [`layers.json`](https://github.com/Lum1104/Understand-Anything/blob/main/layers.json) (architectural grouping)
- [`tour.json`](https://github.com/Lum1104/Understand-Anything/blob/main/tour.json) (guided paths)

Into a single [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json) consumed by the interactive dashboard.

## CLI Usage and Output Artifacts

Run the complete multi-agent pipeline from the root of any repository:

```bash

# Trigger the full analysis sequence

/understand --full

# Optional: Include deep LLM review

/understand --full --review

```

The command produces the following key artefacts in `.understand-anything/`:

```text
.understand-anything/
├── tmp/
│   ├── ua-scan-files.json          # project-scanner output

│   └── ua-import-map-output.json   # import resolution map

├── intermediate/
│   ├── batch-0.json                # file-analyzer batch 0

│   ├── batch-1.json                # file-analyzer batch 1

│   ├── layers.json                 # architecture-analyzer output

│   └── tour.json                   # tour-builder output

└── knowledge-graph.json            # Final merged graph

```

## Key Implementation Files

| Path | Purpose | Link |
|------|---------|------|
| [`README.md`](https://github.com/Lum1104/Understand-Anything/blob/main/README.md) | Multi-agent pipeline overview (lines 83-95) | [View](https://github.com/Lum1104/Understand-Anything/blob/main/README.md#L83-L95) |
| [`understand-anything-plugin/src/understand-chat.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/src/understand-chat.ts) | Orchestrator entry point and agent sequencing | [View](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/src/understand-chat.ts) |
| [`understand-anything-plugin/agents/project-scanner.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/project-scanner.md) | Project discovery agent specification | [View](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/project-scanner.md) |
| [`understand-anything-plugin/agents/file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md) | Structural extraction and enrichment logic | [View](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md) |
| [`understand-anything-plugin/agents/architecture-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/architecture-analyzer.md) | Layer detection and topological analysis | [View](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/architecture-analyzer.md) |
| [`understand-anything-plugin/agents/tour-builder.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/tour-builder.md) | Guided tour generation algorithms | [View](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/tour-builder.md) |
| `understand-anything-plugin/skills/understand/scan-project.mjs` | Deterministic file system scanner | [View](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/scan-project.mjs) |
| `understand-anything-plugin/skills/understand/extract-structure.mjs` | Tree-sitter based AST extractor | [View](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/extract-structure.mjs) |

## Summary

- The **multi-agent pipeline** divides codebase analysis into seven specialized agents to balance deterministic precision with LLM-powered semantic insight.
- **Parallel execution** occurs only during structural analysis (file-analyzer batches), while architectural and tour phases run sequentially to respect data dependencies.
- **Deterministic scripts** (`scan-project.mjs`, `extract-import-map.mjs`, `extract-structure.mjs`) provide reproducible structural data, while LLM agents handle narrative summaries and architectural inference.
- The final output is a single [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) that feeds an interactive dashboard for querying relationships, layers, and guided tours.

## Frequently Asked Questions

### How does the pipeline handle large monorepos with thousands of files?

The **file-analyzer** agent processes files in parallel batches (default 20-30 files per batch, 5 concurrent batches) and supports incremental updates. It caches structural extractions in `batch-<idx>.json` files, comparing file hashes on subsequent runs to skip unchanged files and minimize LLM token consumption.

### What is the difference between the architecture-analyzer and the domain-analyzer?

The **architecture-analyzer** groups files into technical layers (API, service, infrastructure) based on import topology and directory structure. The **domain-analyzer** (triggered by `/understand-domain`) extracts business semantics—identifying domain entities, workflows, and business rules—rather than technical layering.

### Can I run individual agents without executing the full pipeline?

Currently, the orchestrator in [`understand-chat.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-chat.ts) sequences all agents for the `/understand` command. However, the deterministic scripts (`scan-project.mjs`, `extract-structure.mjs`) can be executed independently via CLI for debugging, and intermediate JSON outputs are preserved in `.understand-anything/intermediate/` for inspection between stages.

### How does the graph-reviewer ensure data quality?

The **graph-reviewer** performs integrity validation by checking for dangling edges (relationships pointing to non-existent nodes) and orphaned structural nodes. When invoked with the `--review` flag, it additionally prompts an LLM to flag semantic inconsistencies—such as summaries that contradict code signatures or architectural layers that violate import directionality.