How the Multi-Agent Pipeline in Understand-Anything Orchestrates Codebase Analysis
The multi-agent pipeline in Understand-Anything sequentially executes seven specialized agents—starting with project discovery and ending with graph validation—to transform raw code into an interactive knowledge graph via the /understand command.
The open-source tool Lum1104/Understand-Anything automates codebase comprehension through a deterministic, stage-gated workflow. By combining deterministic static analysis scripts with LLM-powered semantic enrichment, the multi-agent pipeline breaks repository analysis into discrete, scalable phases that generate a queryable knowledge graph.
The Seven Specialized Agents
The pipeline delegates work to focused agents, each defined in dedicated markdown specifications under understand-anything-plugin/agents/:
- project-scanner – Discovers file structure, detects languages, and builds import maps using deterministic scripts.
- file-analyzer – Extracts structural nodes (functions, classes) and edges via Tree-sitter, then enriches them with LLM summaries.
- architecture-analyzer – Groups files into logical layers (API, service, infrastructure) based on import topology and semantic analysis.
- tour-builder – Generates a pedagogical walkthrough of the codebase by ranking entry points and critical paths.
- graph-reviewer – Validates graph integrity, removes dangling edges, and optionally triggers a full LLM consistency review.
- domain-analyzer (optional) – Extracts business-domain concepts and workflows when invoked via
/understand-domain. - article-analyzer (optional) – Parses wiki-style knowledge bases into entities and relationships for
/understand-knowledge.
The high-level architecture is documented in README.md lines 83-95.
Step-by-Step Orchestration Flow
When you run /understand, the orchestrator in understand-chat.ts triggers a five-phase execution sequence. File analysis runs in parallel batches, while architectural and tour phases run sequentially to ensure data dependencies are met.
Phase 1: Project Discovery (project-scanner)
The project-scanner agent executes three deterministic sub-steps:
- Narrative Extraction – Reads
README.md,package.json, and configuration files to extract project context via LLM. - File Enumeration – Executes
scan-project.mjsto producefiles[]with language detection, categorization, and line counts. - Import Mapping – Runs
extract-import-map.mjsto generate a deterministic cross-language import graph.
Outputs are written to .understand-anything/tmp/ua-scan-files.json and ua-import-map-output.json as specified in project-scanner.md lines 51-78.
Phase 2: Structural Analysis (file-analyzer)
The file-analyzer processes the discovered files in parallel batches (default: ~20-30 files per batch, up to 5 concurrent). For each batch:
- It invokes
extract-structure.mjsto parse ASTs and extract raw structural data. - It enriches each node with LLM-generated summaries, complexity metrics, and semantic tags.
- It emits
batch-<idx>.jsonfragments containingnodes[]andedges[]arrays.
Incremental runs cache results, re-processing only modified files to maintain performance on large monorepos.
Phase 3: Architectural Layering (architecture-analyzer)
The architecture-analyzer consumes the complete node and edge set to compute logical layers. As implemented in architecture-analyzer.md lines 14-66, the agent:
- Groups files by directory topology and cross-reference density.
- Calculates adjacency matrices and import-directionality signals.
- Assigns every file node to exactly one layer (e.g.,
layer:api,layer:service,layer:data).
The result is persisted to layers.json for downstream consumption.
Phase 4: Guided Tour Generation (tour-builder)
The tour-builder synthesizes a pedagogical entry point into the codebase. Using topology metrics (fan-in/fan-out, BFS traversal order, coupling clusters) detailed in tour-builder.md lines 12-38, it constructs a 5-15 step tour that progresses from high-level entry points (like main.py or index.js) through critical business logic to infrastructure details.
Phase 5: Graph Review and Final Assembly
The graph-reviewer performs integrity checks—detecting orphaned nodes or dangling edges—before optionally invoking an LLM for semantic consistency validation. Finally, the orchestrator merges all intermediate artefacts:
ua-scan-files.json(project metadata)batch-*.json(structural nodes/edges)layers.json(architectural grouping)tour.json(guided paths)
Into a single .understand-anything/knowledge-graph.json consumed by the interactive dashboard.
CLI Usage and Output Artifacts
Run the complete multi-agent pipeline from the root of any repository:
# Trigger the full analysis sequence
/understand --full
# Optional: Include deep LLM review
/understand --full --review
The command produces the following key artefacts in .understand-anything/:
.understand-anything/
├── tmp/
│ ├── ua-scan-files.json # project-scanner output
│ └── ua-import-map-output.json # import resolution map
├── intermediate/
│ ├── batch-0.json # file-analyzer batch 0
│ ├── batch-1.json # file-analyzer batch 1
│ ├── layers.json # architecture-analyzer output
│ └── tour.json # tour-builder output
└── knowledge-graph.json # Final merged graph
Key Implementation Files
| Path | Purpose | Link |
|---|---|---|
README.md |
Multi-agent pipeline overview (lines 83-95) | View |
understand-anything-plugin/src/understand-chat.ts |
Orchestrator entry point and agent sequencing | View |
understand-anything-plugin/agents/project-scanner.md |
Project discovery agent specification | View |
understand-anything-plugin/agents/file-analyzer.md |
Structural extraction and enrichment logic | View |
understand-anything-plugin/agents/architecture-analyzer.md |
Layer detection and topological analysis | View |
understand-anything-plugin/agents/tour-builder.md |
Guided tour generation algorithms | View |
understand-anything-plugin/skills/understand/scan-project.mjs |
Deterministic file system scanner | View |
understand-anything-plugin/skills/understand/extract-structure.mjs |
Tree-sitter based AST extractor | View |
Summary
- The multi-agent pipeline divides codebase analysis into seven specialized agents to balance deterministic precision with LLM-powered semantic insight.
- Parallel execution occurs only during structural analysis (file-analyzer batches), while architectural and tour phases run sequentially to respect data dependencies.
- Deterministic scripts (
scan-project.mjs,extract-import-map.mjs,extract-structure.mjs) provide reproducible structural data, while LLM agents handle narrative summaries and architectural inference. - The final output is a single
knowledge-graph.jsonthat feeds an interactive dashboard for querying relationships, layers, and guided tours.
Frequently Asked Questions
How does the pipeline handle large monorepos with thousands of files?
The file-analyzer agent processes files in parallel batches (default 20-30 files per batch, 5 concurrent batches) and supports incremental updates. It caches structural extractions in batch-<idx>.json files, comparing file hashes on subsequent runs to skip unchanged files and minimize LLM token consumption.
What is the difference between the architecture-analyzer and the domain-analyzer?
The architecture-analyzer groups files into technical layers (API, service, infrastructure) based on import topology and directory structure. The domain-analyzer (triggered by /understand-domain) extracts business semantics—identifying domain entities, workflows, and business rules—rather than technical layering.
Can I run individual agents without executing the full pipeline?
Currently, the orchestrator in understand-chat.ts sequences all agents for the /understand command. However, the deterministic scripts (scan-project.mjs, extract-structure.mjs) can be executed independently via CLI for debugging, and intermediate JSON outputs are preserved in .understand-anything/intermediate/ for inspection between stages.
How does the graph-reviewer ensure data quality?
The graph-reviewer performs integrity validation by checking for dangling edges (relationships pointing to non-existent nodes) and orphaned structural nodes. When invoked with the --review flag, it additionally prompts an LLM to flag semantic inconsistencies—such as summaries that contradict code signatures or architectural layers that violate import directionality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →