How Lum1104/Understand-Anything Orchestrates Its Multi-Agent Pipeline
The Understand-Anything plugin implements a deterministic, five-stage multi-agent pipeline orchestrated by the /understand skill that transforms raw codebases into interactive knowledge graphs through sequential agent dispatching and JSON-based state passing.
The Lum1104/Understand-Anything repository automates codebase comprehension through a sophisticated multi-agent pipeline. Unlike single-shot analysis tools, this system decomposes the problem into specialized agents that scan, analyze, architect, and review code in strict dependency order. Each agent produces typed JSON artifacts that serve as the deterministic input for the next stage, eliminating prompt-based state pollution.
The Five-Agent Architecture
The pipeline consists of five tightly-coupled agents defined in separate markdown files under understand-anything-plugin/agents/. Each agent handles a specific transformation from raw files to final knowledge graph.
Project-Scanner
The Project-Scanner ([agents/project-scanner.md](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/project-scanner.md)) performs the initial filesystem inventory. It detects every file, identifies languages and frameworks, calculates line counts, and builds an importMap of module dependencies. The agent writes its output to .understand-anything/intermediate/scan-result.json, which includes the project name, description, file list, complexity metrics, and the initial import topology.
File-Analyzer
The File-Analyzer ([agents/file-analyzer.md](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md)) processes code in batches of approximately 20–30 files. It executes a deterministic tree-sitter extraction script (extract-structure.mjs) to generate semantic nodes and edges representing functions, classes, and call relationships. Each batch produces a separate JSON file that gets merged into a raw knowledge graph after all parallel batches complete.
Architecture-Analyzer
The Architecture-Analyzer ([agents/architecture-analyzer.md](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/architecture-analyzer.md)) consumes the full node and edge set to discover logical architectural layers. By analyzing import densities and directory groupings, it assigns each file to a specific layer—such as API, Service, Data, or UI—and outputs layers.json with these classifications.
Tour-Builder
The Tour-Builder ([agents/tour-builder.md](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/tour-builder.md)) generates a pedagogical guided tour through the codebase. It computes entry-point scores, fan-in/out rankings, and tight-coupling clusters to create an ordered sequence of 5–15 steps. The resulting tour.json includes step titles, descriptions, and referenced node IDs to guide newcomers through critical dependency chains.
Graph-Reviewer
The Graph-Reviewer ([agents/graph-reviewer.md](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/graph-reviewer.md)) performs final sanitization. It deduplicates edges, validates consistency, and produces the canonical knowledge-graph.json consumed by the dashboard UI. This agent ensures the output graph meets structural constraints before visualization.
Pipeline Execution Flow
The orchestration logic resides in [skills/understand/SKILL.md](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md), which implements a chat-based dispatch protocol through the buildChatPrompt function.
-
Phase 1 – SCAN: The skill launches the Project-Scanner as a sub-agent and waits for
scan-result.jsonto appear in the intermediate directory. -
Phase 2 – ANALYZE: The skill reads
$FILE_LISTand$IMPORT_MAPfrom the scan results, batches the file list into groups of roughly 25 files, and dispatches up to five File-Analyzer agents in parallel. Each batch receives its file list via environment variables and writes structural JSON back to the intermediate folder. -
Phase 3 – ARCHITECT: Once all batches complete, the accumulated node and edge data is assembled into a raw graph. The Architecture-Analyzer ingests this graph and produces
layers.jsonwith computed architectural assignments. -
Phase 4 – TOUR: The Tour-Builder consumes the nodes, edges, and layer classifications to calculate traversal paths and output
tour.json. -
Phase 5 – REVIEW: Finally, the Graph-Reviewer validates the complete dataset and writes the finalized
knowledge-graph.jsonto the intermediate directory.
State Management and Determinism
The multi-agent pipeline maintains strict separation of concerns through file-based state management. All intermediate artifacts live under .understand-anything/intermediate/ and pass between agents via JSON files and environment variables, never through ad-hoc LLM prompts.
This architecture guarantees determinism: the only non-deterministic step is the LLM-driven semantic enrichment (such as function summarization), which operates on top of the deterministic tree-sitter extraction results. The pipeline respects .understandignore patterns during the initial scan to exclude noise from analysis.
Supporting utilities in [src/understand-chat.ts](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/src/understand-chat.ts) and [src/context-builder.ts](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/src/context-builder.ts) format the knowledge graph context into prompts for the LLM, while [src/index.ts](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/src/index.ts) exports the plugin entry point for Claude Code integration.
Running the Pipeline Locally
Execute the full multi-agent pipeline from the repository root using pnpm with Node.js version 22 or higher:
# Install dependencies
pnpm install
# Build core libraries and the skill plugin
pnpm --filter @understand-anything/core build
pnpm --filter @understand-anything/skill build
# Trigger the pipeline via Claude Code
claude-code run /understand --full
Alternatively, invoke the skill directly through the dashboard UI by clicking "Run /understand". Upon completion, the system outputs a summary:
Project: my-app (Node.js web service)
Total files: 42 (code 35, config 5, docs 2)
Languages: javascript, typescript, yaml, markdown
Estimated complexity: moderate
The dashboard then loads knowledge-graph.json to render the interactive graph and displays the tour sidebar generated by the Tour-Builder.
Summary
- The multi-agent pipeline in Understand-Anything uses five specialized agents to transform code into knowledge graphs.
- SKILL.md orchestrates execution by dispatching agents sequentially and managing batch parallelism for the File-Analyzer stage.
- State passes through JSON files (scan-result.json, layers.json, tour.json, knowledge-graph.json) stored in
.understand-anything/intermediate/, ensuring reproducible runs. - Tree-sitter extraction provides deterministic structural analysis, while LLM enrichment handles semantic summarization.
- The pipeline supports concurrent batch processing (up to 5 File-Analyzer instances) to accelerate large codebase analysis.
Frequently Asked Questions
How does the pipeline handle large codebases with thousands of files?
The File-Analyzer agent processes files in batches of approximately 20–30 files, and the orchestration layer in SKILL.md dispatches up to five concurrent analyzer instances. This parallelization prevents memory pressure while maintaining deterministic output through batched JSON merging.
What makes this multi-agent pipeline deterministic?
Only the semantic enrichment steps (such as generating human-readable summaries) utilize non-deterministic LLM calls. The structural foundation relies on tree-sitter extraction via extract-structure.mjs and strict JSON-based state passing between agents, ensuring that file discovery and dependency mapping produce identical results across runs.
Can I customize which files the Project-Scanner analyzes?
Yes. The Project-Scanner respects .understandignore files in the repository root, allowing you to exclude directories like node_modules, test fixtures, or generated artifacts from the initial scan and subsequent analysis phases.
Where does the final knowledge graph get consumed?
The Graph-Reviewer writes the canonical output to .understand-anything/intermediate/knowledge-graph.json. The dashboard UI loads this file directly to render the interactive graph, while the Tour-Builder's tour.json populates the guided navigation sidebar.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →