# How Lum1104/Understand-Anything Orchestrates Its Multi-Agent Pipeline

> Discover how Lum1104/Understand-Anything orchestrates its five-stage multi-agent pipeline to transform codebases into knowledge graphs using sequential agent dispatching and JSON state passing.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: internals
- Published: 2026-06-06

---

**The Understand-Anything plugin implements a deterministic, five-stage multi-agent pipeline orchestrated by the `/understand` skill that transforms raw codebases into interactive knowledge graphs through sequential agent dispatching and JSON-based state passing.**

The **Lum1104/Understand-Anything** repository automates codebase comprehension through a sophisticated multi-agent pipeline. Unlike single-shot analysis tools, this system decomposes the problem into specialized agents that scan, analyze, architect, and review code in strict dependency order. Each agent produces typed JSON artifacts that serve as the deterministic input for the next stage, eliminating prompt-based state pollution.

## The Five-Agent Architecture

The pipeline consists of five tightly-coupled agents defined in separate markdown files under `understand-anything-plugin/agents/`. Each agent handles a specific transformation from raw files to final knowledge graph.

### Project-Scanner

The **Project-Scanner** ([[`agents/project-scanner.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/project-scanner.md)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/project-scanner.md)) performs the initial filesystem inventory. It detects every file, identifies languages and frameworks, calculates line counts, and builds an `importMap` of module dependencies. The agent writes its output to [`.understand-anything/intermediate/scan-result.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/intermediate/scan-result.json), which includes the project name, description, file list, complexity metrics, and the initial import topology.

### File-Analyzer

The **File-Analyzer** ([[`agents/file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/file-analyzer.md)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md)) processes code in batches of approximately 20–30 files. It executes a deterministic **tree-sitter** extraction script (`extract-structure.mjs`) to generate semantic nodes and edges representing functions, classes, and call relationships. Each batch produces a separate JSON file that gets merged into a raw knowledge graph after all parallel batches complete.

### Architecture-Analyzer

The **Architecture-Analyzer** ([[`agents/architecture-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/architecture-analyzer.md)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/architecture-analyzer.md)) consumes the full node and edge set to discover logical architectural layers. By analyzing import densities and directory groupings, it assigns each file to a specific layer—such as API, Service, Data, or UI—and outputs [`layers.json`](https://github.com/Lum1104/Understand-Anything/blob/main/layers.json) with these classifications.

### Tour-Builder

The **Tour-Builder** ([[`agents/tour-builder.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/tour-builder.md)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/tour-builder.md)) generates a pedagogical guided tour through the codebase. It computes entry-point scores, fan-in/out rankings, and tight-coupling clusters to create an ordered sequence of 5–15 steps. The resulting [`tour.json`](https://github.com/Lum1104/Understand-Anything/blob/main/tour.json) includes step titles, descriptions, and referenced node IDs to guide newcomers through critical dependency chains.

### Graph-Reviewer

The **Graph-Reviewer** ([[`agents/graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/graph-reviewer.md)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/graph-reviewer.md)) performs final sanitization. It deduplicates edges, validates consistency, and produces the canonical [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) consumed by the dashboard UI. This agent ensures the output graph meets structural constraints before visualization.

## Pipeline Execution Flow

The orchestration logic resides in [[`skills/understand/SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/skills/understand/SKILL.md)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md), which implements a chat-based dispatch protocol through the `buildChatPrompt` function.

1. **Phase 1 – SCAN**: The skill launches the Project-Scanner as a sub-agent and waits for [`scan-result.json`](https://github.com/Lum1104/Understand-Anything/blob/main/scan-result.json) to appear in the intermediate directory.

2. **Phase 2 – ANALYZE**: The skill reads `$FILE_LIST` and `$IMPORT_MAP` from the scan results, batches the file list into groups of roughly 25 files, and dispatches up to five File-Analyzer agents in parallel. Each batch receives its file list via environment variables and writes structural JSON back to the intermediate folder.

3. **Phase 3 – ARCHITECT**: Once all batches complete, the accumulated node and edge data is assembled into a raw graph. The Architecture-Analyzer ingests this graph and produces [`layers.json`](https://github.com/Lum1104/Understand-Anything/blob/main/layers.json) with computed architectural assignments.

4. **Phase 4 – TOUR**: The Tour-Builder consumes the nodes, edges, and layer classifications to calculate traversal paths and output [`tour.json`](https://github.com/Lum1104/Understand-Anything/blob/main/tour.json).

5. **Phase 5 – REVIEW**: Finally, the Graph-Reviewer validates the complete dataset and writes the finalized [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) to the intermediate directory.

## State Management and Determinism

The multi-agent pipeline maintains strict separation of concerns through file-based state management. All intermediate artifacts live under `.understand-anything/intermediate/` and pass between agents via **JSON files** and **environment variables**, never through ad-hoc LLM prompts.

This architecture guarantees determinism: the only non-deterministic step is the LLM-driven semantic enrichment (such as function summarization), which operates on top of the deterministic tree-sitter extraction results. The pipeline respects `.understandignore` patterns during the initial scan to exclude noise from analysis.

Supporting utilities in [[`src/understand-chat.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/understand-chat.ts)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/src/understand-chat.ts) and [[`src/context-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/context-builder.ts)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/src/context-builder.ts) format the knowledge graph context into prompts for the LLM, while [[`src/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/index.ts)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/src/index.ts) exports the plugin entry point for Claude Code integration.

## Running the Pipeline Locally

Execute the full multi-agent pipeline from the repository root using pnpm with Node.js version 22 or higher:

```bash

# Install dependencies

pnpm install

# Build core libraries and the skill plugin

pnpm --filter @understand-anything/core build
pnpm --filter @understand-anything/skill build

# Trigger the pipeline via Claude Code

claude-code run /understand --full

```

Alternatively, invoke the skill directly through the dashboard UI by clicking "Run /understand". Upon completion, the system outputs a summary:

```

Project: my-app (Node.js web service)
Total files: 42 (code 35, config 5, docs 2)
Languages: javascript, typescript, yaml, markdown
Estimated complexity: moderate

```

The dashboard then loads [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) to render the interactive graph and displays the tour sidebar generated by the Tour-Builder.

## Summary

- The **multi-agent pipeline** in Understand-Anything uses five specialized agents to transform code into knowledge graphs.
- **SKILL.md** orchestrates execution by dispatching agents sequentially and managing batch parallelism for the File-Analyzer stage.
- **State passes through JSON files** (scan-result.json, layers.json, tour.json, knowledge-graph.json) stored in `.understand-anything/intermediate/`, ensuring reproducible runs.
- **Tree-sitter extraction** provides deterministic structural analysis, while LLM enrichment handles semantic summarization.
- The pipeline supports **concurrent batch processing** (up to 5 File-Analyzer instances) to accelerate large codebase analysis.

## Frequently Asked Questions

### How does the pipeline handle large codebases with thousands of files?

The **File-Analyzer** agent processes files in batches of approximately 20–30 files, and the orchestration layer in **SKILL.md** dispatches up to five concurrent analyzer instances. This parallelization prevents memory pressure while maintaining deterministic output through batched JSON merging.

### What makes this multi-agent pipeline deterministic?

Only the semantic enrichment steps (such as generating human-readable summaries) utilize non-deterministic LLM calls. The structural foundation relies on **tree-sitter** extraction via `extract-structure.mjs` and strict JSON-based state passing between agents, ensuring that file discovery and dependency mapping produce identical results across runs.

### Can I customize which files the Project-Scanner analyzes?

Yes. The **Project-Scanner** respects `.understandignore` files in the repository root, allowing you to exclude directories like `node_modules`, test fixtures, or generated artifacts from the initial scan and subsequent analysis phases.

### Where does the final knowledge graph get consumed?

The **Graph-Reviewer** writes the canonical output to [`.understand-anything/intermediate/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/intermediate/knowledge-graph.json). The dashboard UI loads this file directly to render the interactive graph, while the **Tour-Builder**'s [`tour.json`](https://github.com/Lum1104/Understand-Anything/blob/main/tour.json) populates the guided navigation sidebar.