# The 7 Phases of the Understand Anything Analysis Pipeline: From Scan to Save

> Explore the 7 phases of the Understand Anything analysis pipeline. Learn how this process transforms code into a queryable knowledge graph from scan to save.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-03

---

**The Understand Anything analysis pipeline consists of seven sequential phases—Scan, Semantic Batching (Phase 1.5), Analyze, Assemble Review, Architecture, Tour, and Final Review with Save—that transform raw source code into a queryable knowledge graph.**

The **Understand Anything** repository by `Lum1104` defines a deterministic, LLM‑augmented pipeline in [`skills/understand/SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/skills/understand/SKILL.md). Each phase exposes user‑visible progress markers like `[Phase X/7]` and maps to a specific agent or script responsible for a discrete transformation of the codebase.

## Phase 1: Scan (Project Discovery)

The pipeline begins with **Phase 1 – SCAN**, which performs comprehensive project discovery. This phase enumerates every file in the repository, detects programming languages and frameworks, and computes a lightweight file inventory.

According to the source code in [`skills/understand/SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/skills/understand/SKILL.md) (line 233), this phase runs the bundled **project‑scanner** agent defined in [`agents/project-scanner.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/project-scanner.md). The scanner outputs a complete manifest of the codebase, which serves as the foundation for downstream analysis.

## Phase 2: Semantic Batching (Phase 1.5)

Immediately following discovery, the pipeline executes **Phase 1.5 – BATCH** for semantic batching. This step groups the scanned files into balanced batches to optimize downstream LLM processing and prevent context window overflow.

The batching logic is implemented in `compute-batches.mjs` (referenced at line 278 in [`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md)), which consumes the file inventory and produces [`batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batches.json). Despite its decimal naming, this constitutes the second phase of the seven‑step pipeline, reported in the CLI as `[Phase 1.5/7]`.

## Phase 3: Analyze (Structure Extraction)

**Phase 2 – ANALYZE** parses each batch using Tree‑sitter or regex parsers to extract structural nodes and edges. This phase transforms raw text into a graph representation of functions, classes, imports, and dependencies.

The **file‑analyzer** agent ([`agents/file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/file-analyzer.md)) handles this work, as defined at line 295 of [`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md). It processes each batch serially, emitting partial graphs that contain the Abstract Syntax Tree (AST) relationships required for later assembly.

## Phase 4: Assemble Review (Graph Stitching)

In **Phase 3 – ASSEMBLE REVIEW**, the pipeline merges the per‑batch sub‑graphs, deduplicates redundant nodes, and runs a fast deterministic validation script. This step ensures the graph is internally consistent before moving to higher‑level analysis.

The **graph‑reviewer** validation script ([`agents/graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/graph-reviewer.md)) executes this logic at line 389. It checks for orphaned edges, duplicate identifiers, and schema compliance without yet invoking the LLM for semantic quality checks.

## Phase 5: Architecture (Layer Detection)

**Phase 4 – ARCHITECTURE** classifies nodes into high‑level layers such as application code, configuration, infrastructure, and CI/CD. This phase generates a top‑level architectural view that helps users navigate large monorepos.

The **architecture‑analyzer** agent ([`agents/architecture-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/architecture-analyzer.md)), referenced at line 416, applies heuristics and LLM reasoning to categorize files and modules into logical layers that appear in the final dashboard visualization.

## Phase 6: Tour (Narrative Generation)

**Phase 5 – TOUR** synthesizes a guided walkthrough that explains the project to human users. This narrative generation phase creates a contextual tour highlighting entry points, data flows, and critical business logic.

Driven by the **tour‑builder** agent ([`agents/tour-builder.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/tour-builder.md)) at line 499, this step produces markdown or structured JSON narratives that are embedded into the knowledge graph metadata for IDE and dashboard consumption.

## Phase 7: Final Review and Save (QA and Persistence)

The final phase combines **Phase 6 – REVIEW** and **Phase 7 – SAVE** into the seventh pipeline step. First, an LLM‑powered reviewer checks the assembled graph for completeness, consistency, and missing edges (line 572, [`agents/graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/graph-reviewer.md)). Then, a deterministic script serializes the in‑memory graph to disk.

The **SAVE** operation writes the final artifact to [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json) (line 734), making the enriched graph consumable by the dashboard and CLI search tools.

## Working with the Pipeline Output

After Phase 7 completes, you can access the knowledge graph programmatically:

```typescript
import { readFile } from 'node:fs/promises';
import { KnowledgeGraph } from '@understand-anything/core';

async function loadGraph() {
  const raw = await readFile('.understand-anything/knowledge-graph.json', 'utf-8');
  const graph: KnowledgeGraph = JSON.parse(raw);
  console.log(`Loaded ${graph.nodes.length} nodes and ${graph.edges.length} edges`);
}
loadGraph();

```

The `KnowledgeGraph` type is exported from [`packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/types.ts). You can query the persisted graph using the browser‑safe core API:

```typescript
import { search } from '@understand-anything/core/search';

const results = await search('authentication', {
  limit: 5,
  nodeTypes: ['function', 'class', 'service']
});

```

## Summary

The Understand Anything analysis pipeline orchestrates seven deterministic phases to convert repositories into knowledge graphs:

- **Phase 1 (Scan)**: Discovers all files and frameworks via [`agents/project-scanner.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/project-scanner.md)
- **Phase 1.5 (Batch)**: Groups files into processing batches using `compute-batches.mjs`
- **Phase 2 (Analyze)**: Extracts AST nodes and edges via [`agents/file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/file-analyzer.md)
- **Phase 3 (Assemble Review)**: Merges sub‑graphs and validates structure
- **Phase 4 (Architecture)**: Classifies components into architectural layers
- **Phase 5 (Tour)**: Generates human‑readable project narratives
- **Phases 6–7 (Review & Save)**: Performs LLM quality assurance and persists to [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json)

Each phase reports progress as `[Phase X/7]`, providing clear visibility into the deterministic transformation process.

## Frequently Asked Questions

### What is Phase 1.5 in the Understand Anything pipeline?

Phase 1.5 is the **semantic batching** step that executes between project scanning (Phase 1) and structural analysis (Phase 2). Implemented in `compute-batches.mjs`, it balances file groups to optimize LLM context windows and is reported in the CLI as `[Phase 1.5/7]`.

### How does the pipeline handle large codebases?

The pipeline utilizes **semantic batching** (Phase 1.5) to partition repositories into manageable chunks. The **ANALYZE** phase (Phase 2) processes these batches independently via the file‑analyzer agent, allowing the system to scale to monorepos without exceeding token limits.

### What file format does the final knowledge graph use?

The graph is serialized as a JSON file named [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) stored in the `.understand-anything/` directory. The schema is defined in [`packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/types.ts) and includes strongly typed `Node` and `Edge` arrays.

### Can I run individual phases separately?

The pipeline is designed to run sequentially as defined in [`skills/understand/SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/skills/understand/SKILL.md), with each phase consuming artifacts from the previous step. While the CLI (`npx understand-anything`) executes all seven phases automatically, the modular agent architecture (`agents/*.md`) theoretically allows individual invocation provided the prerequisite batch files or partial graphs are present.