# The 7 Phases of the Understand-Anything Analysis Pipeline Explained

> Explore the 7 phases of the Understand-Anything analysis pipeline. Discover how raw code transforms into a queryable knowledge graph through SCAN, BATCH, ANALYZE, and more.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-01

---

**The Understand-Anything analysis pipeline executes seven sequential phases—SCAN, BATCH, ANALYZE, ASSEMBLE REVIEW, ARCHITECTURE, TOUR, REVIEW, and SAVE—to transform raw source code into a queryable knowledge graph, with each phase managed by dedicated agents defined in [`skills/understand/SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/skills/understand/SKILL.md).**

The Understand-Anything repository implements a deterministic, LLM-augmented pipeline that converts source-code repositories into enriched knowledge graphs. Defined in [`skills/understand/SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/skills/understand/SKILL.md), this analysis pipeline processes files through seven primary milestones (Phase 1 through Phase 7), including an intermediate semantic batching step designated Phase 1.5. Each phase reports progress via user-visible tokens like `[Phase X/7]` in the CLI output.

## The Seven Phases of Analysis

The pipeline progresses sequentially from file discovery to final persistence, with each phase invoking specific agent scripts responsible for deterministic processing or LLM-powered analysis.

### Phase 1 – SCAN (Project Discovery)

**Phase 1** performs comprehensive project discovery by enumerating every file, detecting languages and frameworks, and computing a lightweight file inventory. This phase executes the bundled **project-scanner** agent defined in [`agents/project-scanner.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/project-scanner.md). The scanner traverses the repository root, identifies relevant source files, and establishes the initial project context required for downstream processing.

### Phase 1.5 – BATCH (Semantic Batching)

Between discovery and analysis, **Phase 1.5** handles semantic batching to optimize downstream LLM context windows. This intermediate phase runs `compute-batches.mjs` to group scanned files into balanced, semantically coherent batches, outputting the results to [`batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batches.json). By grouping related files (such as models and their corresponding tests), the pipeline ensures that Phase 2 analyzes code with proper contextual neighbors.

### Phase 2 – ANALYZE (Structure Extraction)

**Phase 2** parses each batch using Tree-sitter or regex-based parsers to extract structural nodes and edges. The **file-analyzer** agent ([`agents/file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/file-analyzer.md)) processes each batch file-by-file, emitting AST-derived entities (functions, classes, imports) as graph nodes and their relationships (calls, contains, imports) as edges. This phase transforms raw syntax into structured graph fragments.

### Phase 3 – ASSEMBLE REVIEW (Graph Stitching)

**Phase 3** merges per-batch sub-graphs into a unified knowledge graph. The deterministic **graph-reviewer** validation script ([`agents/graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/graph-reviewer.md)) deduplicates nodes, resolves references across batch boundaries, and runs fast structural validation. This stitching process ensures that a function defined in one batch and called in another becomes a single connected entity in the assembled graph.

### Phase 4 – ARCHITECTURE (Layer Detection)

**Phase 4** classifies nodes into high-level architectural layers (code, configuration, infrastructure, CI/CD, etc.) to generate a top-level architecture view. The **architecture-analyzer** agent ([`agents/architecture-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/architecture-analyzer.md)) applies heuristics and LLM reasoning to label components according to their architectural role, creating a stratified view of the system’s organization.

### Phase 5 – TOUR (Narrative Generation)

**Phase 5** synthesizes a guided walkthrough of the codebase. The **tour-builder** agent ([`agents/tour-builder.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/tour-builder.md)) generates a narrative tour that explains the project to human users, highlighting entry points, core abstractions, and data flows. This phase produces human-readable documentation that maps the graph structure to conceptual understanding.

### Phase 6 – REVIEW (Quality Assurance)

**Phase 6** performs LLM-powered quality assurance on the assembled graph. Unlike the deterministic validation in Phase 3, this **graph-reviewer** step ([`agents/graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/graph-reviewer.md)) uses the LLM to check for completeness, consistency, and missing edges across the entire graph. The reviewer identifies orphaned nodes, incomplete call chains, and undocumented public APIs.

### Phase 7 – SAVE (Persist and Expose)

**Phase 7** serializes the finalized knowledge graph to disk. A deterministic script writes [`graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/graph.json) to [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json), making the graph consumable by the dashboard and CLI tools. This persistence layer exposes the enriched data for interactive exploration and downstream integrations.

## Pipeline Execution and Progress Tracking

When executed from the CLI, the Understand-Anything analysis pipeline reports real-time progress using the phase markers defined in [`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md):

```text
$ npx understand-anything /path/to/project --full
[Phase 1/7] Scanning project files...
[Phase 1.5/7] Computing semantic batches...
[Phase 2/7] Analyzing batch #1 (12 files)…
[Phase 2/7] Analyzing batch #2 (13 files)…
[Phase 3/7] Assembling partial graphs…
[Phase 4/7] Detecting architectural layers…
[Phase 5/7] Building a tour of the codebase…
[Phase 6/7] Reviewing the assembled graph…
[Phase 7/7] Saving knowledge-graph.json → .understand-anything/

```

These progress tokens are injected directly from the skill description, providing deterministic feedback as each agent completes its responsibility.

## Accessing the Generated Knowledge Graph

After Phase 7 completes, the knowledge graph persists as a JSON file that you can load programmatically using the core types defined in [`packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/types.ts):

```typescript
import { readFile } from 'node:fs/promises';
import { KnowledgeGraph } from '@understand-anything/core';

async function loadGraph() {
  const raw = await readFile('.understand-anything/knowledge-graph.json', 'utf-8');
  const graph: KnowledgeGraph = JSON.parse(raw);
  console.log(`Graph contains ${graph.nodes.length} nodes and ${graph.edges.length} edges`);
}
loadGraph();

```

For browser-safe querying without loading the entire file into memory, use the search module from [`packages/core/src/search.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/search.ts):

```typescript
import { search } from '@understand-anything/core/search';

const results = await search('authentication', {
  limit: 5,
  nodeTypes: ['function', 'class', 'service']
});
console.log(results);

```

This search implementation powers the Understand-Anything dashboard’s search bar and supports filtering by node type, relationship depth, and semantic similarity.

## Summary

- **Seven sequential phases** (1 through 7, including the intermediate 1.5 batching step) transform repositories into knowledge graphs via [`skills/understand/SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/skills/understand/SKILL.md).
- **Specialized agents** handle specific concerns: [`project-scanner.md`](https://github.com/Lum1104/Understand-Anything/blob/main/project-scanner.md) for discovery, [`file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/file-analyzer.md) for parsing, [`graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-reviewer.md) for validation and QA, [`architecture-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/architecture-analyzer.md) for layering, and [`tour-builder.md`](https://github.com/Lum1104/Understand-Anything/blob/main/tour-builder.md) for narrative generation.
- **Deterministic and LLM-augmented steps** alternate: Phases 1, 1.5, 3, and 7 use deterministic scripts, while Phases 2, 4, 5, and 6 leverage LLM reasoning.
- **Final output** persists to [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json) with TypeScript types exported from [`packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/types.ts).

## Frequently Asked Questions

### Why is there a Phase 1.5 in the Understand-Anything pipeline?

Phase 1.5 represents the **semantic batching** step that occurs between scanning and analysis. While the pipeline officially counts seven phases (1–7) for progress reporting, Phase 1.5 is displayed as `[Phase 1.5/7]` because it executes between Phase 1 and Phase 2. This step runs `compute-batches.mjs` to group files into optimal batches for LLM context windows, ensuring Phase 2 receives semantically related files together.

### How does Phase 3 (Assemble Review) differ from Phase 6 (Review)?

**Phase 3** performs deterministic graph stitching using the [`graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-reviewer.md) agent to merge sub-graphs, deduplicate nodes, and validate structural integrity without LLM calls. **Phase 6** uses the same agent file but invokes the LLM for qualitative quality assurance, checking for logical completeness, missing documentation, and inconsistent relationships across the entire assembled graph.

### What file format does the final knowledge graph use?

The pipeline outputs a JSON file to [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json). This file conforms to the `KnowledgeGraph` TypeScript interface defined in [`packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/types.ts), containing arrays of nodes (entities like functions, classes, and files) and edges (relationships like calls, imports, and containment).

### Can I run individual phases of the Understand-Anything pipeline independently?

The pipeline is designed to execute sequentially from Phase 1 through Phase 7 as defined in [`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md). While the CLI entry point orchestrates the full flow, individual agent scripts ([`agents/file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/file-analyzer.md), [`agents/architecture-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/architecture-analyzer.md), etc.) can be invoked in isolation during development or debugging. However, phases depend on artifacts from previous steps (e.g., Phase 2 requires [`batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batches.json) from Phase 1.5), so running mid-pipeline phases requires manually preparing those inputs.