The 7 Phases of the Understand Anything Analysis Pipeline: From Scan to Save
The Understand Anything analysis pipeline consists of seven sequential phases—Scan, Semantic Batching (Phase 1.5), Analyze, Assemble Review, Architecture, Tour, and Final Review with Save—that transform raw source code into a queryable knowledge graph.
The Understand Anything repository by Lum1104 defines a deterministic, LLM‑augmented pipeline in skills/understand/SKILL.md. Each phase exposes user‑visible progress markers like [Phase X/7] and maps to a specific agent or script responsible for a discrete transformation of the codebase.
Phase 1: Scan (Project Discovery)
The pipeline begins with Phase 1 – SCAN, which performs comprehensive project discovery. This phase enumerates every file in the repository, detects programming languages and frameworks, and computes a lightweight file inventory.
According to the source code in skills/understand/SKILL.md (line 233), this phase runs the bundled project‑scanner agent defined in agents/project-scanner.md. The scanner outputs a complete manifest of the codebase, which serves as the foundation for downstream analysis.
Phase 2: Semantic Batching (Phase 1.5)
Immediately following discovery, the pipeline executes Phase 1.5 – BATCH for semantic batching. This step groups the scanned files into balanced batches to optimize downstream LLM processing and prevent context window overflow.
The batching logic is implemented in compute-batches.mjs (referenced at line 278 in SKILL.md), which consumes the file inventory and produces batches.json. Despite its decimal naming, this constitutes the second phase of the seven‑step pipeline, reported in the CLI as [Phase 1.5/7].
Phase 3: Analyze (Structure Extraction)
Phase 2 – ANALYZE parses each batch using Tree‑sitter or regex parsers to extract structural nodes and edges. This phase transforms raw text into a graph representation of functions, classes, imports, and dependencies.
The file‑analyzer agent (agents/file-analyzer.md) handles this work, as defined at line 295 of SKILL.md. It processes each batch serially, emitting partial graphs that contain the Abstract Syntax Tree (AST) relationships required for later assembly.
Phase 4: Assemble Review (Graph Stitching)
In Phase 3 – ASSEMBLE REVIEW, the pipeline merges the per‑batch sub‑graphs, deduplicates redundant nodes, and runs a fast deterministic validation script. This step ensures the graph is internally consistent before moving to higher‑level analysis.
The graph‑reviewer validation script (agents/graph-reviewer.md) executes this logic at line 389. It checks for orphaned edges, duplicate identifiers, and schema compliance without yet invoking the LLM for semantic quality checks.
Phase 5: Architecture (Layer Detection)
Phase 4 – ARCHITECTURE classifies nodes into high‑level layers such as application code, configuration, infrastructure, and CI/CD. This phase generates a top‑level architectural view that helps users navigate large monorepos.
The architecture‑analyzer agent (agents/architecture-analyzer.md), referenced at line 416, applies heuristics and LLM reasoning to categorize files and modules into logical layers that appear in the final dashboard visualization.
Phase 6: Tour (Narrative Generation)
Phase 5 – TOUR synthesizes a guided walkthrough that explains the project to human users. This narrative generation phase creates a contextual tour highlighting entry points, data flows, and critical business logic.
Driven by the tour‑builder agent (agents/tour-builder.md) at line 499, this step produces markdown or structured JSON narratives that are embedded into the knowledge graph metadata for IDE and dashboard consumption.
Phase 7: Final Review and Save (QA and Persistence)
The final phase combines Phase 6 – REVIEW and Phase 7 – SAVE into the seventh pipeline step. First, an LLM‑powered reviewer checks the assembled graph for completeness, consistency, and missing edges (line 572, agents/graph-reviewer.md). Then, a deterministic script serializes the in‑memory graph to disk.
The SAVE operation writes the final artifact to .understand-anything/knowledge-graph.json (line 734), making the enriched graph consumable by the dashboard and CLI search tools.
Working with the Pipeline Output
After Phase 7 completes, you can access the knowledge graph programmatically:
import { readFile } from 'node:fs/promises';
import { KnowledgeGraph } from '@understand-anything/core';
async function loadGraph() {
const raw = await readFile('.understand-anything/knowledge-graph.json', 'utf-8');
const graph: KnowledgeGraph = JSON.parse(raw);
console.log(`Loaded ${graph.nodes.length} nodes and ${graph.edges.length} edges`);
}
loadGraph();
The KnowledgeGraph type is exported from packages/core/src/types.ts. You can query the persisted graph using the browser‑safe core API:
import { search } from '@understand-anything/core/search';
const results = await search('authentication', {
limit: 5,
nodeTypes: ['function', 'class', 'service']
});
Summary
The Understand Anything analysis pipeline orchestrates seven deterministic phases to convert repositories into knowledge graphs:
- Phase 1 (Scan): Discovers all files and frameworks via
agents/project-scanner.md - Phase 1.5 (Batch): Groups files into processing batches using
compute-batches.mjs - Phase 2 (Analyze): Extracts AST nodes and edges via
agents/file-analyzer.md - Phase 3 (Assemble Review): Merges sub‑graphs and validates structure
- Phase 4 (Architecture): Classifies components into architectural layers
- Phase 5 (Tour): Generates human‑readable project narratives
- Phases 6–7 (Review & Save): Performs LLM quality assurance and persists to
.understand-anything/knowledge-graph.json
Each phase reports progress as [Phase X/7], providing clear visibility into the deterministic transformation process.
Frequently Asked Questions
What is Phase 1.5 in the Understand Anything pipeline?
Phase 1.5 is the semantic batching step that executes between project scanning (Phase 1) and structural analysis (Phase 2). Implemented in compute-batches.mjs, it balances file groups to optimize LLM context windows and is reported in the CLI as [Phase 1.5/7].
How does the pipeline handle large codebases?
The pipeline utilizes semantic batching (Phase 1.5) to partition repositories into manageable chunks. The ANALYZE phase (Phase 2) processes these batches independently via the file‑analyzer agent, allowing the system to scale to monorepos without exceeding token limits.
What file format does the final knowledge graph use?
The graph is serialized as a JSON file named knowledge-graph.json stored in the .understand-anything/ directory. The schema is defined in packages/core/src/types.ts and includes strongly typed Node and Edge arrays.
Can I run individual phases separately?
The pipeline is designed to run sequentially as defined in skills/understand/SKILL.md, with each phase consuming artifacts from the previous step. While the CLI (npx understand-anything) executes all seven phases automatically, the modular agent architecture (agents/*.md) theoretically allows individual invocation provided the prerequisite batch files or partial graphs are present.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →