The 7 Phases of the Understand-Anything Analysis Pipeline Explained
The Understand-Anything analysis pipeline executes seven sequential phases—SCAN, BATCH, ANALYZE, ASSEMBLE REVIEW, ARCHITECTURE, TOUR, REVIEW, and SAVE—to transform raw source code into a queryable knowledge graph, with each phase managed by dedicated agents defined in skills/understand/SKILL.md.
The Understand-Anything repository implements a deterministic, LLM-augmented pipeline that converts source-code repositories into enriched knowledge graphs. Defined in skills/understand/SKILL.md, this analysis pipeline processes files through seven primary milestones (Phase 1 through Phase 7), including an intermediate semantic batching step designated Phase 1.5. Each phase reports progress via user-visible tokens like [Phase X/7] in the CLI output.
The Seven Phases of Analysis
The pipeline progresses sequentially from file discovery to final persistence, with each phase invoking specific agent scripts responsible for deterministic processing or LLM-powered analysis.
Phase 1 – SCAN (Project Discovery)
Phase 1 performs comprehensive project discovery by enumerating every file, detecting languages and frameworks, and computing a lightweight file inventory. This phase executes the bundled project-scanner agent defined in agents/project-scanner.md. The scanner traverses the repository root, identifies relevant source files, and establishes the initial project context required for downstream processing.
Phase 1.5 – BATCH (Semantic Batching)
Between discovery and analysis, Phase 1.5 handles semantic batching to optimize downstream LLM context windows. This intermediate phase runs compute-batches.mjs to group scanned files into balanced, semantically coherent batches, outputting the results to batches.json. By grouping related files (such as models and their corresponding tests), the pipeline ensures that Phase 2 analyzes code with proper contextual neighbors.
Phase 2 – ANALYZE (Structure Extraction)
Phase 2 parses each batch using Tree-sitter or regex-based parsers to extract structural nodes and edges. The file-analyzer agent (agents/file-analyzer.md) processes each batch file-by-file, emitting AST-derived entities (functions, classes, imports) as graph nodes and their relationships (calls, contains, imports) as edges. This phase transforms raw syntax into structured graph fragments.
Phase 3 – ASSEMBLE REVIEW (Graph Stitching)
Phase 3 merges per-batch sub-graphs into a unified knowledge graph. The deterministic graph-reviewer validation script (agents/graph-reviewer.md) deduplicates nodes, resolves references across batch boundaries, and runs fast structural validation. This stitching process ensures that a function defined in one batch and called in another becomes a single connected entity in the assembled graph.
Phase 4 – ARCHITECTURE (Layer Detection)
Phase 4 classifies nodes into high-level architectural layers (code, configuration, infrastructure, CI/CD, etc.) to generate a top-level architecture view. The architecture-analyzer agent (agents/architecture-analyzer.md) applies heuristics and LLM reasoning to label components according to their architectural role, creating a stratified view of the system’s organization.
Phase 5 – TOUR (Narrative Generation)
Phase 5 synthesizes a guided walkthrough of the codebase. The tour-builder agent (agents/tour-builder.md) generates a narrative tour that explains the project to human users, highlighting entry points, core abstractions, and data flows. This phase produces human-readable documentation that maps the graph structure to conceptual understanding.
Phase 6 – REVIEW (Quality Assurance)
Phase 6 performs LLM-powered quality assurance on the assembled graph. Unlike the deterministic validation in Phase 3, this graph-reviewer step (agents/graph-reviewer.md) uses the LLM to check for completeness, consistency, and missing edges across the entire graph. The reviewer identifies orphaned nodes, incomplete call chains, and undocumented public APIs.
Phase 7 – SAVE (Persist and Expose)
Phase 7 serializes the finalized knowledge graph to disk. A deterministic script writes graph.json to .understand-anything/knowledge-graph.json, making the graph consumable by the dashboard and CLI tools. This persistence layer exposes the enriched data for interactive exploration and downstream integrations.
Pipeline Execution and Progress Tracking
When executed from the CLI, the Understand-Anything analysis pipeline reports real-time progress using the phase markers defined in SKILL.md:
$ npx understand-anything /path/to/project --full
[Phase 1/7] Scanning project files...
[Phase 1.5/7] Computing semantic batches...
[Phase 2/7] Analyzing batch #1 (12 files)…
[Phase 2/7] Analyzing batch #2 (13 files)…
[Phase 3/7] Assembling partial graphs…
[Phase 4/7] Detecting architectural layers…
[Phase 5/7] Building a tour of the codebase…
[Phase 6/7] Reviewing the assembled graph…
[Phase 7/7] Saving knowledge-graph.json → .understand-anything/
These progress tokens are injected directly from the skill description, providing deterministic feedback as each agent completes its responsibility.
Accessing the Generated Knowledge Graph
After Phase 7 completes, the knowledge graph persists as a JSON file that you can load programmatically using the core types defined in packages/core/src/types.ts:
import { readFile } from 'node:fs/promises';
import { KnowledgeGraph } from '@understand-anything/core';
async function loadGraph() {
const raw = await readFile('.understand-anything/knowledge-graph.json', 'utf-8');
const graph: KnowledgeGraph = JSON.parse(raw);
console.log(`Graph contains ${graph.nodes.length} nodes and ${graph.edges.length} edges`);
}
loadGraph();
For browser-safe querying without loading the entire file into memory, use the search module from packages/core/src/search.ts:
import { search } from '@understand-anything/core/search';
const results = await search('authentication', {
limit: 5,
nodeTypes: ['function', 'class', 'service']
});
console.log(results);
This search implementation powers the Understand-Anything dashboard’s search bar and supports filtering by node type, relationship depth, and semantic similarity.
Summary
- Seven sequential phases (1 through 7, including the intermediate 1.5 batching step) transform repositories into knowledge graphs via
skills/understand/SKILL.md. - Specialized agents handle specific concerns:
project-scanner.mdfor discovery,file-analyzer.mdfor parsing,graph-reviewer.mdfor validation and QA,architecture-analyzer.mdfor layering, andtour-builder.mdfor narrative generation. - Deterministic and LLM-augmented steps alternate: Phases 1, 1.5, 3, and 7 use deterministic scripts, while Phases 2, 4, 5, and 6 leverage LLM reasoning.
- Final output persists to
.understand-anything/knowledge-graph.jsonwith TypeScript types exported frompackages/core/src/types.ts.
Frequently Asked Questions
Why is there a Phase 1.5 in the Understand-Anything pipeline?
Phase 1.5 represents the semantic batching step that occurs between scanning and analysis. While the pipeline officially counts seven phases (1–7) for progress reporting, Phase 1.5 is displayed as [Phase 1.5/7] because it executes between Phase 1 and Phase 2. This step runs compute-batches.mjs to group files into optimal batches for LLM context windows, ensuring Phase 2 receives semantically related files together.
How does Phase 3 (Assemble Review) differ from Phase 6 (Review)?
Phase 3 performs deterministic graph stitching using the graph-reviewer.md agent to merge sub-graphs, deduplicate nodes, and validate structural integrity without LLM calls. Phase 6 uses the same agent file but invokes the LLM for qualitative quality assurance, checking for logical completeness, missing documentation, and inconsistent relationships across the entire assembled graph.
What file format does the final knowledge graph use?
The pipeline outputs a JSON file to .understand-anything/knowledge-graph.json. This file conforms to the KnowledgeGraph TypeScript interface defined in packages/core/src/types.ts, containing arrays of nodes (entities like functions, classes, and files) and edges (relationships like calls, imports, and containment).
Can I run individual phases of the Understand-Anything pipeline independently?
The pipeline is designed to execute sequentially from Phase 1 through Phase 7 as defined in SKILL.md. While the CLI entry point orchestrates the full flow, individual agent scripts (agents/file-analyzer.md, agents/architecture-analyzer.md, etc.) can be invoked in isolation during development or debugging. However, phases depend on artifacts from previous steps (e.g., Phase 2 requires batches.json from Phase 1.5), so running mid-pipeline phases requires manually preparing those inputs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →