# What Happens During the ANALYZE Phase in Understand Anything

> Discover the ANALYZE phase in Understand Anything. This phase generates knowledge graph nodes and edges, validates naming conventions, and assembles normalized results.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: internals
- Published: 2026-06-08

---

**The ANALYZE phase (Phase 2) executes parallel file-analyzer sub-agents on pre-computed file batches to generate knowledge graph nodes and edges, strictly validates output naming conventions, and merges individual batch results into a normalized assembled graph.**

The ANALYZE phase serves as the computational engine of the Understand-Anything tool's `/understand` workflow. Following the batching logic from Phase 1.5, this phase transforms raw source code into structured graph data through concurrent LLM-powered analysis. Whether processing an entire codebase or executing an incremental update, the phase enforces rigorous validation to ensure graph integrity.

## Batch Initialization and Progress Reporting

The phase begins by loading the [`batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batches.json) file generated by the [`compute-batches.mjs`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs) script. This file contains an array of batch descriptors, each specifying the files to analyze, pre-computed import data, and a neighbor-map that provides cross-batch edge hints for relationship resolution ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3001](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3001)).

Immediately upon loading, the system reports execution status to maintain user visibility:

```bash
[Phase 2/7] Analyzing files — <totalFiles> files in <totalBatches> batches (up to 5 concurrent)...

```

This progress indicator confirms the workload distribution before sub-agent execution begins ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3007](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3007)).

## Parallel Sub-Agent Execution

### Dispatching File Analyzers

For each batch defined in [`batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batches.json), the skill launches a **file-analyzer** sub-agent defined in [[`agents/file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/agents/file-analyzer.md)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md). To respect LLM token limits and prevent rate throttling, the system constrains concurrency to **five simultaneous sub-agents** ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3010](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3010)).

The sub-agent prompt includes critical context to ensure accurate graph generation:

- **Project metadata**: Repository name, description, and detected programming languages
- **Language-directive**: Language-specific generation instructions
- **Import data and neighbor map**: Cross-reference data enabling the LLM to boost edges linking to symbols in other batches
- **File descriptors**: Array containing `path`, `language`, `sizeLines`, and `fileCategory` for each source file in the batch

### Output Naming Requirements

Each sub-agent must adhere to a strict file naming convention when writing results. Valid output filenames follow the pattern `batch-<index>.json` for single-file outputs or `batch-<index>-part-<k>.json` when a batch is split across multiple files ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3044](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3044)).

The merge script specifically searches for files matching these patterns; any deviation causes the batch data to be silently excluded from the final graph.

### Result Verification

Upon sub-agent completion, the skill performs immediate validation to ensure data integrity. If a batch's expected output file is missing—whether due to sub-agent failure or incorrect naming—the phase aborts with an error rather than proceeding with incomplete data ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3044](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3044)).

Successful completion triggers a confirmation message:

```bash
Phase 2 complete. All <totalBatches> batches analyzed.

```

([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3045](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3045))

## Graph Merging and Normalization

Following successful analysis of all batches, the phase executes the [[`merge-batch-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-batch-graphs.py)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/merge-batch-graphs.py) script to consolidate individual batch outputs into a unified knowledge graph ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3047](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3047)).

The merge process performs several critical operations:

1. **Aggregation**: Reads all `batch-*.json` files (including split parts) and combines nodes and edges
2. **ID Normalization**: Standardizes entity identifiers across batch boundaries to eliminate duplication
3. **Deduplication**: Removes redundant nodes and edges introduced by parallel processing
4. **Reference Correction**: Fixes dangling references and performs a two-pass "tested_by" edge correction to maintain relationship integrity
5. **Logging**: Outputs adjustment details to STDERR for auditability and debugging

The final output is written to [`assembled-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/assembled-graph.json), representing the complete code knowledge graph.

## Incremental Analysis Workflow

When executed in incremental mode, the ANALYZE phase optimizes performance by processing only modified files. The workflow deviates from full analysis through the following sequence ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3067](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3067)):

1. **Change Detection**: Generates [`changed-files.txt`](https://github.com/Lum1104/Understand-Anything/blob/main/changed-files.txt) using `git diff` to identify modified source files
2. **Selective Batching**: Re-runs `compute-batches.mjs` with the `--changed-files` flag to produce batches containing only changed files
3. **Cleanup**: Removes existing nodes and edges belonging to changed files from the previous graph state
4. **Placeholder Creation**: Writes a [`batch-existing.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batch-existing.json) placeholder representing the unchanged graph components
5. **Unified Merge**: Combines new batch outputs with the existing graph placeholder to produce an updated [`assembled-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/assembled-graph.json)

This approach ensures that large codebases can be updated efficiently without re-analyzing unchanged files.

## Summary

- The **ANALYZE phase** processes file batches generated by Phase 1.5 through up to five concurrent file-analyzer sub-agents.
- **Strict output naming** (`batch-<index>.json` or `batch-<index>-part-<k>.json`) is enforced; missing outputs trigger immediate aborts rather than silent failures.
- The phase utilizes **neighbor maps** and **import data** from [`batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batches.json) to enable cross-batch relationship resolution.
- **Normalization** occurs via [`merge-batch-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-batch-graphs.py), which deduplicates entities, fixes dangling references, and performs two-pass edge correction.
- **Incremental mode** supports efficient updates by analyzing only changed files while preserving existing graph structure through placeholder merging.

## Frequently Asked Questions

### What is the maximum number of concurrent sub-agents in the ANALYZE phase?

The ANALYZE phase limits concurrency to **five sub-agents** running simultaneously. This constraint prevents LLM token limit breaches and ensures stable API rate limiting during large-scale codebase analysis ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3010](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3010)).

### How does the tool handle missing batch outputs?

If a sub-agent fails to produce its expected output file (whether through crash, timeout, or incorrect naming), the phase **immediately aborts** with an error status. This strict validation prevents the merge step from silently losing batch data and ensures graph completeness ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3044](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3044)).

### What information is included in the sub-agent prompt?

Each file-analyzer sub-agent receives comprehensive context including project metadata (name, description, languages), language-specific generation directives, the batch's pre-computed import data and neighbor map for cross-batch edge hints, and detailed file descriptors containing path, language, sizeLines, and fileCategory ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3010](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3010)).

### How does incremental analysis differ from full analysis?

Incremental analysis generates a [`changed-files.txt`](https://github.com/Lum1104/Understand-Anything/blob/main/changed-files.txt) list via git diff, re-runs the batching script with `--changed-files` to target only modified files, removes stale nodes/edges from the existing graph, and merges new results with a [`batch-existing.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batch-existing.json) placeholder. This avoids re-processing unchanged files while maintaining graph consistency ([[`SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/SKILL.md) Line 3067](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3067)).