What Happens During the ANALYZE Phase in Understand Anything
The ANALYZE phase (Phase 2) executes parallel file-analyzer sub-agents on pre-computed file batches to generate knowledge graph nodes and edges, strictly validates output naming conventions, and merges individual batch results into a normalized assembled graph.
The ANALYZE phase serves as the computational engine of the Understand-Anything tool's /understand workflow. Following the batching logic from Phase 1.5, this phase transforms raw source code into structured graph data through concurrent LLM-powered analysis. Whether processing an entire codebase or executing an incremental update, the phase enforces rigorous validation to ensure graph integrity.
Batch Initialization and Progress Reporting
The phase begins by loading the batches.json file generated by the compute-batches.mjs script. This file contains an array of batch descriptors, each specifying the files to analyze, pre-computed import data, and a neighbor-map that provides cross-batch edge hints for relationship resolution ([SKILL.md Line 3001](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3001)).
Immediately upon loading, the system reports execution status to maintain user visibility:
[Phase 2/7] Analyzing files — <totalFiles> files in <totalBatches> batches (up to 5 concurrent)...
This progress indicator confirms the workload distribution before sub-agent execution begins ([SKILL.md Line 3007](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3007)).
Parallel Sub-Agent Execution
Dispatching File Analyzers
For each batch defined in batches.json, the skill launches a file-analyzer sub-agent defined in [agents/file-analyzer.md](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md). To respect LLM token limits and prevent rate throttling, the system constrains concurrency to five simultaneous sub-agents ([SKILL.md Line 3010](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3010)).
The sub-agent prompt includes critical context to ensure accurate graph generation:
- Project metadata: Repository name, description, and detected programming languages
- Language-directive: Language-specific generation instructions
- Import data and neighbor map: Cross-reference data enabling the LLM to boost edges linking to symbols in other batches
- File descriptors: Array containing
path,language,sizeLines, andfileCategoryfor each source file in the batch
Output Naming Requirements
Each sub-agent must adhere to a strict file naming convention when writing results. Valid output filenames follow the pattern batch-<index>.json for single-file outputs or batch-<index>-part-<k>.json when a batch is split across multiple files ([SKILL.md Line 3044](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3044)).
The merge script specifically searches for files matching these patterns; any deviation causes the batch data to be silently excluded from the final graph.
Result Verification
Upon sub-agent completion, the skill performs immediate validation to ensure data integrity. If a batch's expected output file is missing—whether due to sub-agent failure or incorrect naming—the phase aborts with an error rather than proceeding with incomplete data ([SKILL.md Line 3044](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3044)).
Successful completion triggers a confirmation message:
Phase 2 complete. All <totalBatches> batches analyzed.
([SKILL.md Line 3045](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3045))
Graph Merging and Normalization
Following successful analysis of all batches, the phase executes the [merge-batch-graphs.py](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/merge-batch-graphs.py) script to consolidate individual batch outputs into a unified knowledge graph ([SKILL.md Line 3047](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3047)).
The merge process performs several critical operations:
- Aggregation: Reads all
batch-*.jsonfiles (including split parts) and combines nodes and edges - ID Normalization: Standardizes entity identifiers across batch boundaries to eliminate duplication
- Deduplication: Removes redundant nodes and edges introduced by parallel processing
- Reference Correction: Fixes dangling references and performs a two-pass "tested_by" edge correction to maintain relationship integrity
- Logging: Outputs adjustment details to STDERR for auditability and debugging
The final output is written to assembled-graph.json, representing the complete code knowledge graph.
Incremental Analysis Workflow
When executed in incremental mode, the ANALYZE phase optimizes performance by processing only modified files. The workflow deviates from full analysis through the following sequence ([SKILL.md Line 3067](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3067)):
- Change Detection: Generates
changed-files.txtusinggit diffto identify modified source files - Selective Batching: Re-runs
compute-batches.mjswith the--changed-filesflag to produce batches containing only changed files - Cleanup: Removes existing nodes and edges belonging to changed files from the previous graph state
- Placeholder Creation: Writes a
batch-existing.jsonplaceholder representing the unchanged graph components - Unified Merge: Combines new batch outputs with the existing graph placeholder to produce an updated
assembled-graph.json
This approach ensures that large codebases can be updated efficiently without re-analyzing unchanged files.
Summary
- The ANALYZE phase processes file batches generated by Phase 1.5 through up to five concurrent file-analyzer sub-agents.
- Strict output naming (
batch-<index>.jsonorbatch-<index>-part-<k>.json) is enforced; missing outputs trigger immediate aborts rather than silent failures. - The phase utilizes neighbor maps and import data from
batches.jsonto enable cross-batch relationship resolution. - Normalization occurs via
merge-batch-graphs.py, which deduplicates entities, fixes dangling references, and performs two-pass edge correction. - Incremental mode supports efficient updates by analyzing only changed files while preserving existing graph structure through placeholder merging.
Frequently Asked Questions
What is the maximum number of concurrent sub-agents in the ANALYZE phase?
The ANALYZE phase limits concurrency to five sub-agents running simultaneously. This constraint prevents LLM token limit breaches and ensures stable API rate limiting during large-scale codebase analysis ([SKILL.md Line 3010](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3010)).
How does the tool handle missing batch outputs?
If a sub-agent fails to produce its expected output file (whether through crash, timeout, or incorrect naming), the phase immediately aborts with an error status. This strict validation prevents the merge step from silently losing batch data and ensures graph completeness ([SKILL.md Line 3044](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3044)).
What information is included in the sub-agent prompt?
Each file-analyzer sub-agent receives comprehensive context including project metadata (name, description, languages), language-specific generation directives, the batch's pre-computed import data and neighbor map for cross-batch edge hints, and detailed file descriptors containing path, language, sizeLines, and fileCategory ([SKILL.md Line 3010](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3010)).
How does incremental analysis differ from full analysis?
Incremental analysis generates a changed-files.txt list via git diff, re-runs the batching script with --changed-files to target only modified files, removes stale nodes/edges from the existing graph, and merges new results with a batch-existing.json placeholder. This avoids re-processing unchanged files while maintaining graph consistency ([SKILL.md Line 3067](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/SKILL.md#L3067)).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →