What Is the Role of the File-Analyzer Agent in Understand-Anything?
The file-analyzer agent in Understand-Anything is the second-stage pipeline worker that converts raw source code into structured knowledge-graph nodes and edges by performing deterministic structural extraction and LLM-powered semantic analysis on batched files.
In the Lum1104/Understand-Anything repository, the file-analyzer agent sits at the heart of the /understand skill's pipeline. After the project-scanner inventories every file and builds an import graph, this agent receives batched files and translates them into the structured elements that power the final knowledge graph. Its output serves as the backbone for downstream components such as the architecture-analyzer and tour-builder.
Where the File-Analyzer Agent Sits in the Pipeline
The /understand skill executes as a multi-phase pipeline. The file-analyzer agent takes over immediately after file collection and batching are complete.
- project-scanner produces
scan-result.json, which contains the full file list and import graph. - compute-batches groups files into batches and writes
batches.json. - file-analyzer processes each batch in parallel, launching up to 5 concurrent sub-agents to generate
batch-*.jsonfiles. - architecture-analyzer consumes the merged graph to infer higher-level components and modules.
- tour-builder and graph-reviewer generate UI-ready tours and validate the final graph.
The file-analyzer is therefore the engine that translates raw source code into the structured nodes and edges that form the backbone of the knowledge graph.
Core Responsibilities of the File-Analyzer Agent
The agent handles every batch through four tightly coupled responsibilities.
Phase 1: Structural Extraction via extract-structure.mjs
The agent executes the bundled extract-structure.mjs script, located under the skill directory, on every file in its assigned batch. This tree-sitter-based script deterministically extracts raw AST information: functions, classes, exported symbols, line counts, and other language-aware metadata.
This structural pass provides a deterministic, language-aware view of each file without relying on ad-hoc regex scripts.
Cross-Batch Context with neighborMap
The dispatch payload includes a neighborMap that carries symbols from files assigned to other batches. The agent uses this map to emit confident cross-batch edges such as calls, inherits, implements, and related.
This ensures the final graph captures relationships that span batch boundaries, preventing a fragmented view of the codebase.
Phase 2: Semantic Analysis and Enrichment
After structural extraction, the LLM consumes the extraction results and enriches them. It adds human-readable summaries, tags, complexity scores, and semantic edges that go beyond purely syntactic information.
This phase supplies intent, purpose, and high-level architectural insight that raw AST data alone cannot convey.
Output Files and Naming Conventions
The agent writes exactly one JSON file per batch, following the strict convention batch-<index>.json or batch-<index>-part-<k>.json. The downstream merge step, implemented in scripts/merge-batch-graphs.py, globs all batch-*.json files and aggregates them into a single unified knowledge graph.
Dispatching and Input Schema
The orchestration logic that launches file-analyzer sub-agents is defined in skills/understand/SKILL.md.
Sub-Agent Dispatch Configuration
For every entry in batches.json, the skill dispatches a file-analyzer instance with concurrency capped at five:
# For each batch in batches.json
- name: file-analyzer
description: Analyze a batch of source files
input:
projectRoot: "{{projectRoot}}"
batchFiles: "{{batch.files}}"
batchImportData: "{{batch.importData}}"
neighborMap: "{{batch.neighborMap}}"
output: "batch-{{batch.index}}.json"
concurrency: 5
The output field enforces the naming contract that merge-batch-graphs.py depends on.
Input Payload Preparation
Before the agent runs, the system serializes the batch context to a temporary JSON file. The schema is documented in agents/file-analyzer.md and includes projectRoot, batchFiles, batchImportData, and neighborMap:
cat > $PROJECT_ROOT/.understand-anything/tmp/ua-file-analyzer-input-<batchIndex>.json << 'ENDJSON'
{
"projectRoot": "<project-root>",
"batchFiles": [
{"path":"src/index.ts","language":"typescript","sizeLines":150,"fileCategory":"code"}
],
"batchImportData": <batchImportData JSON>,
"neighborMap": <neighborMap JSON>
}
ENDJSON
Executing the Bundled Extraction Script
In Phase 1, the agent invokes the bundled tree-sitter script by passing the prepared input path and a results output path:
node <SKILL_DIR>/extract-structure.mjs \
$PROJECT_ROOT/.understand-anything/tmp/ua-file-analyzer-input-<batchIndex>.json \
$PROJECT_ROOT/.understand-anything/tmp/ua-file-extract-results-<batchIndex>.json
The script reads the batched file list from the first argument and writes raw AST data to the second argument.
Sample Batch Output
After Phase 2 enrichment, the agent emits a JSON document containing nodes and edges for that batch. A minimal fragment looks like this:
{
"nodes": [
{"type":"file","id":"src/index.ts"},
{"type":"function","id":"src/index.ts:main","name":"main"}
],
"edges": [
{"source":"src/index.ts:main","target":"src/utils.ts:parse","type":"calls"}
]
}
Production outputs include additional LLM-generated metadata such as summaries, tags, and complexity scores.
Key Implementation Files
The behavior of the file-analyzer agent is defined across five primary files in the Lum1104/Understand-Anything source tree:
agents/file-analyzer.md— Defines the agent, its input schema, and step-by-step execution logic.skills/understand/SKILL.md— Orchestrates batching and dispatches up to five concurrent file-analyzer sub-agents.skills/understand/file-analyzer-prompt.md— Provides the prompt template, including language and framework hints.skills/understand/extract-structure.mjs— The bundled tree-sitter script that performs deterministic structural extraction.scripts/merge-batch-graphs.py— Aggregates allbatch-*.jsonoutputs into the final knowledge graph.
Summary
- The file-analyzer agent is the second-stage worker in the Understand-Anything
/understandpipeline. - It processes batches of files in parallel, with a maximum concurrency of five sub-agents.
- Phase 1 runs the
extract-structure.mjstree-sitter script to extract deterministic AST data. - Phase 2 uses an LLM to add semantic summaries, tags, complexity scores, and cross-batch edges via the neighborMap.
- Output follows a strict
batch-<index>.jsonnaming convention so thatscripts/merge-batch-graphs.pycan reliably assemble the final graph.
Frequently Asked Questions
What is the role of the file-analyzer agent in Understand-Anything?
The file-analyzer agent transforms raw source code into structured knowledge-graph elements. It acts as the bridge between the initial project scan and the final merged graph by extracting both syntactic structure and semantic meaning from batched files.
How does the file-analyzer agent handle relationships that cross batch boundaries?
The agent receives a neighborMap in its input payload that contains symbols from files in other batches. It uses this map to emit edges such as calls, inherits, implements, and related, ensuring the merged graph remains internally consistent.
What is the difference between Phase 1 and Phase 2 of the file-analyzer agent?
Phase 1 executes the deterministic, tree-sitter-based extract-structure.mjs script to capture raw AST facts. Phase 2 feeds those facts to an LLM, which generates human-readable summaries, complexity scores, and semantic edges that express architectural intent rather than pure syntax.
Which source files define the file-analyzer agent and its execution steps?
The agent definition lives in agents/file-analyzer.md, orchestration logic resides in skills/understand/SKILL.md, prompt templates are stored in skills/understand/file-analyzer-prompt.md, structural extraction is handled by skills/understand/extract-structure.mjs, and final graph merging is performed by scripts/merge-batch-graphs.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →