What Is the Role of the File-Analyzer Agent in Understand-Anything?

The file-analyzer agent in Understand-Anything is the second-stage pipeline worker that converts raw source code into structured knowledge-graph nodes and edges by performing deterministic structural extraction and LLM-powered semantic analysis on batched files.

In the Lum1104/Understand-Anything repository, the file-analyzer agent sits at the heart of the /understand skill's pipeline. After the project-scanner inventories every file and builds an import graph, this agent receives batched files and translates them into the structured elements that power the final knowledge graph. Its output serves as the backbone for downstream components such as the architecture-analyzer and tour-builder.

Where the File-Analyzer Agent Sits in the Pipeline

The /understand skill executes as a multi-phase pipeline. The file-analyzer agent takes over immediately after file collection and batching are complete.

  1. project-scanner produces scan-result.json, which contains the full file list and import graph.
  2. compute-batches groups files into batches and writes batches.json.
  3. file-analyzer processes each batch in parallel, launching up to 5 concurrent sub-agents to generate batch-*.json files.
  4. architecture-analyzer consumes the merged graph to infer higher-level components and modules.
  5. tour-builder and graph-reviewer generate UI-ready tours and validate the final graph.

The file-analyzer is therefore the engine that translates raw source code into the structured nodes and edges that form the backbone of the knowledge graph.

Core Responsibilities of the File-Analyzer Agent

The agent handles every batch through four tightly coupled responsibilities.

Phase 1: Structural Extraction via extract-structure.mjs

The agent executes the bundled extract-structure.mjs script, located under the skill directory, on every file in its assigned batch. This tree-sitter-based script deterministically extracts raw AST information: functions, classes, exported symbols, line counts, and other language-aware metadata.

This structural pass provides a deterministic, language-aware view of each file without relying on ad-hoc regex scripts.

Cross-Batch Context with neighborMap

The dispatch payload includes a neighborMap that carries symbols from files assigned to other batches. The agent uses this map to emit confident cross-batch edges such as calls, inherits, implements, and related.

This ensures the final graph captures relationships that span batch boundaries, preventing a fragmented view of the codebase.

Phase 2: Semantic Analysis and Enrichment

After structural extraction, the LLM consumes the extraction results and enriches them. It adds human-readable summaries, tags, complexity scores, and semantic edges that go beyond purely syntactic information.

This phase supplies intent, purpose, and high-level architectural insight that raw AST data alone cannot convey.

Output Files and Naming Conventions

The agent writes exactly one JSON file per batch, following the strict convention batch-<index>.json or batch-<index>-part-<k>.json. The downstream merge step, implemented in scripts/merge-batch-graphs.py, globs all batch-*.json files and aggregates them into a single unified knowledge graph.

Dispatching and Input Schema

The orchestration logic that launches file-analyzer sub-agents is defined in skills/understand/SKILL.md.

Sub-Agent Dispatch Configuration

For every entry in batches.json, the skill dispatches a file-analyzer instance with concurrency capped at five:


# For each batch in batches.json

- name: file-analyzer
  description: Analyze a batch of source files
  input:
    projectRoot: "{{projectRoot}}"
    batchFiles: "{{batch.files}}"
    batchImportData: "{{batch.importData}}"
    neighborMap: "{{batch.neighborMap}}"
  output: "batch-{{batch.index}}.json"
  concurrency: 5

The output field enforces the naming contract that merge-batch-graphs.py depends on.

Input Payload Preparation

Before the agent runs, the system serializes the batch context to a temporary JSON file. The schema is documented in agents/file-analyzer.md and includes projectRoot, batchFiles, batchImportData, and neighborMap:

cat > $PROJECT_ROOT/.understand-anything/tmp/ua-file-analyzer-input-<batchIndex>.json << 'ENDJSON'
{
  "projectRoot": "<project-root>",
  "batchFiles": [
    {"path":"src/index.ts","language":"typescript","sizeLines":150,"fileCategory":"code"}
  ],
  "batchImportData": <batchImportData JSON>,
  "neighborMap": <neighborMap JSON>
}
ENDJSON

Executing the Bundled Extraction Script

In Phase 1, the agent invokes the bundled tree-sitter script by passing the prepared input path and a results output path:

node <SKILL_DIR>/extract-structure.mjs \
  $PROJECT_ROOT/.understand-anything/tmp/ua-file-analyzer-input-<batchIndex>.json \
  $PROJECT_ROOT/.understand-anything/tmp/ua-file-extract-results-<batchIndex>.json

The script reads the batched file list from the first argument and writes raw AST data to the second argument.

Sample Batch Output

After Phase 2 enrichment, the agent emits a JSON document containing nodes and edges for that batch. A minimal fragment looks like this:

{
  "nodes": [
    {"type":"file","id":"src/index.ts"},
    {"type":"function","id":"src/index.ts:main","name":"main"}
  ],
  "edges": [
    {"source":"src/index.ts:main","target":"src/utils.ts:parse","type":"calls"}
  ]
}

Production outputs include additional LLM-generated metadata such as summaries, tags, and complexity scores.

Key Implementation Files

The behavior of the file-analyzer agent is defined across five primary files in the Lum1104/Understand-Anything source tree:

Summary

  • The file-analyzer agent is the second-stage worker in the Understand-Anything /understand pipeline.
  • It processes batches of files in parallel, with a maximum concurrency of five sub-agents.
  • Phase 1 runs the extract-structure.mjs tree-sitter script to extract deterministic AST data.
  • Phase 2 uses an LLM to add semantic summaries, tags, complexity scores, and cross-batch edges via the neighborMap.
  • Output follows a strict batch-<index>.json naming convention so that scripts/merge-batch-graphs.py can reliably assemble the final graph.

Frequently Asked Questions

What is the role of the file-analyzer agent in Understand-Anything?

The file-analyzer agent transforms raw source code into structured knowledge-graph elements. It acts as the bridge between the initial project scan and the final merged graph by extracting both syntactic structure and semantic meaning from batched files.

How does the file-analyzer agent handle relationships that cross batch boundaries?

The agent receives a neighborMap in its input payload that contains symbols from files in other batches. It uses this map to emit edges such as calls, inherits, implements, and related, ensuring the merged graph remains internally consistent.

What is the difference between Phase 1 and Phase 2 of the file-analyzer agent?

Phase 1 executes the deterministic, tree-sitter-based extract-structure.mjs script to capture raw AST facts. Phase 2 feeds those facts to an LLM, which generates human-readable summaries, complexity scores, and semantic edges that express architectural intent rather than pure syntax.

Which source files define the file-analyzer agent and its execution steps?

The agent definition lives in agents/file-analyzer.md, orchestration logic resides in skills/understand/SKILL.md, prompt templates are stored in skills/understand/file-analyzer-prompt.md, structural extraction is handled by skills/understand/extract-structure.mjs, and final graph merging is performed by scripts/merge-batch-graphs.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →