How Semantic Batches Are Computed for Optimal File-Analyzer Parallelization in Understand-Anything

Semantic batches are computed by building an import graph of code files, applying Louvain community detection to identify tightly coupled modules, and enforcing a maximum size of 35 files per batch while merging small groups to minimize sub-agent overhead.

In the Understand-Anything repository, semantic batching occurs as Phase 1.5 of the analysis pipeline. The compute-batches.mjs script processes the scan-result.json output to partition files into semantically coherent batches that respect module boundaries while enabling parallel processing by LLM-powered file-analyzer agents.

The Semantic Batching Pipeline

The batching strategy centers on import relationships to ensure that files with tight logical coupling are analyzed together, reducing cross-batch edge loss.

Building the Import Graph

Only files categorized as code (fileCategory === "code") participate in the graph construction. The script creates an undirected graph using graphology where each node represents a file path and edges represent import relationships. Both "imports" and "imported-by" relationships are added to create bidirectional connections. This construction happens within the runLouvain function at lines 94‑100.

Louvain Community Detection

Once the graph is built, the script runs Louvain community detection via graphology-communities-louvain at lines 32‑34. This produces a mapping of file → communityId that groups files into semantic modules based on dense import connectivity. Files that import each other frequently are clustered into the same batch, ensuring context locality for the LLM.

Enforcing Maximum Community Size

To prevent LLM context window overflow, large communities are split so that no batch exceeds 35 files. This limit is controlled by the MAX_COMMUNITY_SIZE constant defined at lines 69‑71. When communities exceed this threshold, the script applies simple alphabetical chunking at lines 83‑89, emitting a warning to stderr at lines 78‑81.

Consolidating Small and Non-Code Batches

Very small batches (fewer than 3 files) are inefficient for parallel dispatch. The mergeSmallBatches function at lines 45‑53 pools these into "misc" batches to reduce sub-agent overhead. Non-code groups such as Dockerfiles, CI configs, and SQL migrations are handled separately by buildNonCodeBatches (lines 90‑100) and are explicitly marked with mergeable: false to preserve their semantic boundaries.

Batch Output Structure and Metadata

The final output is written to .understand-anything/intermediate/batches.json at lines 225‑232. Each batch object contains:

  • files: Full FileMeta objects including path, language, size, and category.
  • batchImportData: The slice of the original importMap belonging to the batch, allowing the file-analyzer to resolve imports without re-parsing.
  • neighborMap: One-hop cross-batch neighbors with exported symbols, trimmed to the 50 most popular neighbors per file (lines 78‑89).
  • batchIndex: A stable 1-based numeric identifier.

Pre-Resolved Import Data

The batchImportData field provides the file-analyzer with pre-resolved import relationships. This allows the agent to emit complete imports edges immediately without performing its own resolution, as specified in file-analyzer.md at lines 84‑92.

Cross-Batch Context for Semantic Edges

The neighborMap supplies cross-batch context containing symbols exported by neighboring batches. This enables the LLM to safely emit high-confidence semantic edges such as calls, inherits, and implements that reference code defined in other batches, documented in the "Cross-batch context" section at lines 30‑38.

Fallback Chunking Strategy

If the Louvain algorithm fails (for example, due to missing native dependencies), the script falls back to deterministic alphabetical chunking. The countBasedAssignment function at lines 22‑28 creates batches of 12 files each, and the output records algorithm: "count-fallback" at lines 47‑59. This ensures the pipeline remains functional even when advanced graph analysis is unavailable.

Running the Batch Computation

Execute the Phase 1.5 batching step from the repository root:

node understand-anything-plugin/skills/understand/compute-batches.mjs $PROJECT_ROOT

This generates .understand-anything/intermediate/batches.json containing the optimized batch definitions. Diagnostics are written to stderr, while the JSON output is written to the intermediate directory.

Inspecting Batch Content

A typical batch entry in batches.json appears as:

{
  "batchIndex": 3,
  "files": [
    { "path": "src/auth/login.ts", "language": "typescript", "sizeLines": 120, "fileCategory": "code" },
    { "path": "src/auth/session.ts", "language": "typescript", "sizeLines": 95, "fileCategory": "code" }
  ],
  "batchImportData": {
    "src/auth/login.ts": ["src/auth/session.ts", "src/db/users.ts"],
    "src/auth/session.ts": ["src/db/users.ts"]
  },
  "neighborMap": {
    "src/auth/login.ts": [
      { "path": "src/db/users.ts", "batchIndex": 7, "symbols": ["User","findById","createUser"] }
    ]
  }
}

Consuming Batches in File-Analyzer

The file-analyzer agent receives a single batch payload and produces batch-<index>.json (or batch-<index>-part-<k>.json if output exceeds LLM limits). After all batches complete, the merge script assembles the final graph:

python understand-anything-plugin/skills/understand/merge-batch-graphs.py $PROJECT_ROOT

This produces .understand-anything/graph/assembled-graph.json by stitching all batch parts and running a deterministic safety-net to recover any missing import edges.

Summary

  • Semantic batching in Understand-Anything uses Louvain community detection on an import graph to group tightly coupled code files.
  • The system enforces a hard limit of 35 files per batch via alphabetical splitting and consolidates batches smaller than 3 files to optimize parallel dispatch.
  • Each batch exports batchImportData for internal imports and a neighborMap (limited to 50 symbols) for safe cross-batch semantic edge generation.
  • A fallback mode using 12-file alphabetical chunks ensures pipeline reliability if graph analysis dependencies are missing.
  • The architecture minimizes cross-batch edge loss while keeping individual LLM tasks within token limits.

Frequently Asked Questions

Why does Understand-Anything use Louvain community detection for semantic batching?

Louvain community detection identifies densely connected subgraphs in the import graph, which correspond to logical software modules. Grouping these tightly coupled files together ensures that the LLM has full context for internal relationships within a batch, reducing the need for uncertain cross-batch inferences and improving the accuracy of generated semantic edges.

What happens when a semantic community exceeds 35 files?

When the Louvain algorithm produces a community larger than MAX_COMMUNITY_SIZE (35 files), the compute-batches.mjs script splits it using alphabetical chunking at lines 83‑89. A warning is emitted to stderr indicating that semantic boundaries were breached due to size constraints, ensuring operators are aware of the fragmentation.

How does the file-analyzer handle dependencies on code in other batches?

The file-analyzer receives a neighborMap containing the top 50 most frequently referenced symbols from adjacent batches. This map provides sufficient context for the LLM to emit semantic edges like calls, implements, or inherits targeting code outside the current batch without requiring the full source of those external files, as detailed in file-analyzer.md at lines 30‑38.

What is the fallback if the Louvain algorithm fails to run?

If the graphology-communities-louvain library throws an error (typically due to missing native dependencies), the script automatically falls back to countBasedAssignment at lines 22‑28. This creates deterministic batches of 12 files each in alphabetical order, recording algorithm: "count-fallback" in the output metadata to indicate that semantic grouping was not applied.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →