# How Semantic Batches Are Computed for Optimal File-Analyzer Parallelization in Understand-Anything

> Discover how Understand Anything computes semantic batches for optimal file-analyzer parallelization using import graphs and Louvain community detection, ensuring efficient code analysis.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-05-31

---

**Semantic batches are computed by building an import graph of code files, applying Louvain community detection to identify tightly coupled modules, and enforcing a maximum size of 35 files per batch while merging small groups to minimize sub-agent overhead.**

In the **Understand-Anything** repository, semantic batching occurs as Phase 1.5 of the analysis pipeline. The `compute-batches.mjs` script processes the [`scan-result.json`](https://github.com/Lum1104/Understand-Anything/blob/main/scan-result.json) output to partition files into **semantically coherent batches** that respect module boundaries while enabling parallel processing by LLM-powered file-analyzer agents.

## The Semantic Batching Pipeline

The batching strategy centers on import relationships to ensure that files with tight logical coupling are analyzed together, reducing cross-batch edge loss.

### Building the Import Graph

Only files categorized as code (`fileCategory === "code"`) participate in the graph construction. The script creates an undirected graph using **graphology** where each node represents a file path and edges represent import relationships. Both "imports" and "imported-by" relationships are added to create bidirectional connections. This construction happens within the `runLouvain` function at lines [94‑100](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L94-L100).

### Louvain Community Detection

Once the graph is built, the script runs **Louvain community detection** via `graphology-communities-louvain` at lines [32‑34](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L32-L34). This produces a mapping of `file → communityId` that groups files into **semantic modules** based on dense import connectivity. Files that import each other frequently are clustered into the same batch, ensuring context locality for the LLM.

### Enforcing Maximum Community Size

To prevent LLM context window overflow, large communities are split so that no batch exceeds **35 files**. This limit is controlled by the `MAX_COMMUNITY_SIZE` constant defined at lines [69‑71](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L69-L71). When communities exceed this threshold, the script applies simple alphabetical chunking at lines [83‑89](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L83-L89), emitting a warning to stderr at lines [78‑81](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L78-L81).

### Consolidating Small and Non-Code Batches

Very small batches (fewer than 3 files) are inefficient for parallel dispatch. The `mergeSmallBatches` function at lines [45‑53](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L45-L53) pools these into "misc" batches to reduce sub-agent overhead. Non-code groups such as Dockerfiles, CI configs, and SQL migrations are handled separately by `buildNonCodeBatches` (lines [90‑100](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L90-L100)) and are explicitly marked with `mergeable: false` to preserve their semantic boundaries.

## Batch Output Structure and Metadata

The final output is written to [`.understand-anything/intermediate/batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/intermediate/batches.json) at lines [225‑232](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L225-L232). Each batch object contains:

- **`files`**: Full `FileMeta` objects including path, language, size, and category.
- **`batchImportData`**: The slice of the original `importMap` belonging to the batch, allowing the file-analyzer to resolve imports without re-parsing.
- **`neighborMap`**: One-hop cross-batch neighbors with exported symbols, trimmed to the **50 most popular** neighbors per file (lines [78‑89](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L78-L89)).
- **`batchIndex`**: A stable 1-based numeric identifier.

### Pre-Resolved Import Data

The `batchImportData` field provides the file-analyzer with pre-resolved import relationships. This allows the agent to emit complete `imports` edges immediately without performing its own resolution, as specified in [`file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/file-analyzer.md) at lines [84‑92](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md#L84-L92).

### Cross-Batch Context for Semantic Edges

The **neighborMap** supplies cross-batch context containing symbols exported by neighboring batches. This enables the LLM to safely emit high-confidence **semantic edges** such as `calls`, `inherits`, and `implements` that reference code defined in other batches, documented in the "Cross-batch context" section at lines [30‑38](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md#L30-L38).

## Fallback Chunking Strategy

If the Louvain algorithm fails (for example, due to missing native dependencies), the script falls back to deterministic alphabetical chunking. The `countBasedAssignment` function at lines [22‑28](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L22-L28) creates batches of **12 files** each, and the output records `algorithm: "count-fallback"` at lines [47‑59](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L47-L59). This ensures the pipeline remains functional even when advanced graph analysis is unavailable.

## Running the Batch Computation

Execute the Phase 1.5 batching step from the repository root:

```bash
node understand-anything-plugin/skills/understand/compute-batches.mjs $PROJECT_ROOT

```

This generates [`.understand-anything/intermediate/batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/intermediate/batches.json) containing the optimized batch definitions. Diagnostics are written to stderr, while the JSON output is written to the intermediate directory.

### Inspecting Batch Content

A typical batch entry in [`batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batches.json) appears as:

```json
{
  "batchIndex": 3,
  "files": [
    { "path": "src/auth/login.ts", "language": "typescript", "sizeLines": 120, "fileCategory": "code" },
    { "path": "src/auth/session.ts", "language": "typescript", "sizeLines": 95, "fileCategory": "code" }
  ],
  "batchImportData": {
    "src/auth/login.ts": ["src/auth/session.ts", "src/db/users.ts"],
    "src/auth/session.ts": ["src/db/users.ts"]
  },
  "neighborMap": {
    "src/auth/login.ts": [
      { "path": "src/db/users.ts", "batchIndex": 7, "symbols": ["User","findById","createUser"] }
    ]
  }
}

```

### Consuming Batches in File-Analyzer

The file-analyzer agent receives a single batch payload and produces `batch-<index>.json` (or `batch-<index>-part-<k>.json` if output exceeds LLM limits). After all batches complete, the merge script assembles the final graph:

```bash
python understand-anything-plugin/skills/understand/merge-batch-graphs.py $PROJECT_ROOT

```

This produces [`.understand-anything/graph/assembled-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/graph/assembled-graph.json) by stitching all batch parts and running a deterministic safety-net to recover any missing import edges.

## Summary

- **Semantic batching** in Understand-Anything uses **Louvain community detection** on an import graph to group tightly coupled code files.
- The system enforces a **hard limit of 35 files** per batch via alphabetical splitting and consolidates batches smaller than 3 files to optimize parallel dispatch.
- Each batch exports `batchImportData` for internal imports and a `neighborMap` (limited to 50 symbols) for safe cross-batch semantic edge generation.
- A **fallback mode** using 12-file alphabetical chunks ensures pipeline reliability if graph analysis dependencies are missing.
- The architecture minimizes cross-batch edge loss while keeping individual LLM tasks within token limits.

## Frequently Asked Questions

### Why does Understand-Anything use Louvain community detection for semantic batching?

Louvain community detection identifies densely connected subgraphs in the import graph, which correspond to logical software modules. Grouping these tightly coupled files together ensures that the LLM has full context for internal relationships within a batch, reducing the need for uncertain cross-batch inferences and improving the accuracy of generated semantic edges.

### What happens when a semantic community exceeds 35 files?

When the Louvain algorithm produces a community larger than `MAX_COMMUNITY_SIZE` (35 files), the `compute-batches.mjs` script splits it using alphabetical chunking at lines [83‑89](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L83-L89). A warning is emitted to stderr indicating that semantic boundaries were breached due to size constraints, ensuring operators are aware of the fragmentation.

### How does the file-analyzer handle dependencies on code in other batches?

The file-analyzer receives a `neighborMap` containing the top 50 most frequently referenced symbols from adjacent batches. This map provides sufficient context for the LLM to emit semantic edges like `calls`, `implements`, or `inherits` targeting code outside the current batch without requiring the full source of those external files, as detailed in [`file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/file-analyzer.md) at lines [30‑38](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md#L30-L38).

### What is the fallback if the Louvain algorithm fails to run?

If the `graphology-communities-louvain` library throws an error (typically due to missing native dependencies), the script automatically falls back to `countBasedAssignment` at lines [22‑28](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L22-L28). This creates deterministic batches of 12 files each in alphabetical order, recording `algorithm: "count-fallback"` in the output metadata to indicate that semantic grouping was not applied.