# Batch Processing Strategy in Understand-Anything: Concurrency, Sizing, and Parallel Execution

> Explore the Understand-Anything batch processing strategy. Learn how concurrency, file limits, and parallel execution preserve semantic relationships for efficient code analysis.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-02

---

**The Understand-Anything repository processes codebases through deterministic, size-capped batches that enable parallel agent execution while preserving semantic relationships through Louvain community detection and strict file-per-batch limits.**

This repository analyzes large-scale projects by dividing them into semantically coherent batches that can be processed concurrently without filesystem collisions. The batch processing strategy implemented in `compute-batches.mjs` combines graph-based community detection with hard size constraints to optimize LLM token usage and minimize agent startup overhead.

## The Three-Stage Batch Processing Pipeline

The system transforms a raw project scan into parallelizable units through three tightly coupled stages defined in `understand-anything-plugin/skills/understand/compute-batches.mjs`.

### Stage 1: Community Detection via Louvain Algorithm

The pipeline first groups code files by import topology using the **Louvain method** for community detection. The `runLouvain` function ([lines 200-229](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L200-L229)) analyzes export/import relationships extracted via Tree-Sitter to identify tightly coupled modules. If the algorithm fails or the environment variable `UA_COMPUTE_BATCHES_FORCE_LOUVAIN_THROW` is set, the system falls back to `countBasedAssignment` with a deterministic batch size of **12 files**.

### Stage 2: Semantic Non-Code Grouping

Non-code files bypass the graph analysis and are routed through `buildNonCodeBatches` ([lines 98-180](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L98-L180)), which forces specific file types into dedicated, non-mergeable batches:

- **Group A**: Dockerfile clusters (including `.dockerignore` and `docker-compose.*` files in the same directory)
- **Group B**: GitHub Actions workflows (`.github/workflows/*`)
- **Group C**: GitLab CI and CircleCI configs ([`.gitlab-ci.yml`](https://github.com/Lum1104/Understand-Anything/blob/main/.gitlab-ci.yml), `.circleci/*`)
- **Group D**: SQL migrations under `migrations/` directories
- **Group E**: Remaining files grouped by immediate parent directory (max 20 per sub-batch)

### Stage 3: Size Enforcement and Small-Batch Merging

The `mergeSmallBatches` function ([lines 45-78](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/compute-batches.mjs#L45-L78)) enforces the final constraints:

- **Louvain communities are capped at 35 files** to keep LLM prompts within token budgets
- **Batches smaller than 3 files** marked as `mergeable: true` are pooled into "misc" batches of up to **25 files**
- **Non-code groups (A-D) remain unmerged** to preserve deployment and CI/CD semantics

## Files-Per-Batch Limits and Configuration

The batch processing strategy relies on specific numeric thresholds to balance semantic cohesion against computational efficiency:

| Parameter | Value | Location | Purpose |
|-----------|-------|----------|---------|
| **Max Louvain size** | 35 files | `compute-batches.mjs` | Prevents token limit exhaustion while preserving module relationships |
| **Min batch size** | 3 files | `compute-batches.mjs` | Threshold for pooling tiny batches to reduce agent startup overhead |
| **Max merge target** | 25 files | `compute-batches.mjs` | Upper bound for pooled "misc" batches |
| **Neighbor map cap** | 50 entries | `compute-batches.mjs` | Limits cross-batch reference payload size |
| **Fallback size** | 12 files | `countBasedAssignment` | Deterministic assignment when Louvain fails |

Groups A through D are explicitly marked `mergeable: false` because splitting Docker configurations or CI workflows would destroy their logical deployment semantics.

## Concurrency and Parallel Execution Model

The architecture enables safe parallel processing through **filesystem isolation by batch index**. Each batch receives a deterministic integer identifier (`batchIndex`) assigned via sorted community size then alphabetically, ensuring consistent cross-references even when processing out of order.

### Isolation Mechanism

Every file-analyzer agent writes to unique temporary files named with the batch index:

```bash
.understand-anything/tmp/ua-file-analyzer-input-<batchIndex>.json
.understand-anything/tmp/ua-file-extract-results-<batchIndex>.json

```

Because filenames embed the original `batchIndex`, multiple agents can run concurrently on separate CPU cores or LLM instances without race conditions or filesystem collisions.

### Dispatcher Implementation

The dispatcher launches agents in parallel using the batch index for isolation:

```typescript
import { readFileSync } from 'node:fs';
import { spawn } from 'node:child_process';

const batches = JSON.parse(
  readFileSync('./.understand-anything/intermediate/batches.json', 'utf8')
).batches;

// Fire off a file-analyzer for each batch – they run in parallel
for (const b of batches) {
  const args = [
    '--batch-index', String(b.batchIndex),
    '--project-root', projectRoot,
  ];
  spawn('node', ['run-file-analyzer.mjs', ...args], {
    stdio: 'inherit',
    detached: true,  // OS schedules them independently
  });
}

```

## Handling Token Limits and Output Chunking

Even after batch sizing, individual files may generate extraction graphs exceeding LLM context windows. The file-analyzer agent implements secondary chunking based on structural complexity:

If `nodeCount > 60` or `edgeCount > 120`, the batch splits into parts (`batch-<idx>-part-<k>.json`) while preserving the original `batchIndex` for downstream merging ([file-analyzer.md lines 480-508](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md#L480-L508)).

The generated [`batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batches.json) structure includes cross-batch references necessary for reconstruction:

```json
{
  "batchIndex": 4,
  "files": [
    {"path":"src/auth/login.ts","language":"typescript","sizeLines":42}
  ],
  "batchImportData": {
    "src/auth/login.ts": ["src/auth/session.ts"]
  },
  "neighborMap": {
    "src/auth/login.ts": [
      {"path":"src/auth/session.ts","batchIndex":4,"symbols":["User"]},
      {"path":"src/auth/tokens.ts","batchIndex":5,"symbols":["generateToken"]}
    ]
  }
}

```

## Summary

- **Louvain community detection** groups related code files by import topology, with a hard cap of **35 files per batch** to respect LLM token limits.
- **Semantic non-code groups** (A-E) isolate Docker, CI/CD, and migration files into logical units that cannot be merged with unrelated content.
- **Small-batch pooling** aggregates batches under **3 files** into "misc" groups of up to **25 files**, reducing agent startup overhead.
- **Concurrency isolation** is achieved through deterministic `batchIndex` values embedded in temporary filenames, enabling parallel execution without filesystem collisions.
- **Secondary chunking** splits individual batches by node/edge count (60/120 thresholds) if structural extraction exceeds token budgets.

## Frequently Asked Questions

### How does Understand-Anything prevent race conditions when running multiple batch analyzers in parallel?

The system eliminates race conditions through **index-based filesystem isolation**. Each batch writes to uniquely named temporary files (`ua-file-analyzer-input-<batchIndex>.json`) where the `batchIndex` is assigned deterministically during the `compute-batches.mjs` phase. Because no two batches share the same index, parallel agents operate on completely separate file handles, allowing the OS to schedule them across multiple CPU cores or distributed LLM instances safely.

### What happens if the Louvain algorithm fails to detect communities in a project?

If the Louvain method throws an exception or if the `UA_COMPUTE_BATCHES_FORCE_LOUVAIN_THROW` environment variable is set, the pipeline falls back to `countBasedAssignment`. This deterministic function assigns files into fixed-size batches of **12 files** each, ensuring the analysis can proceed even when the import graph is too sparse or corrupted for community detection.

### Why are some batches limited to 35 files while others can reach 25 after merging?

The **35-file limit** applies specifically to Louvain-detected code communities to preserve semantic cohesion while respecting LLM token budgets. The **25-file limit** applies only to "misc" batches created by `mergeSmallBatches`, which pool tiny, mergeable batches (under 3 files) that would otherwise waste agent startup resources. Non-code groups (A-D) and large Louvain communities follow different constraints because their boundaries are semantically significant rather than arbitrary clusterings.

### Can I run the batch computation independently of the analysis agents?

Yes. Execute the batch builder directly via Node.js to generate [`batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/batches.json) without launching analyzers:

```bash
node understand-anything-plugin/skills/understand/compute-batches.mjs /path/to/project

```

This writes [`./.understand-anything/intermediate/batches.json`](https://github.com/Lum1104/Understand-Anything/blob/main/./.understand-anything/intermediate/batches.json) containing all file assignments, neighbor maps, and import data. You can then inspect the batch distribution or selectively rerun specific indices using `--batch-index` without recomputing the entire partition.