Batch Processing Strategy in Understand-Anything: Concurrency, Sizing, and Parallel Execution
The Understand-Anything repository processes codebases through deterministic, size-capped batches that enable parallel agent execution while preserving semantic relationships through Louvain community detection and strict file-per-batch limits.
This repository analyzes large-scale projects by dividing them into semantically coherent batches that can be processed concurrently without filesystem collisions. The batch processing strategy implemented in compute-batches.mjs combines graph-based community detection with hard size constraints to optimize LLM token usage and minimize agent startup overhead.
The Three-Stage Batch Processing Pipeline
The system transforms a raw project scan into parallelizable units through three tightly coupled stages defined in understand-anything-plugin/skills/understand/compute-batches.mjs.
Stage 1: Community Detection via Louvain Algorithm
The pipeline first groups code files by import topology using the Louvain method for community detection. The runLouvain function (lines 200-229) analyzes export/import relationships extracted via Tree-Sitter to identify tightly coupled modules. If the algorithm fails or the environment variable UA_COMPUTE_BATCHES_FORCE_LOUVAIN_THROW is set, the system falls back to countBasedAssignment with a deterministic batch size of 12 files.
Stage 2: Semantic Non-Code Grouping
Non-code files bypass the graph analysis and are routed through buildNonCodeBatches (lines 98-180), which forces specific file types into dedicated, non-mergeable batches:
- Group A: Dockerfile clusters (including
.dockerignoreanddocker-compose.*files in the same directory) - Group B: GitHub Actions workflows (
.github/workflows/*) - Group C: GitLab CI and CircleCI configs (
.gitlab-ci.yml,.circleci/*) - Group D: SQL migrations under
migrations/directories - Group E: Remaining files grouped by immediate parent directory (max 20 per sub-batch)
Stage 3: Size Enforcement and Small-Batch Merging
The mergeSmallBatches function (lines 45-78) enforces the final constraints:
- Louvain communities are capped at 35 files to keep LLM prompts within token budgets
- Batches smaller than 3 files marked as
mergeable: trueare pooled into "misc" batches of up to 25 files - Non-code groups (A-D) remain unmerged to preserve deployment and CI/CD semantics
Files-Per-Batch Limits and Configuration
The batch processing strategy relies on specific numeric thresholds to balance semantic cohesion against computational efficiency:
| Parameter | Value | Location | Purpose |
|---|---|---|---|
| Max Louvain size | 35 files | compute-batches.mjs |
Prevents token limit exhaustion while preserving module relationships |
| Min batch size | 3 files | compute-batches.mjs |
Threshold for pooling tiny batches to reduce agent startup overhead |
| Max merge target | 25 files | compute-batches.mjs |
Upper bound for pooled "misc" batches |
| Neighbor map cap | 50 entries | compute-batches.mjs |
Limits cross-batch reference payload size |
| Fallback size | 12 files | countBasedAssignment |
Deterministic assignment when Louvain fails |
Groups A through D are explicitly marked mergeable: false because splitting Docker configurations or CI workflows would destroy their logical deployment semantics.
Concurrency and Parallel Execution Model
The architecture enables safe parallel processing through filesystem isolation by batch index. Each batch receives a deterministic integer identifier (batchIndex) assigned via sorted community size then alphabetically, ensuring consistent cross-references even when processing out of order.
Isolation Mechanism
Every file-analyzer agent writes to unique temporary files named with the batch index:
.understand-anything/tmp/ua-file-analyzer-input-<batchIndex>.json
.understand-anything/tmp/ua-file-extract-results-<batchIndex>.json
Because filenames embed the original batchIndex, multiple agents can run concurrently on separate CPU cores or LLM instances without race conditions or filesystem collisions.
Dispatcher Implementation
The dispatcher launches agents in parallel using the batch index for isolation:
import { readFileSync } from 'node:fs';
import { spawn } from 'node:child_process';
const batches = JSON.parse(
readFileSync('./.understand-anything/intermediate/batches.json', 'utf8')
).batches;
// Fire off a file-analyzer for each batch – they run in parallel
for (const b of batches) {
const args = [
'--batch-index', String(b.batchIndex),
'--project-root', projectRoot,
];
spawn('node', ['run-file-analyzer.mjs', ...args], {
stdio: 'inherit',
detached: true, // OS schedules them independently
});
}
Handling Token Limits and Output Chunking
Even after batch sizing, individual files may generate extraction graphs exceeding LLM context windows. The file-analyzer agent implements secondary chunking based on structural complexity:
If nodeCount > 60 or edgeCount > 120, the batch splits into parts (batch-<idx>-part-<k>.json) while preserving the original batchIndex for downstream merging (file-analyzer.md lines 480-508).
The generated batches.json structure includes cross-batch references necessary for reconstruction:
{
"batchIndex": 4,
"files": [
{"path":"src/auth/login.ts","language":"typescript","sizeLines":42}
],
"batchImportData": {
"src/auth/login.ts": ["src/auth/session.ts"]
},
"neighborMap": {
"src/auth/login.ts": [
{"path":"src/auth/session.ts","batchIndex":4,"symbols":["User"]},
{"path":"src/auth/tokens.ts","batchIndex":5,"symbols":["generateToken"]}
]
}
}
Summary
- Louvain community detection groups related code files by import topology, with a hard cap of 35 files per batch to respect LLM token limits.
- Semantic non-code groups (A-E) isolate Docker, CI/CD, and migration files into logical units that cannot be merged with unrelated content.
- Small-batch pooling aggregates batches under 3 files into "misc" groups of up to 25 files, reducing agent startup overhead.
- Concurrency isolation is achieved through deterministic
batchIndexvalues embedded in temporary filenames, enabling parallel execution without filesystem collisions. - Secondary chunking splits individual batches by node/edge count (60/120 thresholds) if structural extraction exceeds token budgets.
Frequently Asked Questions
How does Understand-Anything prevent race conditions when running multiple batch analyzers in parallel?
The system eliminates race conditions through index-based filesystem isolation. Each batch writes to uniquely named temporary files (ua-file-analyzer-input-<batchIndex>.json) where the batchIndex is assigned deterministically during the compute-batches.mjs phase. Because no two batches share the same index, parallel agents operate on completely separate file handles, allowing the OS to schedule them across multiple CPU cores or distributed LLM instances safely.
What happens if the Louvain algorithm fails to detect communities in a project?
If the Louvain method throws an exception or if the UA_COMPUTE_BATCHES_FORCE_LOUVAIN_THROW environment variable is set, the pipeline falls back to countBasedAssignment. This deterministic function assigns files into fixed-size batches of 12 files each, ensuring the analysis can proceed even when the import graph is too sparse or corrupted for community detection.
Why are some batches limited to 35 files while others can reach 25 after merging?
The 35-file limit applies specifically to Louvain-detected code communities to preserve semantic cohesion while respecting LLM token budgets. The 25-file limit applies only to "misc" batches created by mergeSmallBatches, which pool tiny, mergeable batches (under 3 files) that would otherwise waste agent startup resources. Non-code groups (A-D) and large Louvain communities follow different constraints because their boundaries are semantically significant rather than arbitrary clusterings.
Can I run the batch computation independently of the analysis agents?
Yes. Execute the batch builder directly via Node.js to generate batches.json without launching analyzers:
node understand-anything-plugin/skills/understand/compute-batches.mjs /path/to/project
This writes ./.understand-anything/intermediate/batches.json containing all file assignments, neighbor maps, and import data. You can then inspect the batch distribution or selectively rerun specific indices using --batch-index without recomputing the entire partition.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →