How Understand-Anything Batches Files for Parallel Analysis: A Two-Phase Deep Dive
Understand-Anything splits discovered source files into fixed-size chunks and processes each chunk concurrently in separate worker threads, then merges the structural results into a single knowledge graph.
Understand-Anything is an open-source knowledge-graph builder for codebases. To keep analysis fast and memory-safe, it uses a deterministic pipeline that batches files for parallel analysis before aggregating their structural data into one coherent graph.
The Two-Phase Analysis Pipeline
The repository implements a strict separation between discovery and execution.
- Discovery – The engine walks the project directory and emits a flat list of every supported source file.
- Batch execution – That list is sliced into chunks, each assigned to a worker thread that runs the appropriate language extractor.
This design keeps I/O concerns separate from CPU-bound parsing and ensures the system can scale from small libraries to large monorepos.
File Discovery and List Generation
The first phase lives in packages/core/src/plugins/discovery.ts.
The FileDiscovery logic recursively scans the filesystem and collects all matching source paths into a single array. No parsing happens here; the goal is simply to build a complete, ordered file list that the next stage can consume.
How Files Are Batched for Parallel Analysis
The core batching algorithm is defined in packages/core/src/analyzer/graph-builder.ts. Inside the GraphBuilder class, a static helper splits the flat file array into equally sized chunks:
// packages/core/src/analyzer/graph-builder.ts (excerpt)
private readonly batchSize = 50; // default batch size
/** Split the flat file list into equally‑sized chunks */
private static batchFiles(files: string[], size: number): string[][] {
const batches: string[][] = [];
for (let i = 0; i < files.length; i += size) {
batches.push(files.slice(i, i + size));
}
return batches;
}
/** Orchestrate parallel extraction */
async analyzeFiles(filePaths: string[]) {
const batches = GraphBuilder.batchFiles(filePaths, this.batchSize);
// Run each batch concurrently (max = #CPU cores)
await Promise.all(
batches.map((batch) => this.runExtractionBatch(batch)),
);
}
The batchFiles method uses a simple loop with Array.prototype.slice to create non-overlapping segments. The analyzeFiles entry point then invokes runExtractionBatch for every segment inside a Promise.all call.
Why the Default Batch Size Is 50
The hard-coded batchSize = 50 balances messaging overhead against memory pressure.
- Too small – Excessive batches spawn too many workers and pay repeatedly for inter-process messaging.
- Too large – A single worker can hog memory and block the thread pool.
You can override this default through the public GraphBuilder constructor on lines 60-66 of the same file.
Worker Thread Execution and Concurrency Limits
Each batch is handled by runExtractionBatch, which spins up a worker thread via Node’s worker_threads API.
Because the batches are launched under Promise.all, they all start simultaneously up to the physical limit of os.cpus().length. The pool is created lazily: if the number of batches exceeds the core count, excess batches queue until a thread frees. This prevents context-switching thrashing while still keeping every core busy.
Individual workers load the relevant language extractor—such as typescript-extractor.ts or python-extractor.ts from packages/core/src/plugins/extractors/*.ts—and return an array of StructuralAnalysis objects.
Aggregating Results into the Knowledge Graph
Result merging is intentionally single-threaded to avoid race conditions.
As each worker resolves, the main thread feeds its output into GraphBuilder methods such as addFileWithAnalysis, addImportEdge, and addCallEdge. This sequential aggregation populates the shared nodes and edges collections safely, guaranteeing a coherent final graph even when hundreds of files were analyzed in parallel.
Practical Code Examples
Example 1: Manually Invoke the Batch Logic
import { GraphBuilder } from '@understand-anything/core';
// Suppose we already have a list of file paths:
const files = [
'src/app.ts',
'src/utils/helper.ts',
// … many more …
];
// Create a builder (project name & git hash are arbitrary here)
const builder = new GraphBuilder('my-project', 'deadbeef');
// Run the parallel analysis step
await builder.analyzeFiles(files);
// When done, retrieve the graph
const graph = builder.build();
Example 2: Adjust the Batch Size
class MyBuilder extends GraphBuilder {
// Reduce batch size for low-memory environments
private readonly batchSize = 20;
}
Example 3: Hook into the Worker Pool
// Inside a custom extractor you can access the per-batch context:
export async function extractBatch(batch: string[]) {
// This runs inside a worker thread; you can safely use heavy
// CPU-bound parsers without blocking the main process.
}
Testing the Batching Logic
Unit coverage for this behavior lives in packages/core/src/analyzer/graph-builder.test.ts. Look for the test titled “should process files in parallel batches”, which validates that the batchFiles splitter and the Promise.all orchestration behave correctly under varying file counts.
Summary
- Discovery first –
packages/core/src/plugins/discovery.tsbuilds a flat list of every supported source file. - Fixed-size chunks –
GraphBuilder.batchFilesinpackages/core/src/analyzer/graph-builder.tssplits the list into batches of 50 files by default. - Worker parallelism –
analyzeFileslaunches each batch viarunExtractionBatchinsidePromise.all, constrained by the number of CPU cores. - Safe aggregation – The main thread merges
StructuralAnalysisresults sequentially usingaddFileWithAnalysis,addImportEdge, andaddCallEdgeto prevent race conditions. - Configurable – Override the default batch size through the
GraphBuilderconstructor or by extending the class.
Frequently Asked Questions
What is the default batch size when Understand-Anything batches files for parallel analysis?
The default batch size is 50 files, defined as a private constant in packages/core/src/analyzer/graph-builder.ts. This value is chosen to minimize worker-spawn overhead while preventing any single thread from consuming excessive memory.
How does Understand-Anything limit the number of concurrent workers?
Concurrency is naturally capped by the system’s CPU core count via os.cpus().length. The worker pool is created lazily, so if there are more batches than cores, the excess tasks queue until a thread finishes its current batch.
Can I change the batch size for parallel analysis?
Yes. The default can be overridden through the public GraphBuilder constructor (lines 60-66 of packages/core/src/analyzer/graph-builder.ts). Alternatively, you can subclass GraphBuilder and declare your own batchSize value.
Is the result aggregation step thread-safe?
Yes, because aggregation is intentionally single-threaded. Although extraction runs in parallel worker threads, the main thread collects each batch’s StructuralAnalysis objects sequentially and feeds them into GraphBuilder methods such as addFileWithAnalysis and addImportEdge, which avoids race conditions on the shared graph structures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →