# How Understand-Anything Batches Files for Parallel Analysis: A Two-Phase Deep Dive

> Learn how Understand Anything batches files for parallel analysis. Discover its two-phase process for efficient concurrent processing and knowledge graph generation.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-08

---

**Understand-Anything splits discovered source files into fixed-size chunks and processes each chunk concurrently in separate worker threads, then merges the structural results into a single knowledge graph.**

Understand-Anything is an open-source knowledge-graph builder for codebases. To keep analysis fast and memory-safe, it uses a deterministic pipeline that batches files for parallel analysis before aggregating their structural data into one coherent graph.

## The Two-Phase Analysis Pipeline

The repository implements a strict separation between discovery and execution.

1. **Discovery** – The engine walks the project directory and emits a flat list of every supported source file.
2. **Batch execution** – That list is sliced into chunks, each assigned to a worker thread that runs the appropriate language extractor.

This design keeps I/O concerns separate from CPU-bound parsing and ensures the system can scale from small libraries to large monorepos.

## File Discovery and List Generation

The first phase lives in [`packages/core/src/plugins/discovery.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/plugins/discovery.ts).

The **FileDiscovery** logic recursively scans the filesystem and collects all matching source paths into a single array. No parsing happens here; the goal is simply to build a complete, ordered file list that the next stage can consume.

## How Files Are Batched for Parallel Analysis

The core batching algorithm is defined in [`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts). Inside the `GraphBuilder` class, a static helper splits the flat file array into equally sized chunks:

```ts
// packages/core/src/analyzer/graph-builder.ts (excerpt)
private readonly batchSize = 50;                // default batch size

/** Split the flat file list into equally‑sized chunks */
private static batchFiles(files: string[], size: number): string[][] {
  const batches: string[][] = [];
  for (let i = 0; i < files.length; i += size) {
    batches.push(files.slice(i, i + size));
  }
  return batches;
}

/** Orchestrate parallel extraction */
async analyzeFiles(filePaths: string[]) {
  const batches = GraphBuilder.batchFiles(filePaths, this.batchSize);
  // Run each batch concurrently (max = #CPU cores)
  await Promise.all(
    batches.map((batch) => this.runExtractionBatch(batch)),
  );
}

```

The **`batchFiles`** method uses a simple loop with `Array.prototype.slice` to create non-overlapping segments. The **`analyzeFiles`** entry point then invokes `runExtractionBatch` for every segment inside a `Promise.all` call.

### Why the Default Batch Size Is 50

The hard-coded **`batchSize = 50`** balances messaging overhead against memory pressure.

- **Too small** – Excessive batches spawn too many workers and pay repeatedly for inter-process messaging.
- **Too large** – A single worker can hog memory and block the thread pool.

You can override this default through the public `GraphBuilder` constructor on lines 60-66 of the same file.

## Worker Thread Execution and Concurrency Limits

Each batch is handled by **`runExtractionBatch`**, which spins up a **worker thread** via Node’s `worker_threads` API.

Because the batches are launched under `Promise.all`, they all start simultaneously up to the physical limit of `os.cpus().length`. The pool is created lazily: if the number of batches exceeds the core count, excess batches queue until a thread frees. This prevents context-switching thrashing while still keeping every core busy.

Individual workers load the relevant language extractor—such as [`typescript-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/typescript-extractor.ts) or [`python-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/python-extractor.ts) from `packages/core/src/plugins/extractors/*.ts`—and return an array of `StructuralAnalysis` objects.

## Aggregating Results into the Knowledge Graph

Result merging is intentionally **single-threaded** to avoid race conditions.

As each worker resolves, the main thread feeds its output into `GraphBuilder` methods such as **`addFileWithAnalysis`**, **`addImportEdge`**, and **`addCallEdge`**. This sequential aggregation populates the shared `nodes` and `edges` collections safely, guaranteeing a coherent final graph even when hundreds of files were analyzed in parallel.

## Practical Code Examples

### Example 1: Manually Invoke the Batch Logic

```ts
import { GraphBuilder } from '@understand-anything/core';

// Suppose we already have a list of file paths:
const files = [
  'src/app.ts',
  'src/utils/helper.ts',
  // … many more …
];

// Create a builder (project name & git hash are arbitrary here)
const builder = new GraphBuilder('my-project', 'deadbeef');

// Run the parallel analysis step
await builder.analyzeFiles(files);

// When done, retrieve the graph
const graph = builder.build();

```

### Example 2: Adjust the Batch Size

```ts
class MyBuilder extends GraphBuilder {
  // Reduce batch size for low-memory environments
  private readonly batchSize = 20;
}

```

### Example 3: Hook into the Worker Pool

```ts
// Inside a custom extractor you can access the per-batch context:
export async function extractBatch(batch: string[]) {
  // This runs inside a worker thread; you can safely use heavy
  // CPU-bound parsers without blocking the main process.
}

```

## Testing the Batching Logic

Unit coverage for this behavior lives in **[`packages/core/src/analyzer/graph-builder.test.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.test.ts)**. Look for the test titled *“should process files in parallel batches”*, which validates that the `batchFiles` splitter and the `Promise.all` orchestration behave correctly under varying file counts.

## Summary

- **Discovery first** – [`packages/core/src/plugins/discovery.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/plugins/discovery.ts) builds a flat list of every supported source file.
- **Fixed-size chunks** – `GraphBuilder.batchFiles` in [`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts) splits the list into batches of **50 files** by default.
- **Worker parallelism** – `analyzeFiles` launches each batch via `runExtractionBatch` inside `Promise.all`, constrained by the number of CPU cores.
- **Safe aggregation** – The main thread merges `StructuralAnalysis` results sequentially using `addFileWithAnalysis`, `addImportEdge`, and `addCallEdge` to prevent race conditions.
- **Configurable** – Override the default batch size through the `GraphBuilder` constructor or by extending the class.

## Frequently Asked Questions

### What is the default batch size when Understand-Anything batches files for parallel analysis?

The default batch size is **50 files**, defined as a private constant in [`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts). This value is chosen to minimize worker-spawn overhead while preventing any single thread from consuming excessive memory.

### How does Understand-Anything limit the number of concurrent workers?

Concurrency is naturally capped by the system’s CPU core count via `os.cpus().length`. The worker pool is created lazily, so if there are more batches than cores, the excess tasks queue until a thread finishes its current batch.

### Can I change the batch size for parallel analysis?

Yes. The default can be overridden through the public `GraphBuilder` constructor (lines 60-66 of [`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts)). Alternatively, you can subclass `GraphBuilder` and declare your own `batchSize` value.

### Is the result aggregation step thread-safe?

Yes, because aggregation is intentionally **single-threaded**. Although extraction runs in parallel worker threads, the main thread collects each batch’s `StructuralAnalysis` objects sequentially and feeds them into `GraphBuilder` methods such as `addFileWithAnalysis` and `addImportEdge`, which avoids race conditions on the shared graph structures.