# How Parallel File Analyses Are Managed in the Understand‑Anything Pipeline

> Discover how the Understand-Anything pipeline manages parallel file analyses. Learn how indexed batches are processed concurrently and merged into a unified knowledge graph.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: internals
- Published: 2026-06-01

---

**The Understand‑Anything pipeline achieves parallel file analyses by splitting source files into indexed batches that are processed concurrently by isolated agents, then deterministically merging the per‑batch JSON outputs into a unified knowledge graph while validating completeness and recovering any missing import edges.**

The **Understand‑Anything** repository implements a sophisticated parallel processing architecture to analyze large codebases efficiently. Rather than processing files sequentially, the pipeline distributes workloads across multiple file‑analyzer agents that operate simultaneously on distinct batches, ensuring scalability without race conditions. This article examines the exact mechanisms—from batch indexing to deterministic merging—that enable safe parallel file analyses in the pipeline.

## Batch Isolation Strategies for Parallel File Analyses

The orchestrator initiates parallel processing by splitting the full file list into logical batches, assigning each a numeric `batchIndex`. This index permeates the temporary file naming convention to prevent disk‑level collisions between concurrent agents.

According to the agent definition in [`understand-anything-plugin/agents/file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md) (lines 31‑38), every temporary file must embed the `batchIndex`:

- Input manifest: `ua-file-analyzer-input-<batchIndex>.json`
- Extraction results: `ua-file-extract-results-<batchIndex>.json`

This naming guarantees that concurrently running agents never clash on disk, even when hundreds of batches execute simultaneously.

## Deterministic Structural Extraction per Batch

Each agent executes the bundled `extract-structure.mjs` script on its assigned batch. The script performs a one‑time initialization of all Tree‑Sitter grammars via `await tsPlugin.init()`, then iterates over the batch’s files calling `registry.analyzeFile` (and `registry.extractCallGraph` for code or script files).

As implemented in `understand-anything-plugin/skills/understand/extract-structure.mjs` (lines 78‑101), this approach is pure JavaScript/Node and stateless between files, allowing multiple instances to run in parallel processes without interference. The grammars are loaded once per batch (via `Promise.all(loadPromises)` in [`understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts)), avoiding repeated I/O while ensuring each batch has its own isolated context.

## Strict Output Naming and Validation

After processing, each agent writes **exactly one** JSON file per batch, or multiple parts when node/edge counts exceed thresholds. The orchestrator enforces strict filename patterns; any deviation causes the file to be silently dropped during assembly.

As defined in [`understand-anything-plugin/agents/file-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/file-analyzer.md) (lines 82‑88), valid outputs must match:
- `batch-<batchIndex>.json`
- `batch-<batchIndex>-part-<k>.json`

This convention allows the downstream merge step to unambiguously identify and sequence batch results.

## The Merge Phase: Deterministic Assembly

Once all agents finish, the [`merge-batch-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-batch-graphs.py) script discovers every `batch-*.json` file, sorts them by numeric index, and merges nodes and edges into a single knowledge graph. The script validates that multipart files are contiguous and that no unrecognized filenames were produced.

The logic in [`understand-anything-plugin/skills/understand/merge-batch-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/merge-batch-graphs.py) (lines 15‑30 and 65‑78) generates explicit warnings for missing parts or malformed filenames, enabling early detection of parallel‑execution glitches in automated pipelines.

## Import Edge Recovery for Graph Completeness

To guarantee graph integrity regardless of batch processing errors, the merge script rescues any `imports` edges that individual agents might have omitted. It consults the deterministic `importMap` produced by the project‑scanner to recover missing relationships.

The `recover_imports_from_scan` function in [`merge-batch-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-batch-graphs.py) (lines 40‑48) runs after the initial merge, ensuring the final graph contains a complete set of internal import relationships even when specific batches fail to enumerate them correctly.

## Practical Implementation Examples

The following example demonstrates the end‑to‑end flow for a single batch, which scales horizontally when executed in parallel across many agents:

```bash

# 1️⃣ Write the batch input (the orchestrator does this)

cat > $PROJECT_ROOT/.understand-anything/tmp/ua-file-analyzer-input-3.json <<'ENDJSON'
{
  "projectRoot": "/path/to/project",
  "batchFiles": [
    {"path":"src/index.ts","language":"typescript","sizeLines":120,"fileCategory":"code"},
    {"path":"README.md","language":"markdown","sizeLines":30,"fileCategory":"docs"}
  ],
  "batchImportData": {
    "src/index.ts": ["src/util.ts","src/config.ts"]
  }
}
ENDJSON

# 2️⃣ Execute the deterministic extractor for this batch

node extract-structure.mjs \
  $PROJECT_ROOT/.understand-anything/tmp/ua-file-analyzer-input-3.json \
  $PROJECT_ROOT/.understand-anything/tmp/ua-file-extract-results-3.json

# 3️⃣ The agent converts the extractor output into the required batch-N JSON

# (handled automatically by the agent logic – see file-analyzer.md Step 2)

# 4️⃣ After all batches finish, merge them

python merge-batch-graphs.py $PROJECT_ROOT

```

Programmatically, the merge script groups files by numeric index to ensure deterministic ordering:

```python
import json, re
from pathlib import Path

intermediate = Path("/my/project/.understand-anything/intermediate")
batch_files = sorted(
    intermediate.glob("batch-*.json"),
    key=lambda p: int(re.search(r"batch-(\d+)", p.stem).group(1))
)

# Load each batch, ignoring any file that doesn't match the strict pattern

batches = [json.loads(p.read_text()) for p in batch_files
           if re.match(r"batch-(\d+)(?:-part-(\d+))?\.json", p.name)]

assembled, report = merge_and_normalize(batches)   # from merge-batch-graphs.py

print("\n".join(report))                         # Shows any missing parts / dropped files

```

## Summary

- **Indexed batch files** prevent race conditions by embedding the `batchIndex` into every temporary file path.
- **One‑time grammar initialization** in `extract-structure.mjs` minimizes I/O overhead while keeping batches isolated.
- **Strict naming conventions** (`batch-<N>.json` or `batch-<N>-part-<K>.json`) ensure the merge script can deterministically assemble results.
- **Post‑merge validation** surfaces missing parts or malformed outputs as explicit warnings.
- **Import edge recovery** guarantees graph completeness by consulting the central `importMap` after merging.

## Frequently Asked Questions

### What prevents race conditions when running parallel file analyses?

Race conditions are avoided through **disk isolation via batch indexing**. Each parallel agent writes to temporary files that embed its unique `batchIndex` (e.g., [`ua-file-analyzer-input-3.json`](https://github.com/Lum1104/Understand-Anything/blob/main/ua-file-analyzer-input-3.json)), ensuring no two agents attempt to write to the same path during concurrent execution.

### How does the pipeline handle partial or failed batch outputs?

The [`merge-batch-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-batch-graphs.py) script validates all discovered files against the strict `batch-<N>.json` or `batch-<N>-part-<K>.json` pattern. Missing parts or unrecognized filenames generate explicit warnings in the final report, allowing the orchestrator to identify which batches failed without corrupting the assembled graph.

### Why does each batch load Tree‑Sitter grammars only once?

Each `extract-structure.mjs` instance calls `await tsPlugin.init()` once to load all grammars via `Promise.all(loadPromises)` in [`tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/tree-sitter-plugin.ts) before iterating over its assigned files. This minimizes I/O overhead while maintaining process isolation, as each batch runs in its own independent Node.js process.

### What happens if import edges are missing from a batch?

The merge phase executes `recover_imports_from_scan` to rescue omitted `imports` edges by consulting the deterministic `importMap` produced by the project scanner. This ensures the final knowledge graph contains complete internal import relationships regardless of individual batch processing errors.