How Parallel File Analyses Are Managed in the Understand‑Anything Pipeline
The Understand‑Anything pipeline achieves parallel file analyses by splitting source files into indexed batches that are processed concurrently by isolated agents, then deterministically merging the per‑batch JSON outputs into a unified knowledge graph while validating completeness and recovering any missing import edges.
The Understand‑Anything repository implements a sophisticated parallel processing architecture to analyze large codebases efficiently. Rather than processing files sequentially, the pipeline distributes workloads across multiple file‑analyzer agents that operate simultaneously on distinct batches, ensuring scalability without race conditions. This article examines the exact mechanisms—from batch indexing to deterministic merging—that enable safe parallel file analyses in the pipeline.
Batch Isolation Strategies for Parallel File Analyses
The orchestrator initiates parallel processing by splitting the full file list into logical batches, assigning each a numeric batchIndex. This index permeates the temporary file naming convention to prevent disk‑level collisions between concurrent agents.
According to the agent definition in understand-anything-plugin/agents/file-analyzer.md (lines 31‑38), every temporary file must embed the batchIndex:
- Input manifest:
ua-file-analyzer-input-<batchIndex>.json - Extraction results:
ua-file-extract-results-<batchIndex>.json
This naming guarantees that concurrently running agents never clash on disk, even when hundreds of batches execute simultaneously.
Deterministic Structural Extraction per Batch
Each agent executes the bundled extract-structure.mjs script on its assigned batch. The script performs a one‑time initialization of all Tree‑Sitter grammars via await tsPlugin.init(), then iterates over the batch’s files calling registry.analyzeFile (and registry.extractCallGraph for code or script files).
As implemented in understand-anything-plugin/skills/understand/extract-structure.mjs (lines 78‑101), this approach is pure JavaScript/Node and stateless between files, allowing multiple instances to run in parallel processes without interference. The grammars are loaded once per batch (via Promise.all(loadPromises) in understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts), avoiding repeated I/O while ensuring each batch has its own isolated context.
Strict Output Naming and Validation
After processing, each agent writes exactly one JSON file per batch, or multiple parts when node/edge counts exceed thresholds. The orchestrator enforces strict filename patterns; any deviation causes the file to be silently dropped during assembly.
As defined in understand-anything-plugin/agents/file-analyzer.md (lines 82‑88), valid outputs must match:
batch-<batchIndex>.jsonbatch-<batchIndex>-part-<k>.json
This convention allows the downstream merge step to unambiguously identify and sequence batch results.
The Merge Phase: Deterministic Assembly
Once all agents finish, the merge-batch-graphs.py script discovers every batch-*.json file, sorts them by numeric index, and merges nodes and edges into a single knowledge graph. The script validates that multipart files are contiguous and that no unrecognized filenames were produced.
The logic in understand-anything-plugin/skills/understand/merge-batch-graphs.py (lines 15‑30 and 65‑78) generates explicit warnings for missing parts or malformed filenames, enabling early detection of parallel‑execution glitches in automated pipelines.
Import Edge Recovery for Graph Completeness
To guarantee graph integrity regardless of batch processing errors, the merge script rescues any imports edges that individual agents might have omitted. It consults the deterministic importMap produced by the project‑scanner to recover missing relationships.
The recover_imports_from_scan function in merge-batch-graphs.py (lines 40‑48) runs after the initial merge, ensuring the final graph contains a complete set of internal import relationships even when specific batches fail to enumerate them correctly.
Practical Implementation Examples
The following example demonstrates the end‑to‑end flow for a single batch, which scales horizontally when executed in parallel across many agents:
# 1️⃣ Write the batch input (the orchestrator does this)
cat > $PROJECT_ROOT/.understand-anything/tmp/ua-file-analyzer-input-3.json <<'ENDJSON'
{
"projectRoot": "/path/to/project",
"batchFiles": [
{"path":"src/index.ts","language":"typescript","sizeLines":120,"fileCategory":"code"},
{"path":"README.md","language":"markdown","sizeLines":30,"fileCategory":"docs"}
],
"batchImportData": {
"src/index.ts": ["src/util.ts","src/config.ts"]
}
}
ENDJSON
# 2️⃣ Execute the deterministic extractor for this batch
node extract-structure.mjs \
$PROJECT_ROOT/.understand-anything/tmp/ua-file-analyzer-input-3.json \
$PROJECT_ROOT/.understand-anything/tmp/ua-file-extract-results-3.json
# 3️⃣ The agent converts the extractor output into the required batch-N JSON
# (handled automatically by the agent logic – see file-analyzer.md Step 2)
# 4️⃣ After all batches finish, merge them
python merge-batch-graphs.py $PROJECT_ROOT
Programmatically, the merge script groups files by numeric index to ensure deterministic ordering:
import json, re
from pathlib import Path
intermediate = Path("/my/project/.understand-anything/intermediate")
batch_files = sorted(
intermediate.glob("batch-*.json"),
key=lambda p: int(re.search(r"batch-(\d+)", p.stem).group(1))
)
# Load each batch, ignoring any file that doesn't match the strict pattern
batches = [json.loads(p.read_text()) for p in batch_files
if re.match(r"batch-(\d+)(?:-part-(\d+))?\.json", p.name)]
assembled, report = merge_and_normalize(batches) # from merge-batch-graphs.py
print("\n".join(report)) # Shows any missing parts / dropped files
Summary
- Indexed batch files prevent race conditions by embedding the
batchIndexinto every temporary file path. - One‑time grammar initialization in
extract-structure.mjsminimizes I/O overhead while keeping batches isolated. - Strict naming conventions (
batch-<N>.jsonorbatch-<N>-part-<K>.json) ensure the merge script can deterministically assemble results. - Post‑merge validation surfaces missing parts or malformed outputs as explicit warnings.
- Import edge recovery guarantees graph completeness by consulting the central
importMapafter merging.
Frequently Asked Questions
What prevents race conditions when running parallel file analyses?
Race conditions are avoided through disk isolation via batch indexing. Each parallel agent writes to temporary files that embed its unique batchIndex (e.g., ua-file-analyzer-input-3.json), ensuring no two agents attempt to write to the same path during concurrent execution.
How does the pipeline handle partial or failed batch outputs?
The merge-batch-graphs.py script validates all discovered files against the strict batch-<N>.json or batch-<N>-part-<K>.json pattern. Missing parts or unrecognized filenames generate explicit warnings in the final report, allowing the orchestrator to identify which batches failed without corrupting the assembled graph.
Why does each batch load Tree‑Sitter grammars only once?
Each extract-structure.mjs instance calls await tsPlugin.init() once to load all grammars via Promise.all(loadPromises) in tree-sitter-plugin.ts before iterating over its assigned files. This minimizes I/O overhead while maintaining process isolation, as each batch runs in its own independent Node.js process.
What happens if import edges are missing from a batch?
The merge phase executes recover_imports_from_scan to rescue omitted imports edges by consulting the deterministic importMap produced by the project scanner. This ensures the final knowledge graph contains complete internal import relationships regardless of individual batch processing errors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →