How Cross-Batch Edges Work in Understand-Anything: Merging Analysis Batches
Cross-batch edges survive the merge process because the normalization pipeline rewrites edge references to canonical IDs after all batches complete, discarding only truly dangling edges that reference non-existent nodes.
When analyzing large codebases in the Understand-Anything repository, the system processes files incrementally using separate batches. Understanding how cross-batch edges—relationships that span files analyzed in different batches—are preserved is critical for building accurate knowledge graphs of distributed code analysis.
Why Batch Processing Requires Cross-Batch Edge Handling
Large projects cannot fit into memory or processing constraints as a single unit. The analyzer therefore partitions work into discrete batches, where each batch builds a partial graph containing nodes (files, functions, classes) and edges (imports, calls) for its specific subset of files.
However, software dependencies do not respect batch boundaries. A file processed in batch 1 might import a file processed in batch 2. Without a merge strategy, these cross-batch edges would fragment into disconnected subgraphs. The system solves this through a normalization pipeline that reconciles all batch outputs into a single coherent graph.
How GraphBuilder Creates Edges Per Batch
Each batch instantiates a GraphBuilder that constructs nodes and edges independently. According to the source code in packages/core/src/analyzer/graph-builder.ts, the builder provides specific methods for creating relationships:
addImportEdge(lines 79-87): Creates animportsrelationship between files.addCallEdge(lines 92-100): Creates acallsrelationship between functions or methods.
When a batch finishes, it emits a raw payload containing { nodes, edges }. At this stage, edge endpoints reference node IDs that may be reformatted during the final merge.
The Normalization Pipeline: Merging Batch Outputs
The core logic resides in packages/core/src/analyzer/normalize-graph.ts within the normalizeBatchOutput function. This utility aggregates all batch results and performs a five-stage rewrite to reconcile cross-batch references.
Step 1: Building the Canonical ID Map
The function first constructs an idMap dictionary that translates old node identifiers into canonical type:path formats. The helper normalizeNodeId (lines 61-86) ensures every node ID follows the pattern file:src/path.ts or function:src/path.ts:myFunction, regardless of how individual batches formatted them initially.
This mapping enables the system to recognize that file:src/a.ts in batch 1 refers to the same entity as file:src/a.ts in batch 5.
Step 2: Rewriting Edge References
With the idMap established, the pipeline iterates through every edge and rewrites its source and target properties to point to canonical IDs. This rewrite happens after all batches have been collected, meaning edges created in batch 1 that target nodes from batch 2 simply lookup the target's canonical ID in the completed map.
Step 3: Handling Malformed IDs and Dangling Edges
If an edge references an ID not present in idMap, the system attempts recovery using inferTypeFromId combined with normalizeNodeId (lines 86-92). For example, a malformed ID like src:a.ts (missing the file: prefix) is inferred as a file type and normalized to file:src/a.ts.
Edges that still cannot resolve after this inference step are considered dangling and are removed (lines 101-110). The system categorizes these as missing-source, missing-target, or missing-both and drops them from the final graph.
Finally, the pipeline deduplicates edges using a composite key of source|target|type (lines 114-119), ensuring the same relationship discovered in multiple batches appears only once.
Example: Preserving an Import Across Two Batches
Consider two batches processing related files:
// Batch 1 processes src/a.ts and creates an import edge
builder.addImportEdge('src/a.ts', 'src/b.ts');
// Produces: { source: "file:src/a.ts", target: "file:src/b.ts", type: "imports" }
// Batch 2 later processes src/b.ts and creates nodes
builder.addFileWithAnalysis('src/b.ts', analysisB, metaB);
// Produces: { id: "file:src/b.ts", ... }
After both batches complete, normalizeBatchOutput receives the combined payload:
{
"nodes": [
{ "id": "file:src/a.ts", "type": "file" },
{ "id": "file:src/b.ts", "type": "file" }
],
"edges": [
{ "source": "file:src/a.ts", "target": "file:src/b.ts", "type": "imports" }
]
}
The normalization builds idMap containing both file IDs, rewrites the edge (no change needed here), and retains it in the merged graph. The cross-batch edge survives because both endpoints exist in the final node collection, even though they originated from separate analysis sessions.
If batch 1 had produced a malformed ID like src:a.ts instead of file:src/a.ts, the rewrite logic would infer the type from the prefix and normalize it before re-attaching the edge, preserving the relationship despite the formatting inconsistency.
Summary
- Cross-batch edges are preserved by rewriting edge references to canonical IDs after all batches complete.
- The
normalizeBatchOutputfunction inpackages/core/src/analyzer/normalize-graph.tsmanages the merge process. - Node IDs are normalized to the
type:pathformat usingnormalizeNodeIdto ensure consistency across batches. - Malformed IDs are repaired on-the-fly using
inferTypeFromIdbefore determining if an edge is truly dangling. - Only edges referencing nodes that never appear in any batch are dropped; valid cross-batch connections survive deduplication.
Frequently Asked Questions
What happens if a batch references a file that was never analyzed?
If an edge points to a node ID that does not exist in any batch's output and cannot be normalized from the ID structure itself, the system classifies it as a dangling edge. According to lines 101-110 in normalize-graph.ts, these edges are removed and logged with a specific reason such as missing-target or missing-source.
How does the system handle malformed node IDs in cross-batch edges?
The normalization pipeline attempts to infer the node type from the ID string using inferTypeFromId (lines 86-92). For example, an ID missing its file: prefix can be repaired to the canonical format. This allows edges with inconsistent formatting to reconnect to their proper nodes rather than being discarded as dangling.
Are duplicate edges between the same nodes preserved?
No. The deduplication step (lines 114-119) creates a composite key from source|target|type to ensure that if the same relationship (such as an import) is discovered in multiple batches, it appears only once in the final knowledge graph.
Why not process all files in a single batch?
The batch architecture supports incremental analysis of very large projects that exceed memory constraints or require distributed processing. By processing files in chunks and relying on the normalization pipeline to reconcile cross-batch edges, the system can scale to enterprise-sized codebases without requiring unbounded resources for a single monolithic analysis pass.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →