# How Cross-Batch Edges Work in Understand-Anything: Merging Analysis Batches

> Learn how cross-batch edges in Understand Anything use canonical IDs to merge analysis batches, ensuring robust data integrity by discarding only dangling edges.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: internals
- Published: 2026-05-31

---

**Cross-batch edges survive the merge process because the normalization pipeline rewrites edge references to canonical IDs after all batches complete, discarding only truly dangling edges that reference non-existent nodes.**

When analyzing large codebases in the [Understand-Anything](https://github.com/Lum1104/Understand-Anything) repository, the system processes files incrementally using separate batches. Understanding how **cross-batch edges**—relationships that span files analyzed in different batches—are preserved is critical for building accurate knowledge graphs of distributed code analysis.

## Why Batch Processing Requires Cross-Batch Edge Handling

Large projects cannot fit into memory or processing constraints as a single unit. The analyzer therefore partitions work into discrete batches, where each batch builds a partial graph containing nodes (files, functions, classes) and edges (imports, calls) for its specific subset of files.

However, software dependencies do not respect batch boundaries. A file processed in batch 1 might import a file processed in batch 2. Without a merge strategy, these **cross-batch edges** would fragment into disconnected subgraphs. The system solves this through a normalization pipeline that reconciles all batch outputs into a single coherent graph.

## How GraphBuilder Creates Edges Per Batch

Each batch instantiates a **`GraphBuilder`** that constructs nodes and edges independently. According to the source code in [`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts), the builder provides specific methods for creating relationships:

- **`addImportEdge`** (lines 79-87): Creates an `imports` relationship between files.
- **`addCallEdge`** (lines 92-100): Creates a `calls` relationship between functions or methods.

When a batch finishes, it emits a raw payload containing `{ nodes, edges }`. At this stage, edge endpoints reference node IDs that may be reformatted during the final merge.

## The Normalization Pipeline: Merging Batch Outputs

The core logic resides in [`packages/core/src/analyzer/normalize-graph.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/normalize-graph.ts) within the **`normalizeBatchOutput`** function. This utility aggregates all batch results and performs a five-stage rewrite to reconcile cross-batch references.

### Step 1: Building the Canonical ID Map

The function first constructs an `idMap` dictionary that translates old node identifiers into canonical `type:path` formats. The helper `normalizeNodeId` (lines 61-86) ensures every node ID follows the pattern `file:src/path.ts` or `function:src/path.ts:myFunction`, regardless of how individual batches formatted them initially.

This mapping enables the system to recognize that `file:src/a.ts` in batch 1 refers to the same entity as `file:src/a.ts` in batch 5.

### Step 2: Rewriting Edge References

With the `idMap` established, the pipeline iterates through every edge and rewrites its `source` and `target` properties to point to canonical IDs. This rewrite happens **after all batches have been collected**, meaning edges created in batch 1 that target nodes from batch 2 simply lookup the target's canonical ID in the completed map.

### Step 3: Handling Malformed IDs and Dangling Edges

If an edge references an ID not present in `idMap`, the system attempts recovery using **`inferTypeFromId`** combined with `normalizeNodeId` (lines 86-92). For example, a malformed ID like `src:a.ts` (missing the `file:` prefix) is inferred as a file type and normalized to `file:src/a.ts`.

Edges that still cannot resolve after this inference step are considered dangling and are removed (lines 101-110). The system categorizes these as `missing-source`, `missing-target`, or `missing-both` and drops them from the final graph.

Finally, the pipeline deduplicates edges using a composite key of `source|target|type` (lines 114-119), ensuring the same relationship discovered in multiple batches appears only once.

## Example: Preserving an Import Across Two Batches

Consider two batches processing related files:

```typescript
// Batch 1 processes src/a.ts and creates an import edge
builder.addImportEdge('src/a.ts', 'src/b.ts');
// Produces: { source: "file:src/a.ts", target: "file:src/b.ts", type: "imports" }

// Batch 2 later processes src/b.ts and creates nodes
builder.addFileWithAnalysis('src/b.ts', analysisB, metaB);
// Produces: { id: "file:src/b.ts", ... }

```

After both batches complete, `normalizeBatchOutput` receives the combined payload:

```json
{
  "nodes": [
    { "id": "file:src/a.ts", "type": "file" },
    { "id": "file:src/b.ts", "type": "file" }
  ],
  "edges": [
    { "source": "file:src/a.ts", "target": "file:src/b.ts", "type": "imports" }
  ]
}

```

The normalization builds `idMap` containing both file IDs, rewrites the edge (no change needed here), and retains it in the merged graph. The **cross-batch edge** survives because both endpoints exist in the final node collection, even though they originated from separate analysis sessions.

If batch 1 had produced a malformed ID like `src:a.ts` instead of `file:src/a.ts`, the rewrite logic would infer the type from the prefix and normalize it before re-attaching the edge, preserving the relationship despite the formatting inconsistency.

## Summary

- **Cross-batch edges** are preserved by rewriting edge references to canonical IDs after all batches complete.
- The **`normalizeBatchOutput`** function in [`packages/core/src/analyzer/normalize-graph.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/normalize-graph.ts) manages the merge process.
- Node IDs are normalized to the `type:path` format using `normalizeNodeId` to ensure consistency across batches.
- Malformed IDs are repaired on-the-fly using `inferTypeFromId` before determining if an edge is truly dangling.
- Only edges referencing nodes that never appear in any batch are dropped; valid cross-batch connections survive deduplication.

## Frequently Asked Questions

### What happens if a batch references a file that was never analyzed?

If an edge points to a node ID that does not exist in any batch's output and cannot be normalized from the ID structure itself, the system classifies it as a dangling edge. According to lines 101-110 in [`normalize-graph.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/normalize-graph.ts), these edges are removed and logged with a specific reason such as `missing-target` or `missing-source`.

### How does the system handle malformed node IDs in cross-batch edges?

The normalization pipeline attempts to infer the node type from the ID string using **`inferTypeFromId`** (lines 86-92). For example, an ID missing its `file:` prefix can be repaired to the canonical format. This allows edges with inconsistent formatting to reconnect to their proper nodes rather than being discarded as dangling.

### Are duplicate edges between the same nodes preserved?

No. The deduplication step (lines 114-119) creates a composite key from `source|target|type` to ensure that if the same relationship (such as an import) is discovered in multiple batches, it appears only once in the final knowledge graph.

### Why not process all files in a single batch?

The batch architecture supports incremental analysis of very large projects that exceed memory constraints or require distributed processing. By processing files in chunks and relying on the normalization pipeline to reconcile **cross-batch edges**, the system can scale to enterprise-sized codebases without requiring unbounded resources for a single monolithic analysis pass.