# How the File-Analyzer Handles Cross-Batch Edges in Understand-Anything

> Learn how the file-analyzer resolves cross-batch edges by normalizing node IDs, rewriting endpoints, and safely dropping missing references. Understand the process clearly.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: internals
- Published: 2026-06-10

---

**The file-analyzer resolves cross-batch edges by normalizing node IDs across all batches, maintaining a lookup map to rewrite edge endpoints, and safely dropping any edges that reference missing nodes after ID canonicalization.**

The Understand-Anything repository processes large codebases by splitting analysis into independent batches of approximately three code nodes each. Because each batch operates in isolation, edges often reference nodes that reside in different batches. The file-analyzer's `normalizeBatchOutput` function, implemented in [`packages/core/src/analyzer/normalize-graph.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/analyzer/normalize-graph.ts), stitches these fragmented outputs into a coherent knowledge graph through a three-phase normalization pipeline.

## The Challenge of Cross-Batch References

When a batch completes, it emits a set of **nodes** (files, functions, classes, steps) and **edges** linking those nodes. Since batches run independently, an edge in batch A might target a node that only exists in batch B. Without intervention, these references would remain dangling or use inconsistent identifiers. The analyzer solves this by collecting all nodes first, building a canonical ID mapping, then rewriting every edge to use those canonical IDs.

## Three-Phase Resolution Process

The `normalizeBatchOutput` routine processes cross-batch edges through three distinct phases before passing the result to the higher-level sanitization pipeline (`sanitizeGraph` → `autoFixGraph` → `normalizeGraph`).

### Phase 1: Collect Flow-Step Information

Before altering any identifiers, the analyzer scans all nodes and edges to construct a `stepToFlowSlug` map. This map records which **flow** each step node belongs to, ensuring that subsequent ID normalization can embed the flow slug into step identifiers. This prevents collisions when two different flows contain steps with identical names.

```typescript
// stepToFlowSlug is built by scanning all nodes
const stepToFlowSlug = new Map<string, string>();
// Maps step node IDs to their parent flow slugs

```

### Phase 2: Normalize Node IDs

Each node's `id` is passed to the `normalizeNodeId` helper function. This process:

- **Preserves known prefixes**: IDs already containing `file:`, `func:`, or `class:` remain unchanged.
- **Prefixes bare paths**: Raw file paths like [`src/utils.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/utils.ts) become `file:src/utils.ts`.
- **Embeds flow context**: Step nodes are transformed into `step:<flowSlug>:<filePath>:<stepSlug>` using the map from Phase 1.
- **Fixes malformed IDs**: Removes double prefixes and project-name prefixes, normalizes complexity values.

A map called `idMap` records every old-to-new ID mapping, while a `idsFixed` counter tracks how many identifiers required modification.

```typescript
const idMap = new Map<string, string>();
let idsFixed = 0;

// Example transformation:
// 'src/foo.ts' → 'file:src/foo.ts'
// 'func:bar' → 'func:src/bar.ts:bar'

```

### Phase 3: Rewrite and Deduplicate Edges

In the final phase, the analyzer iterates through every edge and rewrites its `source` and `target` properties using the `idMap`. When an endpoint is missing from the map (because the referring batch did not contain that node), the analyzer:

1. Infers the node type from the ID prefix or defaults to `file`.
2. Runs `normalizeNodeId` on the raw endpoint.
3. Checks if the newly normalized ID exists in the valid node set.

If a rewrite occurs, the `edgesRewritten` counter increments. If after normalization an endpoint still cannot be resolved, the edge is dropped and recorded in `droppedEdges` with a reason (`missing-source`, `missing-target`, or `missing-both`). Valid edges are deduplicated using the composite key `source|target|type`.

## Handling Missing Endpoints

Not all cross-batch edges survive normalization. When the analyzer encounters an edge referencing a node that never appears in any batch (or whose ID cannot be canonicalized), it removes the edge rather than leave a dangling reference. The `droppedEdges` array captures these casualties along with their failure reasons, allowing developers to audit which relationships could not be resolved.

## Practical Example: Merging Batch Outputs

The following example demonstrates how two independent batches merge into a single coherent graph with resolved cross-batch edges:

```typescript
import { normalizeBatchOutput } from '@understand-anything/core/src/analyzer/normalize-graph.js';

// Batch A knows about a file but references a function it doesn't contain
const batchA = {
  nodes: [{ id: 'src/foo.ts', type: 'file' }],
  edges: [{ source: 'src/foo.ts', target: 'func:bar', type: 'contains' }],
};

// Batch B contains the actual function node
const batchB = {
  nodes: [{ id: 'func:bar', type: 'function', filePath: 'src/bar.ts', name: 'bar' }],
  edges: [],
};

// Merge and normalize
const merged = normalizeBatchOutput({
  nodes: [...batchA.nodes, ...batchB.nodes],
  edges: [...batchA.edges, ...batchB.edges],
});

console.log(merged.nodes);
// → [
//   { id: 'file:src/foo.ts', type: 'file' },
//   { id: 'func:src/bar.ts:bar', type: 'function', ... }
// ]

console.log(merged.edges);
// → [{
//   source: 'file:src/foo.ts',
//   target: 'func:src/bar.ts:bar',
//   type: 'contains'
// }]

```

In this example, the edge from batch A referencing `func:bar` is **rewritten** to point to the canonical ID of the function node discovered in batch B. The statistics object in `merged.stats` would show `edgesRewritten: 1` and any dropped edges would appear in `merged.stats.droppedEdges`.

## Summary

- **Node ID normalization** happens in [`packages/core/src/analyzer/normalize-graph.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/analyzer/normalize-graph.ts) via `normalizeNodeId`, which canonicalizes prefixes and embeds flow context for steps.
- **Cross-batch edge resolution** uses an `idMap` to rewrite `source` and `target` properties, with the `normalizeBatchOutput` function coordinating the three-phase pipeline.
- **Missing endpoints** are handled gracefully by inferring types, attempting late normalization, and recording failures in `droppedEdges` with specific reason codes.
- **Deduplication** occurs automatically using the composite key `source|target|type` before the graph enters the sanitization pipeline.

## Frequently Asked Questions

### What happens when a cross-batch edge references a node that doesn't exist in any batch?

The analyzer attempts to normalize the missing endpoint by inferring its type from the ID prefix or defaulting to `file`. If the normalized ID still does not exist in the valid node set, the edge is **dropped** and recorded in `droppedEdges` with a reason such as `missing-source` or `missing-target`. This prevents dangling references from polluting the final graph.

### How does the analyzer prevent ID collisions between steps in different flows?

During Phase 1, the analyzer builds a `stepToFlowSlug` map that associates each step node with its parent flow. When normalizing step IDs in Phase 2, it embeds the flow slug into the identifier using the format `step:<flowSlug>:<filePath>:<stepSlug>`. This ensures that identically named steps in different flows receive unique, canonical IDs.

### Where can I find the tests that validate cross-batch edge handling?

The test suite in [`packages/core/__tests__/tour-generator.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/__tests__/tour-generator.test.ts) contains scenarios that exercise cross-batch edge resolution. These tests verify that the `normalizeBatchOutput` function correctly rewrites edges, handles missing nodes, and maintains the `idMap` and statistics counters (`idsFixed`, `edgesRewritten`, `droppedEdges`) as expected when merging multiple batch outputs.

### What statistics does the file-analyzer track during the normalization process?

The analyzer maintains three primary counters: `idsFixed` (incremented when node IDs require canonicalization), `edgesRewritten` (incremented when edge endpoints are remapped), and `droppedEdges` (an array tracking removed edges with their failure reasons). These metrics provide visibility into how many cross-batch references required intervention versus how many were pruned as unresolvable.