How the File-Analyzer Handles Cross-Batch Edges in Understand-Anything
The file-analyzer resolves cross-batch edges by normalizing node IDs across all batches, maintaining a lookup map to rewrite edge endpoints, and safely dropping any edges that reference missing nodes after ID canonicalization.
The Understand-Anything repository processes large codebases by splitting analysis into independent batches of approximately three code nodes each. Because each batch operates in isolation, edges often reference nodes that reside in different batches. The file-analyzer's normalizeBatchOutput function, implemented in packages/core/src/analyzer/normalize-graph.ts, stitches these fragmented outputs into a coherent knowledge graph through a three-phase normalization pipeline.
The Challenge of Cross-Batch References
When a batch completes, it emits a set of nodes (files, functions, classes, steps) and edges linking those nodes. Since batches run independently, an edge in batch A might target a node that only exists in batch B. Without intervention, these references would remain dangling or use inconsistent identifiers. The analyzer solves this by collecting all nodes first, building a canonical ID mapping, then rewriting every edge to use those canonical IDs.
Three-Phase Resolution Process
The normalizeBatchOutput routine processes cross-batch edges through three distinct phases before passing the result to the higher-level sanitization pipeline (sanitizeGraph → autoFixGraph → normalizeGraph).
Phase 1: Collect Flow-Step Information
Before altering any identifiers, the analyzer scans all nodes and edges to construct a stepToFlowSlug map. This map records which flow each step node belongs to, ensuring that subsequent ID normalization can embed the flow slug into step identifiers. This prevents collisions when two different flows contain steps with identical names.
// stepToFlowSlug is built by scanning all nodes
const stepToFlowSlug = new Map<string, string>();
// Maps step node IDs to their parent flow slugs
Phase 2: Normalize Node IDs
Each node's id is passed to the normalizeNodeId helper function. This process:
- Preserves known prefixes: IDs already containing
file:,func:, orclass:remain unchanged. - Prefixes bare paths: Raw file paths like
src/utils.tsbecomefile:src/utils.ts. - Embeds flow context: Step nodes are transformed into
step:<flowSlug>:<filePath>:<stepSlug>using the map from Phase 1. - Fixes malformed IDs: Removes double prefixes and project-name prefixes, normalizes complexity values.
A map called idMap records every old-to-new ID mapping, while a idsFixed counter tracks how many identifiers required modification.
const idMap = new Map<string, string>();
let idsFixed = 0;
// Example transformation:
// 'src/foo.ts' → 'file:src/foo.ts'
// 'func:bar' → 'func:src/bar.ts:bar'
Phase 3: Rewrite and Deduplicate Edges
In the final phase, the analyzer iterates through every edge and rewrites its source and target properties using the idMap. When an endpoint is missing from the map (because the referring batch did not contain that node), the analyzer:
- Infers the node type from the ID prefix or defaults to
file. - Runs
normalizeNodeIdon the raw endpoint. - Checks if the newly normalized ID exists in the valid node set.
If a rewrite occurs, the edgesRewritten counter increments. If after normalization an endpoint still cannot be resolved, the edge is dropped and recorded in droppedEdges with a reason (missing-source, missing-target, or missing-both). Valid edges are deduplicated using the composite key source|target|type.
Handling Missing Endpoints
Not all cross-batch edges survive normalization. When the analyzer encounters an edge referencing a node that never appears in any batch (or whose ID cannot be canonicalized), it removes the edge rather than leave a dangling reference. The droppedEdges array captures these casualties along with their failure reasons, allowing developers to audit which relationships could not be resolved.
Practical Example: Merging Batch Outputs
The following example demonstrates how two independent batches merge into a single coherent graph with resolved cross-batch edges:
import { normalizeBatchOutput } from '@understand-anything/core/src/analyzer/normalize-graph.js';
// Batch A knows about a file but references a function it doesn't contain
const batchA = {
nodes: [{ id: 'src/foo.ts', type: 'file' }],
edges: [{ source: 'src/foo.ts', target: 'func:bar', type: 'contains' }],
};
// Batch B contains the actual function node
const batchB = {
nodes: [{ id: 'func:bar', type: 'function', filePath: 'src/bar.ts', name: 'bar' }],
edges: [],
};
// Merge and normalize
const merged = normalizeBatchOutput({
nodes: [...batchA.nodes, ...batchB.nodes],
edges: [...batchA.edges, ...batchB.edges],
});
console.log(merged.nodes);
// → [
// { id: 'file:src/foo.ts', type: 'file' },
// { id: 'func:src/bar.ts:bar', type: 'function', ... }
// ]
console.log(merged.edges);
// → [{
// source: 'file:src/foo.ts',
// target: 'func:src/bar.ts:bar',
// type: 'contains'
// }]
In this example, the edge from batch A referencing func:bar is rewritten to point to the canonical ID of the function node discovered in batch B. The statistics object in merged.stats would show edgesRewritten: 1 and any dropped edges would appear in merged.stats.droppedEdges.
Summary
- Node ID normalization happens in
packages/core/src/analyzer/normalize-graph.tsvianormalizeNodeId, which canonicalizes prefixes and embeds flow context for steps. - Cross-batch edge resolution uses an
idMapto rewritesourceandtargetproperties, with thenormalizeBatchOutputfunction coordinating the three-phase pipeline. - Missing endpoints are handled gracefully by inferring types, attempting late normalization, and recording failures in
droppedEdgeswith specific reason codes. - Deduplication occurs automatically using the composite key
source|target|typebefore the graph enters the sanitization pipeline.
Frequently Asked Questions
What happens when a cross-batch edge references a node that doesn't exist in any batch?
The analyzer attempts to normalize the missing endpoint by inferring its type from the ID prefix or defaulting to file. If the normalized ID still does not exist in the valid node set, the edge is dropped and recorded in droppedEdges with a reason such as missing-source or missing-target. This prevents dangling references from polluting the final graph.
How does the analyzer prevent ID collisions between steps in different flows?
During Phase 1, the analyzer builds a stepToFlowSlug map that associates each step node with its parent flow. When normalizing step IDs in Phase 2, it embeds the flow slug into the identifier using the format step:<flowSlug>:<filePath>:<stepSlug>. This ensures that identically named steps in different flows receive unique, canonical IDs.
Where can I find the tests that validate cross-batch edge handling?
The test suite in packages/core/__tests__/tour-generator.test.ts contains scenarios that exercise cross-batch edge resolution. These tests verify that the normalizeBatchOutput function correctly rewrites edges, handles missing nodes, and maintains the idMap and statistics counters (idsFixed, edgesRewritten, droppedEdges) as expected when merging multiple batch outputs.
What statistics does the file-analyzer track during the normalization process?
The analyzer maintains three primary counters: idsFixed (incremented when node IDs require canonicalization), edgesRewritten (incremented when edge endpoints are remapped), and droppedEdges (an array tracking removed edges with their failure reasons). These metrics provide visibility into how many cross-batch references required intervention versus how many were pruned as unresolvable.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →