How the Graph Normalization Process Handles Deduplication and Edge Rewriting in Egonex-AI
The graph normalization process in Egonex-AI eliminates duplicate nodes by retaining the last occurrence of each ID, rewrites edge endpoints using a canonical ID mapping with fallback type inference, and removes dangling edges while deduplicating identical source-target-type triples.
The normalizeBatchOutput() function in understand-anything-plugin/packages/core/src/analyzer/normalize-graph.ts serves as the critical sanitization layer within the Egonex-AI/Understand-Anything repository. This graph normalization process transforms raw merged graphs—produced by the scan-project and merge-batch-graphs skills—into clean, validated structures by resolving identifier collisions and repairing broken references before the schema validator runs.
Node Deduplication Strategy
Collecting Canonical IDs
The deduplication algorithm begins by building a Map that records the last index of every node ID encountered in the input array. As implemented in lines 66-73 of normalize-graph.ts, the code iterates through the nodes array and overwrites the stored index for each encountered ID:
const seenIds = new Map<string, number>();
for (let i = 0; i < nodes.length; i++) {
seenIds.set(String(nodes[i].id), i);
}
Filtering Duplicate Entries
After mapping IDs to their final positions, the function filters the array to retain only those nodes whose current index matches the stored index. This last-write-wins approach effectively preserves the final duplicate while discarding all earlier occurrences:
const deduped = nodes.filter((_, i) => seenIds.get(String(nodes[i].id)) === i);
The resulting validNodeIds Set then serves as the authoritative registry for subsequent edge validation.
Edge Rewriting and Validation
Mapping Legacy IDs to Canonical References
Edge rewriting occurs in a second pass that translates stale source and target identifiers using an idMap populated during the node normalization phase. The rewriter attempts to resolve each endpoint through the canonical mapping:
let newSource = idMap.get(oldSource) ?? oldSource;
let newTarget = idMap.get(oldTarget) ?? oldTarget;
Fallback Normalization for Stale Endpoints
When a mapped ID does not exist in validNodeIds, the system attempts on-the-fly normalization using type inference. According to lines 86-95 in the source file, the code extracts the inferred type from the malformed ID, applies normalizeNodeId(), and accepts the result only if it matches a known valid node:
if (!validNodeIds.has(newSource)) {
const inferredType = inferTypeFromId(newSource);
const normalized = normalizeNodeId(newSource, { type: inferredType });
if (validNodeIds.has(normalized)) newSource = normalized;
}
Removing Dangling Edges
Edges that cannot resolve to valid nodes after the fallback attempt are recorded in stats.droppedEdges with a specific reason code (such as "missing-source") and excluded from the final graph. This prevents orphaned references from corrupting the knowledge graph integrity.
Deduplicating Edge Triples
The final edge-processing stage eliminates redundant connections using a composite key strategy. Lines 156-163 construct a unique signature from the source, target, and edge type, storing it in a Set called seenEdges:
const edgeKey = `${newSource}|${newTarget}|${edgeType}`;
if (seenEdges.has(edgeKey)) continue; // duplicate → drop
seenEdges.add(edgeKey);
Implementation Walkthrough
The following example demonstrates how normalizeBatchOutput() processes a raw graph containing duplicate nodes and malformed edges:
import { normalizeBatchOutput } from "understand-anything-plugin/packages/core/src/analyzer/normalize-graph";
const raw = {
nodes: [
{ id: "file:src/foo.ts", type: "file", filePath: "src/foo.ts" },
{ id: "file:src/foo.ts", type: "file", filePath: "src/foo.ts", extra: true }, // duplicate
{ id: "func:src/bar.ts:oldName", type: "function", filePath: "src/bar.ts", name: "oldName" },
],
edges: [
{ source: "file:src/foo.ts", target: "func:src/bar.ts:oldName", type: "calls" },
{ source: "src/missing.ts", target: "func:src/bar.ts:oldName", type: "calls" }, // malformed
],
};
const { nodes, edges, stats } = normalizeBatchOutput(raw);
console.log(nodes.length); // 2 (duplicate removed)
console.log(stats.edgesRewritten); // 2
console.log(stats.danglingEdgesDropped); // 1
Integration with the Analysis Pipeline
Within the merge-batch-graphs skill, the normalizer acts as the bridge between raw graph construction and schema validation. After flattening batch outputs into a merged structure, the system invokes normalizeBatchOutput() to produce the cleaned dataset passed to sanitizeGraph in packages/core/src/schema.ts:
export async function mergeBatchGraphs(batchOutputs) {
const merged = {
nodes: batchOutputs.flatMap(b => b.nodes),
edges: batchOutputs.flatMap(b => b.edges),
};
const { nodes, edges, idMap, stats } = normalizeBatchOutput(merged);
return { nodes, edges, idMap, stats };
}
Summary
- The graph normalization process retains the last occurrence of duplicate nodes by indexing IDs against their final positions in the array.
- Edge rewriting relies on an
idMapfor primary translation, with fallback normalization usinginferTypeFromId()andnormalizeNodeId()for orphaned references. - Dangling edges—those with unresolvable endpoints—are logged to
stats.droppedEdgesand excluded from output. - Edge deduplication uses a composite key (
source|target|type) to ensure unique relationships in the final graph. - The system returns comprehensive statistics including
idsFixed,edgesRewritten, anddanglingEdgesDroppedto audit the transformation.
Frequently Asked Questions
How does the deduplication algorithm decide which node to keep?
The algorithm uses a last-write-wins strategy. By storing the index of each encountered ID in a Map and overwriting previous entries, then filtering to keep only nodes whose current index matches the stored index, the system preserves the final duplicate while discarding all earlier instances.
What happens to edges that reference deleted or non-existent nodes?
Edges with endpoints that cannot be resolved through the idMap or fallback normalization are classified as dangling. These edges are recorded in stats.droppedEdges with a reason code (such as "missing-source") and are omitted from the final edges array.
Can the normalization process fix malformed node IDs during edge rewriting?
Yes. When an edge references an ID not present in validNodeIds, the system attempts repair by inferring the type from the ID string, applying normalizeNodeId() with the inferred type, and validating the result against the known node set. If successful, the edge endpoint is rewritten to the corrected canonical ID.
Where does normalizeBatchOutput fit in the Understand-Anything plugin pipeline?
The function executes immediately after the merge-batch-graphs skill combines raw subgraphs and before the schema sanitizer runs. Located in packages/core/src/analyzer/normalize-graph.ts, it serves as the critical sanitization layer that ensures graph integrity prior to dashboard consumption and persistence.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →