How merge-batch-graphs.py Deduplicates Nodes and Normalizes IDs in Understand Anything
The merge-batch-graphs.py script in Egonex-AI/Understand-Anything canonicalizes malformed node identifiers using normalize_node_id and removes duplicates by keeping the last occurrence in a dictionary-based last-wins strategy.
merge-batch-graphs.py serves as the final aggregation phase of the /understand pipeline in the Egonex-AI/Understand-Anything repository. It consumes the per-batch JSON files generated by file-analyzer agents, corrects identifier inconsistencies, and emits a single assembled knowledge graph. This article examines exactly how the script handles ID normalization and node deduplication based on the source implementation.
ID Normalization Strategy
The script fixes malformed node IDs before deduplication to ensure canonical identifiers across all batches. In understand-anything-plugin/skills/understand/merge-batch-graphs.py, the normalize_node_id function (lines 777-822) applies a series of regex-based transformations to handle four specific corruption patterns.
Handling Double Prefixes and Project Names
When agents emit nodes with duplicate type prefixes or project-specific prefixes, the function strips the redundant segments:
- Double prefix (e.g.,
file:file:src/foo.ts): The redundantfile:is removed, leavingfile:src/foo.ts. - Project-name prefix (e.g.,
my-project:file:src/foo.ts): A regex matches the pattern<project-name>:<valid-prefix>:...and drops the leading project name, retaining onlyfile:src/foo.ts.
Legacy Prefix Migration and Missing Prefix Resolution
The normalizer addresses historical identifier formats and incomplete data:
- Legacy
func:prefix: Converted to the currentfunction:standard. - Missing prefix (bare paths like
src/main.py): The script consults the node'stypefield against theTYPE_TO_PREFIXmapping. If the node represents a function or class but lacks a file path, the placeholder__nofilepath__is inserted to prevent silent collisions.
Rewriting Edge References
After normalization, the script records any ID changes in an id_mapping dictionary (lines 996-1004). This mapping is then applied to all edge source and target fields within the merge_and_normalize function, ensuring graph connectivity remains intact after identifiers are corrected.
Node Deduplication Mechanism
Once IDs are normalized, the script eliminates redundant nodes that share the same canonical identifier across different batches.
The Last-Wins Dictionary Approach
The deduplication logic (lines 996-1004) uses a simple but deterministic dictionary accumulation:
nodes_by_id: dict[str, dict] = {}
for node in nodes_with_ids:
nid = node.get("id", "")
if nid in nodes_by_id:
duplicate_count += 1
nodes_by_id[nid] = node # Last occurrence wins
This approach guarantees that when multiple batches contain the same node (e.g., file:src/index.ts), the version from the last processed batch is retained.
Duplicate Counting and Reporting
The script tracks the number of duplicates removed and includes this metric in the final console report. As shown in the deduplication block, duplicate_count increments whenever an existing key is overwritten, providing transparency about data reduction during the merge process.
Running the Merge Process
Execute the script from the repository root to process intermediate batch files:
python understand-anything-plugin/skills/understand/merge-batch-graphs.py /path/to/project
The script reads .understand-anything/intermediate/batch-*.json, applies the normalization and deduplication logic, and writes the result to .understand-anything/intermediate/assembled-graph.json.
Example Transformations
Consider these malformed input nodes:
[
{"id": "file:file:src/app.js", "type": "file"},
{"id": "myproj:file:src/app.js", "type": "file"},
{"id": "func:myFunction", "type": "function", "filePath": "src/app.js", "name": "myFunction"}
]
After passing through normalize_node_id, they become:
[
{"id": "file:src/app.js", "type": "file"},
{"id": "file:src/app.js", "type": "file"},
{"id": "function:file:src/app.js:myFunction", "type": "function"}
]
The deduplication step then keeps only the second file:src/app.js node, removing the earlier duplicate.
Summary
normalize_node_id(lines 777-822) fixes double prefixes, project prefixes, legacyfunc:labels, and missing prefixes using theTYPE_TO_PREFIXmapping.id_mappingtracks all ID corrections so edges can be rewritten to point to normalized targets.- Deduplication (lines 996-1004) uses a dictionary with a last-wins policy to ensure unique nodes while counting removals for the final report.
- The script outputs
assembled-graph.jsoncontaining clean, canonical nodes consistent across all batches.
Frequently Asked Questions
What is the difference between ID normalization and node deduplication in merge-batch-graphs.py?
ID normalization corrects malformed identifier strings (e.g., removing duplicate prefixes like file:file:) to ensure every node has a canonical ID. Node deduplication occurs after normalization and removes multiple nodes that share the same canonical ID, keeping only the last occurrence. Normalization happens first because deduplication relies on consistent identifier strings to detect true duplicates.
How does merge-batch-graphs.py handle conflicting node IDs from different batches?
The script processes batches sequentially and uses a dictionary-based accumulation where the last occurrence of a node ID overwrites previous entries. According to the implementation in lines 996-1004, when a duplicate ID is encountered, the duplicate_count increments, but the new node data replaces the old entry in nodes_by_id. This deterministic "last-wins" approach ensures predictable output regardless of batch ordering.
Why does the script use a "last-wins" strategy for duplicate nodes?
The last-wins strategy in merge-batch-graphs.py provides deterministic behavior without requiring complex merge logic for node attributes. Since the script cannot automatically reconcile conflicting properties between duplicate nodes (e.g., different metadata from two analyzers), it assumes later batches contain more accurate or complete information. This approach also matches the sequential processing model of the pipeline.
What happens to edges when node IDs are normalized?
During the merge_and_normalize phase, the script builds an id_mapping dictionary that tracks every original ID and its normalized form. After all nodes are processed, the script iterates through all edges and rewrites their source and target fields using this mapping. This ensures that edges always point to the corrected, canonical node identifiers even after aggressive normalization of malformed prefixes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →