How merge-batch-graphs.py Deduplicates Nodes and Normalizes IDs in Understand Anything

The merge-batch-graphs.py script in Egonex-AI/Understand-Anything canonicalizes malformed node identifiers using normalize_node_id and removes duplicates by keeping the last occurrence in a dictionary-based last-wins strategy.

merge-batch-graphs.py serves as the final aggregation phase of the /understand pipeline in the Egonex-AI/Understand-Anything repository. It consumes the per-batch JSON files generated by file-analyzer agents, corrects identifier inconsistencies, and emits a single assembled knowledge graph. This article examines exactly how the script handles ID normalization and node deduplication based on the source implementation.

ID Normalization Strategy

The script fixes malformed node IDs before deduplication to ensure canonical identifiers across all batches. In understand-anything-plugin/skills/understand/merge-batch-graphs.py, the normalize_node_id function (lines 777-822) applies a series of regex-based transformations to handle four specific corruption patterns.

Handling Double Prefixes and Project Names

When agents emit nodes with duplicate type prefixes or project-specific prefixes, the function strips the redundant segments:

  • Double prefix (e.g., file:file:src/foo.ts): The redundant file: is removed, leaving file:src/foo.ts.
  • Project-name prefix (e.g., my-project:file:src/foo.ts): A regex matches the pattern <project-name>:<valid-prefix>:... and drops the leading project name, retaining only file:src/foo.ts.

Legacy Prefix Migration and Missing Prefix Resolution

The normalizer addresses historical identifier formats and incomplete data:

  • Legacy func: prefix: Converted to the current function: standard.
  • Missing prefix (bare paths like src/main.py): The script consults the node's type field against the TYPE_TO_PREFIX mapping. If the node represents a function or class but lacks a file path, the placeholder __nofilepath__ is inserted to prevent silent collisions.

Rewriting Edge References

After normalization, the script records any ID changes in an id_mapping dictionary (lines 996-1004). This mapping is then applied to all edge source and target fields within the merge_and_normalize function, ensuring graph connectivity remains intact after identifiers are corrected.

Node Deduplication Mechanism

Once IDs are normalized, the script eliminates redundant nodes that share the same canonical identifier across different batches.

The Last-Wins Dictionary Approach

The deduplication logic (lines 996-1004) uses a simple but deterministic dictionary accumulation:

nodes_by_id: dict[str, dict] = {}
for node in nodes_with_ids:
    nid = node.get("id", "")
    if nid in nodes_by_id:
        duplicate_count += 1
    nodes_by_id[nid] = node  # Last occurrence wins

This approach guarantees that when multiple batches contain the same node (e.g., file:src/index.ts), the version from the last processed batch is retained.

Duplicate Counting and Reporting

The script tracks the number of duplicates removed and includes this metric in the final console report. As shown in the deduplication block, duplicate_count increments whenever an existing key is overwritten, providing transparency about data reduction during the merge process.

Running the Merge Process

Execute the script from the repository root to process intermediate batch files:

python understand-anything-plugin/skills/understand/merge-batch-graphs.py /path/to/project

The script reads .understand-anything/intermediate/batch-*.json, applies the normalization and deduplication logic, and writes the result to .understand-anything/intermediate/assembled-graph.json.

Example Transformations

Consider these malformed input nodes:

[
  {"id": "file:file:src/app.js", "type": "file"},
  {"id": "myproj:file:src/app.js", "type": "file"},
  {"id": "func:myFunction", "type": "function", "filePath": "src/app.js", "name": "myFunction"}
]

After passing through normalize_node_id, they become:

[
  {"id": "file:src/app.js", "type": "file"},
  {"id": "file:src/app.js", "type": "file"},
  {"id": "function:file:src/app.js:myFunction", "type": "function"}
]

The deduplication step then keeps only the second file:src/app.js node, removing the earlier duplicate.

Summary

  • normalize_node_id (lines 777-822) fixes double prefixes, project prefixes, legacy func: labels, and missing prefixes using the TYPE_TO_PREFIX mapping.
  • id_mapping tracks all ID corrections so edges can be rewritten to point to normalized targets.
  • Deduplication (lines 996-1004) uses a dictionary with a last-wins policy to ensure unique nodes while counting removals for the final report.
  • The script outputs assembled-graph.json containing clean, canonical nodes consistent across all batches.

Frequently Asked Questions

What is the difference between ID normalization and node deduplication in merge-batch-graphs.py?

ID normalization corrects malformed identifier strings (e.g., removing duplicate prefixes like file:file:) to ensure every node has a canonical ID. Node deduplication occurs after normalization and removes multiple nodes that share the same canonical ID, keeping only the last occurrence. Normalization happens first because deduplication relies on consistent identifier strings to detect true duplicates.

How does merge-batch-graphs.py handle conflicting node IDs from different batches?

The script processes batches sequentially and uses a dictionary-based accumulation where the last occurrence of a node ID overwrites previous entries. According to the implementation in lines 996-1004, when a duplicate ID is encountered, the duplicate_count increments, but the new node data replaces the old entry in nodes_by_id. This deterministic "last-wins" approach ensures predictable output regardless of batch ordering.

Why does the script use a "last-wins" strategy for duplicate nodes?

The last-wins strategy in merge-batch-graphs.py provides deterministic behavior without requiring complex merge logic for node attributes. Since the script cannot automatically reconcile conflicting properties between duplicate nodes (e.g., different metadata from two analyzers), it assumes later batches contain more accurate or complete information. This approach also matches the sequential processing model of the pipeline.

What happens to edges when node IDs are normalized?

During the merge_and_normalize phase, the script builds an id_mapping dictionary that tracks every original ID and its normalized form. After all nodes are processed, the script iterates through all edges and rewrites their source and target fields using this mapping. This ensures that edges always point to the corrected, canonical node identifiers even after aggressive normalization of malformed prefixes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →