# How merge-batch-graphs.py Deduplicates Nodes and Normalizes IDs in Understand Anything

> Learn how merge-batch-graphs.py in Egonex AI Understand Anything deduplicates nodes and normalizes IDs using a last-wins strategy and ID canonicalization for cleaner graph data.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: internals
- Published: 2026-06-13

---

**The [`merge-batch-graphs.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/merge-batch-graphs.py) script in Egonex-AI/Understand-Anything canonicalizes malformed node identifiers using `normalize_node_id` and removes duplicates by keeping the last occurrence in a dictionary-based last-wins strategy.**

[`merge-batch-graphs.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/merge-batch-graphs.py) serves as the final aggregation phase of the `/understand` pipeline in the [Egonex-AI/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything) repository. It consumes the per-batch JSON files generated by file-analyzer agents, corrects identifier inconsistencies, and emits a single assembled knowledge graph. This article examines exactly how the script handles **ID normalization** and **node deduplication** based on the source implementation.

## ID Normalization Strategy

The script fixes malformed node IDs before deduplication to ensure canonical identifiers across all batches. In [`understand-anything-plugin/skills/understand/merge-batch-graphs.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/merge-batch-graphs.py), the `normalize_node_id` function (lines 777-822) applies a series of regex-based transformations to handle four specific corruption patterns.

### Handling Double Prefixes and Project Names

When agents emit nodes with duplicate type prefixes or project-specific prefixes, the function strips the redundant segments:

- **Double prefix** (e.g., `file:file:src/foo.ts`): The redundant `file:` is removed, leaving `file:src/foo.ts`.
- **Project-name prefix** (e.g., `my-project:file:src/foo.ts`): A regex matches the pattern `<project-name>:<valid-prefix>:...` and drops the leading project name, retaining only `file:src/foo.ts`.

### Legacy Prefix Migration and Missing Prefix Resolution

The normalizer addresses historical identifier formats and incomplete data:

- **Legacy `func:` prefix**: Converted to the current `function:` standard.
- **Missing prefix** (bare paths like [`src/main.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/main.py)): The script consults the node's `type` field against the `TYPE_TO_PREFIX` mapping. If the node represents a function or class but lacks a file path, the placeholder `__nofilepath__` is inserted to prevent silent collisions.

### Rewriting Edge References

After normalization, the script records any ID changes in an `id_mapping` dictionary (lines 996-1004). This mapping is then applied to all edge `source` and `target` fields within the `merge_and_normalize` function, ensuring graph connectivity remains intact after identifiers are corrected.

## Node Deduplication Mechanism

Once IDs are normalized, the script eliminates redundant nodes that share the same canonical identifier across different batches.

### The Last-Wins Dictionary Approach

The deduplication logic (lines 996-1004) uses a simple but deterministic dictionary accumulation:

```python
nodes_by_id: dict[str, dict] = {}
for node in nodes_with_ids:
    nid = node.get("id", "")
    if nid in nodes_by_id:
        duplicate_count += 1
    nodes_by_id[nid] = node  # Last occurrence wins

```

This approach guarantees that when multiple batches contain the same node (e.g., `file:src/index.ts`), the version from the last processed batch is retained.

### Duplicate Counting and Reporting

The script tracks the number of duplicates removed and includes this metric in the final console report. As shown in the deduplication block, `duplicate_count` increments whenever an existing key is overwritten, providing transparency about data reduction during the merge process.

## Running the Merge Process

Execute the script from the repository root to process intermediate batch files:

```bash
python understand-anything-plugin/skills/understand/merge-batch-graphs.py /path/to/project

```

The script reads `.understand-anything/intermediate/batch-*.json`, applies the normalization and deduplication logic, and writes the result to [`.understand-anything/intermediate/assembled-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/assembled-graph.json).

## Example Transformations

Consider these malformed input nodes:

```json
[
  {"id": "file:file:src/app.js", "type": "file"},
  {"id": "myproj:file:src/app.js", "type": "file"},
  {"id": "func:myFunction", "type": "function", "filePath": "src/app.js", "name": "myFunction"}
]

```

After passing through `normalize_node_id`, they become:

```json
[
  {"id": "file:src/app.js", "type": "file"},
  {"id": "file:src/app.js", "type": "file"},
  {"id": "function:file:src/app.js:myFunction", "type": "function"}
]

```

The deduplication step then keeps only the second `file:src/app.js` node, removing the earlier duplicate.

## Summary

- **`normalize_node_id`** (lines 777-822) fixes double prefixes, project prefixes, legacy `func:` labels, and missing prefixes using the `TYPE_TO_PREFIX` mapping.
- **`id_mapping`** tracks all ID corrections so edges can be rewritten to point to normalized targets.
- **Deduplication** (lines 996-1004) uses a dictionary with a last-wins policy to ensure unique nodes while counting removals for the final report.
- The script outputs [`assembled-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/assembled-graph.json) containing clean, canonical nodes consistent across all batches.

## Frequently Asked Questions

### What is the difference between ID normalization and node deduplication in merge-batch-graphs.py?

ID normalization corrects malformed identifier strings (e.g., removing duplicate prefixes like `file:file:`) to ensure every node has a canonical ID. Node deduplication occurs after normalization and removes multiple nodes that share the same canonical ID, keeping only the last occurrence. Normalization happens first because deduplication relies on consistent identifier strings to detect true duplicates.

### How does merge-batch-graphs.py handle conflicting node IDs from different batches?

The script processes batches sequentially and uses a dictionary-based accumulation where the last occurrence of a node ID overwrites previous entries. According to the implementation in lines 996-1004, when a duplicate ID is encountered, the `duplicate_count` increments, but the new node data replaces the old entry in `nodes_by_id`. This deterministic "last-wins" approach ensures predictable output regardless of batch ordering.

### Why does the script use a "last-wins" strategy for duplicate nodes?

The last-wins strategy in [`merge-batch-graphs.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/merge-batch-graphs.py) provides deterministic behavior without requiring complex merge logic for node attributes. Since the script cannot automatically reconcile conflicting properties between duplicate nodes (e.g., different metadata from two analyzers), it assumes later batches contain more accurate or complete information. This approach also matches the sequential processing model of the pipeline.

### What happens to edges when node IDs are normalized?

During the `merge_and_normalize` phase, the script builds an `id_mapping` dictionary that tracks every original ID and its normalized form. After all nodes are processed, the script iterates through all edges and rewrites their `source` and `target` fields using this mapping. This ensures that edges always point to the corrected, canonical node identifiers even after aggressive normalization of malformed prefixes.