# How Graphify Detects and Merges Ghost Duplicate Nodes: A Technical Deep Dive

> Learn how Graphify detects and merges ghost duplicate nodes using a three-pass algorithm that remaps edges and eliminates stale references for cleaner graph data.

- Repository: [Safi/graphify](https://github.com/safishamsi/graphify)
- Tags: technical-deep-dive
- Published: 2026-06-15

---

**Graphify detects ghost duplicate nodes by identifying stale cached symbols from deleted files and non-AST duplicates, then merges them into canonical AST nodes using a three-pass algorithm in [`build.py`](https://github.com/safishamsi/graphify/blob/main/build.py) that remaps edges and eliminates stale references.**

In the `safishamsi/graphify` repository, **ghost nodes** represent stale or duplicate symbols that appear during graph extraction but should be consolidated into existing canonical AST nodes. The system handles these duplicates through a two-stage pipeline that first detects obsolete entries during incremental builds, then performs a sophisticated three-pass merge operation to preserve graph integrity.

## Detection of Ghost Nodes in Incremental Extraction

The detection process begins in [`graphify/detect.py`](https://github.com/safishamsi/graphify/blob/main/graphify/detect.py) during incremental extraction workflows. When rebuilding a graph, the system compares the current repository state against the manifest of previously indexed files. Any file that no longer exists in the repository but still has cached nodes in the graph represents a ghost node that requires cleanup.

The implementation uses set difference logic to identify deleted files:

```python

# graphify/detect.py – identify deleted files and their cached nodes

# (lines 15-24)

# --------------------------------------------------------------

# Files in manifest that no longer exist - their cached nodes are now ghost nodes

current_files = {f for flist in full["files"].values() for f in flist}
deleted_files = [f for f in manifest if f not in current_files]

# --------------------------------------------------------------

```

The resulting `deleted_files` list is returned to the build step, where the corresponding nodes are flagged for merging or removal. This ensures that **stale cached data** from file deletions does not persist as orphaned nodes in the knowledge graph.

## Merging Ghost Nodes into Canonical AST Nodes

The heavy lifting of ghost node resolution occurs in [`graphify/build.py`](https://github.com/safishamsi/graphify/blob/main/graphify/build.py) through a three-pass algorithm that distinguishes between canonical AST nodes and duplicate ghosts. This process ensures that LLM-generated symbols or stale cached entries merge safely into their AST counterparts.

### Pass 1 – Gather Canonical AST Nodes

The first pass collects all AST-origin nodes into a lookup map keyed by `(basename, label)`. Nodes originating from the AST are identified by the `_origin` attribute set to `"ast"` or by the presence of a `source_location` field. The algorithm tracks collisions when multiple AST nodes share the same key to prevent unsafe merges.

```python

# graphify/build.py – Pass 1 (lines 71-95)

# --------------------------------------------------------------

for nid in node_set:
    attrs = G.nodes[nid]
    label = str(attrs.get("label", "")).strip()
    sf = str(attrs.get("source_file", ""))
    basename = Path(sf).name if sf else ""
    if not label or not basename:
        continue
    is_ast = attrs.get("_origin") == "ast"
    if attrs.get("source_location") or is_ast:
        key = (basename, label)
        if is_ast:
            if key in _loc_nodes and G.nodes[_loc_nodes[key]].get("_origin") == "ast":
                _loc_collisions.add(key)
            _loc_nodes[key] = nid          # AST wins over any non-AST entry

        elif key not in _loc_nodes:
            _loc_nodes[key] = nid

# --------------------------------------------------------------

```

The `_loc_nodes` dictionary stores the canonical node IDs, while `_loc_collisions` tracks ambiguous keys where multiple AST nodes exist for the same symbol.

### Pass 2 – Identify Ghost Nodes

The second pass identifies ghost candidates by examining non-AST nodes that share keys with canonical AST entries. Any node that is not AST-origin but matches a `(basename, label)` key in `_loc_nodes` becomes a ghost, provided the key does not appear in the collisions set.

```python

# graphify/build.py – Pass 2 (lines 96-110)

# --------------------------------------------------------------

for nid in node_set:
    attrs = G.nodes[nid]
    if attrs.get("_origin") == "ast":
        continue                       # AST nodes are never ghosts

    label = str(attrs.get("label", "")).strip()
    sf = str(attrs.get("source_file", ""))
    basename = Path(sf).name if sf else ""
    if not label or not basename:
        continue
    key = (basename, label)
    if key in _loc_collisions:
        continue                       # ambiguous – keep ghost

    if key in _loc_nodes and _loc_nodes[key] != nid:
        _noloc_nodes[key] = nid          # record as ghost

# --------------------------------------------------------------

```

The `_noloc_nodes` dictionary accumulates these ghost mappings, waiting for the final consolidation pass.

### Pass 3 – Remap and Delete Ghosts

The final pass creates remapping entries from ghost IDs to canonical AST IDs, removes the ghost nodes from the graph, and updates a normalized ID map to preserve edge connectivity. This ensures that any relationships originally pointing to ghosts are redirected to their canonical counterparts.

```python

# graphify/build.py – Remap & delete (lines 111-229)

# --------------------------------------------------------------

_ghost_remap: dict[str, str] = {}
for key, sem_id in _noloc_nodes.items():
    ast_id = _loc_nodes.get(key)
    if ast_id is not None:
        _ghost_remap[sem_id] = ast_id

for ghost_id in _ghost_remap:
    G.remove_node(ghost_id)                 # drop ghost

    node_set.discard(ghost_id)

# Normalised-ID map – make edges survive the merge

norm_to_id: dict[str, str] = {_normalize_id(nid): nid for nid in node_set}
for ghost_id, canonical_id in _ghost_remap.items():
    norm_to_id[_normalize_id(ghost_id)] = canonical_id
    norm_to_id[ghost_id] = canonical_id

# --------------------------------------------------------------

```

After this remap, the edge-construction loop resolves all source and target references through `norm_to_id`, guaranteeing that edges originally attached to ghost nodes now point to the canonical AST node.

## Practical Example: Building a Deduplicated Graph

When combining extractions from an AST parser and an LLM semantic extractor, ghost nodes are automatically resolved during the build process:

```python
from pathlib import Path
from graphify import build

# 1️⃣  Extracts from an AST extractor (canonical) and an LLM extractor (semantic)

ast_extraction   = {... "nodes": [...], "edges": [...]}      # contains _origin="ast"

semantic_extraction = {... "nodes": [...], "edges": [...]}   # may contain duplicates

# 2️⃣  Build the merged graph – ghosts are automatically merged/deduped

G = build([ast_extraction, semantic_extraction])

# 3️⃣  Verify removal of a ghost node (example ID "render_ghost")

assert "render_ghost" not in G.nodes

# The node “render” (the AST canonical) now has all edges that previously pointed

# to “render_ghost”.

```

This workflow demonstrates how Graphify maintains a **clean, canonical graph** even when consuming multiple extraction sources that may contain overlapping symbols.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`graphify/build.py`](https://github.com/safishamsi/graphify/blob/main/graphify/build.py) | Implements the three-pass ghost detection and merge logic |
| [`graphify/detect.py`](https://github.com/safishamsi/graphify/blob/main/graphify/detect.py) | Marks nodes from deleted files as ghosts for incremental builds |
| [`graphify/extract.py`](https://github.com/safishamsi/graphify/blob/main/graphify/extract.py) | Produces the raw node/edge JSON that feeds the ghost-merge step |
| [`docs/how-it-works.md`](https://github.com/safishamsi/graphify/blob/main/docs/how-it-works.md) | Provides high-level background on the extraction pipeline |

## Summary

- **Ghost nodes** are stale or duplicate symbols that appear in extracted graphs but should merge into canonical AST nodes.
- **Detection** occurs in [`graphify/detect.py`](https://github.com/safishamsi/graphify/blob/main/graphify/detect.py) by identifying files removed from the repository that still have cached node entries.
- **Merging** happens in [`graphify/build.py`](https://github.com/safishamsi/graphify/blob/main/graphify/build.py) through a three-pass algorithm: gathering AST nodes, identifying ghost candidates, and remapping edges.
- The **(basename, label)** key uniquely identifies symbols, with AST-origin nodes taking precedence over semantic or cached duplicates.
- **Edge preservation** is maintained through a normalized ID map that redirects references from deleted ghosts to their canonical counterparts.

## Frequently Asked Questions

### What causes ghost nodes to appear in Graphify?

Ghost nodes appear when cached graph data references files that have been deleted from the repository, or when semantic extractors (such as LLM-based tools) generate symbols that duplicate existing AST nodes. According to the source code in [`graphify/detect.py`](https://github.com/safishamsi/graphify/blob/main/graphify/detect.py), these stale entries persist in the graph manifest until the incremental extraction process identifies them for cleanup.

### How does Graphify decide which node is the canonical source?

Graphify prioritizes **AST-origin nodes** over all other sources. In [`graphify/build.py`](https://github.com/safishamsi/graphify/blob/main/graphify/build.py) Pass 1, nodes with `_origin == "ast"` or a `source_location` attribute are stored in `_loc_nodes`, and when collisions occur, the AST node always wins over non-AST entries. This ensures that statically analyzed source code serves as the ground truth for symbol identity.

### What happens to edges connected to ghost nodes during the merge?

All edges originally connected to ghost nodes are **remapped to the canonical AST node**. During Pass 3 in [`build.py`](https://github.com/safishamsi/graphify/blob/main/build.py), the system populates a `_ghost_remap` dictionary mapping ghost IDs to canonical IDs, then updates the `norm_to_id` lookup table so that subsequent edge construction automatically resolves ghost references to their canonical targets.

### Can the ghost detection handle ambiguous symbols with multiple AST definitions?

Yes. The algorithm tracks ambiguous keys in the `_loc_collisions` set during Pass 1. If multiple AST nodes share the same `(basename, label)` key, that key is excluded from ghost merging in Pass 2. This safety mechanism prevents the system from incorrectly merging distinct AST symbols that happen to share a name, keeping the graph semantically accurate.