How Graphify Detects and Merges Ghost Duplicate Nodes: A Technical Deep Dive
Graphify detects ghost duplicate nodes by identifying stale cached symbols from deleted files and non-AST duplicates, then merges them into canonical AST nodes using a three-pass algorithm in build.py that remaps edges and eliminates stale references.
In the safishamsi/graphify repository, ghost nodes represent stale or duplicate symbols that appear during graph extraction but should be consolidated into existing canonical AST nodes. The system handles these duplicates through a two-stage pipeline that first detects obsolete entries during incremental builds, then performs a sophisticated three-pass merge operation to preserve graph integrity.
Detection of Ghost Nodes in Incremental Extraction
The detection process begins in graphify/detect.py during incremental extraction workflows. When rebuilding a graph, the system compares the current repository state against the manifest of previously indexed files. Any file that no longer exists in the repository but still has cached nodes in the graph represents a ghost node that requires cleanup.
The implementation uses set difference logic to identify deleted files:
# graphify/detect.py – identify deleted files and their cached nodes
# (lines 15-24)
# --------------------------------------------------------------
# Files in manifest that no longer exist - their cached nodes are now ghost nodes
current_files = {f for flist in full["files"].values() for f in flist}
deleted_files = [f for f in manifest if f not in current_files]
# --------------------------------------------------------------
The resulting deleted_files list is returned to the build step, where the corresponding nodes are flagged for merging or removal. This ensures that stale cached data from file deletions does not persist as orphaned nodes in the knowledge graph.
Merging Ghost Nodes into Canonical AST Nodes
The heavy lifting of ghost node resolution occurs in graphify/build.py through a three-pass algorithm that distinguishes between canonical AST nodes and duplicate ghosts. This process ensures that LLM-generated symbols or stale cached entries merge safely into their AST counterparts.
Pass 1 – Gather Canonical AST Nodes
The first pass collects all AST-origin nodes into a lookup map keyed by (basename, label). Nodes originating from the AST are identified by the _origin attribute set to "ast" or by the presence of a source_location field. The algorithm tracks collisions when multiple AST nodes share the same key to prevent unsafe merges.
# graphify/build.py – Pass 1 (lines 71-95)
# --------------------------------------------------------------
for nid in node_set:
attrs = G.nodes[nid]
label = str(attrs.get("label", "")).strip()
sf = str(attrs.get("source_file", ""))
basename = Path(sf).name if sf else ""
if not label or not basename:
continue
is_ast = attrs.get("_origin") == "ast"
if attrs.get("source_location") or is_ast:
key = (basename, label)
if is_ast:
if key in _loc_nodes and G.nodes[_loc_nodes[key]].get("_origin") == "ast":
_loc_collisions.add(key)
_loc_nodes[key] = nid # AST wins over any non-AST entry
elif key not in _loc_nodes:
_loc_nodes[key] = nid
# --------------------------------------------------------------
The _loc_nodes dictionary stores the canonical node IDs, while _loc_collisions tracks ambiguous keys where multiple AST nodes exist for the same symbol.
Pass 2 – Identify Ghost Nodes
The second pass identifies ghost candidates by examining non-AST nodes that share keys with canonical AST entries. Any node that is not AST-origin but matches a (basename, label) key in _loc_nodes becomes a ghost, provided the key does not appear in the collisions set.
# graphify/build.py – Pass 2 (lines 96-110)
# --------------------------------------------------------------
for nid in node_set:
attrs = G.nodes[nid]
if attrs.get("_origin") == "ast":
continue # AST nodes are never ghosts
label = str(attrs.get("label", "")).strip()
sf = str(attrs.get("source_file", ""))
basename = Path(sf).name if sf else ""
if not label or not basename:
continue
key = (basename, label)
if key in _loc_collisions:
continue # ambiguous – keep ghost
if key in _loc_nodes and _loc_nodes[key] != nid:
_noloc_nodes[key] = nid # record as ghost
# --------------------------------------------------------------
The _noloc_nodes dictionary accumulates these ghost mappings, waiting for the final consolidation pass.
Pass 3 – Remap and Delete Ghosts
The final pass creates remapping entries from ghost IDs to canonical AST IDs, removes the ghost nodes from the graph, and updates a normalized ID map to preserve edge connectivity. This ensures that any relationships originally pointing to ghosts are redirected to their canonical counterparts.
# graphify/build.py – Remap & delete (lines 111-229)
# --------------------------------------------------------------
_ghost_remap: dict[str, str] = {}
for key, sem_id in _noloc_nodes.items():
ast_id = _loc_nodes.get(key)
if ast_id is not None:
_ghost_remap[sem_id] = ast_id
for ghost_id in _ghost_remap:
G.remove_node(ghost_id) # drop ghost
node_set.discard(ghost_id)
# Normalised-ID map – make edges survive the merge
norm_to_id: dict[str, str] = {_normalize_id(nid): nid for nid in node_set}
for ghost_id, canonical_id in _ghost_remap.items():
norm_to_id[_normalize_id(ghost_id)] = canonical_id
norm_to_id[ghost_id] = canonical_id
# --------------------------------------------------------------
After this remap, the edge-construction loop resolves all source and target references through norm_to_id, guaranteeing that edges originally attached to ghost nodes now point to the canonical AST node.
Practical Example: Building a Deduplicated Graph
When combining extractions from an AST parser and an LLM semantic extractor, ghost nodes are automatically resolved during the build process:
from pathlib import Path
from graphify import build
# 1️⃣ Extracts from an AST extractor (canonical) and an LLM extractor (semantic)
ast_extraction = {... "nodes": [...], "edges": [...]} # contains _origin="ast"
semantic_extraction = {... "nodes": [...], "edges": [...]} # may contain duplicates
# 2️⃣ Build the merged graph – ghosts are automatically merged/deduped
G = build([ast_extraction, semantic_extraction])
# 3️⃣ Verify removal of a ghost node (example ID "render_ghost")
assert "render_ghost" not in G.nodes
# The node “render” (the AST canonical) now has all edges that previously pointed
# to “render_ghost”.
This workflow demonstrates how Graphify maintains a clean, canonical graph even when consuming multiple extraction sources that may contain overlapping symbols.
Key Implementation Files
| File | Purpose |
|---|---|
graphify/build.py |
Implements the three-pass ghost detection and merge logic |
graphify/detect.py |
Marks nodes from deleted files as ghosts for incremental builds |
graphify/extract.py |
Produces the raw node/edge JSON that feeds the ghost-merge step |
docs/how-it-works.md |
Provides high-level background on the extraction pipeline |
Summary
- Ghost nodes are stale or duplicate symbols that appear in extracted graphs but should merge into canonical AST nodes.
- Detection occurs in
graphify/detect.pyby identifying files removed from the repository that still have cached node entries. - Merging happens in
graphify/build.pythrough a three-pass algorithm: gathering AST nodes, identifying ghost candidates, and remapping edges. - The (basename, label) key uniquely identifies symbols, with AST-origin nodes taking precedence over semantic or cached duplicates.
- Edge preservation is maintained through a normalized ID map that redirects references from deleted ghosts to their canonical counterparts.
Frequently Asked Questions
What causes ghost nodes to appear in Graphify?
Ghost nodes appear when cached graph data references files that have been deleted from the repository, or when semantic extractors (such as LLM-based tools) generate symbols that duplicate existing AST nodes. According to the source code in graphify/detect.py, these stale entries persist in the graph manifest until the incremental extraction process identifies them for cleanup.
How does Graphify decide which node is the canonical source?
Graphify prioritizes AST-origin nodes over all other sources. In graphify/build.py Pass 1, nodes with _origin == "ast" or a source_location attribute are stored in _loc_nodes, and when collisions occur, the AST node always wins over non-AST entries. This ensures that statically analyzed source code serves as the ground truth for symbol identity.
What happens to edges connected to ghost nodes during the merge?
All edges originally connected to ghost nodes are remapped to the canonical AST node. During Pass 3 in build.py, the system populates a _ghost_remap dictionary mapping ghost IDs to canonical IDs, then updates the norm_to_id lookup table so that subsequent edge construction automatically resolves ghost references to their canonical targets.
Can the ghost detection handle ambiguous symbols with multiple AST definitions?
Yes. The algorithm tracks ambiguous keys in the _loc_collisions set during Pass 1. If multiple AST nodes share the same (basename, label) key, that key is excluded from ghost merging in Pass 2. This safety mechanism prevents the system from incorrectly merging distinct AST symbols that happen to share a name, keeping the graph semantically accurate.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →