How Nodes Are Uniquely Identified in the code-review-graph Database

Nodes in the code-review-graph database are uniquely identified by the qualified_name column, which combines a POSIX-normalized file path with an optional parent name and identity name, enforced via a SQLite UNIQUE constraint.

The code-review-graph project transforms codebase elements—files, classes, functions, methods, and tests—into a queryable knowledge graph backed by SQLite. Understanding how nodes achieve unique identity is essential for anyone building tools on top of this graph or debugging duplicate node issues. The system guarantees exactly one canonical record per logical symbol through a deterministic naming scheme implemented in code_review_graph/graph.py.

The qualified_name Column and UNIQUE Constraint

The foundation of node uniqueness lies in the nodes table schema defined in code_review_graph/graph.py. The qualified_name column carries a UNIQUE constraint that prevents duplicate insertions at the database level.


# From code_review_graph/graph.py lines 75-80 (approximate schema section)

CREATE TABLE IF NOT EXISTS nodes (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    qualified_name TEXT NOT NULL UNIQUE,  -- ← uniqueness enforced here
    kind TEXT NOT NULL,
    name TEXT NOT NULL,
    file_path TEXT NOT NULL,
    ...
)

This constraint ensures that any attempt to insert a node with an existing qualified_name triggers the ON CONFLICT(qualified_name) DO UPDATE logic, converting potential duplicates into upsert operations.

How qualified_name Is Constructed

The private helper function _make_qualified in code_review_graph/graph.py builds the qualified_name string whenever a node is created or updated (lines 2159–2165). The algorithm follows a consistent pattern based on node type:

File Nodes

For nodes representing files themselves, the qualified_name equals the normalized file path:


qualified_name = <file_path>

Non-File Nodes

For classes, functions, methods, and other symbols, the format incorporates hierarchy:


qualified_name = <file_path>::[<parent_name>.]<identity_name>

The components work as follows:

  • <file_path> – POSIX-normalized path relative to the repository root (via normalize_file_path)
  • [<parent_name>.] – Optional enclosing scope (e.g., class name for methods)
  • <identity_name> – Falls back to the node's name when not explicitly provided

Upsert Behavior and Canonical Records

The upsert_node method in GraphStore leverages this design to maintain idempotency. Calling upsert_node with identical qualified_name components will always return the same node ID, regardless of how many times it is invoked.

from code_review_graph.graph import GraphStore, NodeInfo

# Example: a top-level function in src/util.py

node = NodeInfo(
    kind="Function",
    name="do_work",
    file_path="src/util.py",               # stored POSIX path

    line_start=10,
    line_end=15,
    language="python",
    parent_name=None,                      # no enclosing class

    identity_name=None,                    # falls back to name

    is_test=False,
    extra={}
)

with GraphStore("graph.db") as store:
    node_id = store.upsert_node(node)      # → inserts or updates

    # The generated qualified_name is:

    # "src/util.py::do_work"

    retrieved = store.get_node("src/util.py::do_work")
    assert retrieved.id == node_id

The qualified_name generation is deterministic: the same source symbol will always produce the same identifier, enabling reliable lookups and cross-referencing.

Handling Nested Symbols with Parent Names

When symbols exist within enclosing scopes—such as methods inside classes—the parent_name field captures this hierarchy in the identifier.


# Example: a method inside a class

node = NodeInfo(
    kind="Method",
    name="run",
    file_path="src/app.py",
    line_start=45,
    line_end=50,
    language="python",
    parent_name="App",                     # enclosing class

    identity_name=None,
    is_test=False,
    extra={}
)

with GraphStore("graph.db") as store:
    qual = store.upsert_node(node)        # qualified_name = "src/app.py::App.run"

    print(store.get_node(qual).qualified_name)

This hierarchy encoding prevents collisions between methods with identical names in different classes within the same file.

Key Implementation Files

File Purpose
code_review_graph/graph.py Defines the SQLite schema with qualified_name TEXT NOT NULL UNIQUE and implements _make_qualified for identifier construction
code_review_graph/parser.py Provides NodeInfo, the dataclass used when creating or updating nodes
code_review_graph/constants.py Supplies normalize_file_path for POSIX path normalization

Summary

  • Unique identification relies on the qualified_name column with a UNIQUE constraint in code_review_graph/graph.py
  • Construction algorithm in _make_qualified combines normalized file paths with optional parent and identity names
  • File nodes use bare paths; non-file nodes use the :: separator with hierarchical naming
  • Upsert semantics guarantee canonical records through ON CONFLICT(qualified_name) DO UPDATE
  • Deterministic identifiers enable reliable node lookup and graph traversal across analysis runs

Frequently Asked Questions

What happens if two nodes have the same qualified_name?

SQLite's UNIQUE constraint prevents insertion of duplicate qualified_name values. The upsert_node method catches this conflict and executes DO UPDATE, merging new information into the existing record while preserving the original node ID.

Can qualified_name collisions occur across different languages?

No. The qualified_name includes the file path, and files from different languages typically reside in separate directories or use distinct extensions. Even if two languages define identical symbol names in similarly-named files, the full path—including extension—differentiates them.

How does parent_name differ from identity_name?

parent_name captures the enclosing scope (e.g., the class containing a method), while identity_name serves as an override for the node's canonical name. When identity_name is None, the system falls back to the name field. Both feed into the hierarchical qualified_name construction: parent_name appears before the dot separator, and identity_name (or name) appears as the final segment.

Where is the path normalization performed?

Path normalization occurs through normalize_file_path, defined in code_review_graph/constants.py. This utility ensures Windows paths convert to POSIX format and eliminates redundant path components before the path enters the qualified_name, maintaining cross-platform consistency in node identification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →