# How Nodes Are Uniquely Identified in the code-review-graph Database

> Learn how nodes are uniquely identified in the code-review-graph database using the qualified_name column. Discover the combination of file path, parent, and identity names enforced by SQLite.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: internals
- Published: 2026-08-11

---

**Nodes in the code-review-graph database are uniquely identified by the `qualified_name` column, which combines a POSIX-normalized file path with an optional parent name and identity name, enforced via a SQLite `UNIQUE` constraint.**

The **code-review-graph** project transforms codebase elements—files, classes, functions, methods, and tests—into a queryable knowledge graph backed by SQLite. Understanding how nodes achieve unique identity is essential for anyone building tools on top of this graph or debugging duplicate node issues. The system guarantees exactly one canonical record per logical symbol through a deterministic naming scheme implemented in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py).

## The qualified_name Column and UNIQUE Constraint

The foundation of node uniqueness lies in the `nodes` table schema defined in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py). The `qualified_name` column carries a `UNIQUE` constraint that prevents duplicate insertions at the database level.

```python

# From code_review_graph/graph.py lines 75-80 (approximate schema section)

CREATE TABLE IF NOT EXISTS nodes (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    qualified_name TEXT NOT NULL UNIQUE,  -- ← uniqueness enforced here
    kind TEXT NOT NULL,
    name TEXT NOT NULL,
    file_path TEXT NOT NULL,
    ...
)

```

This constraint ensures that any attempt to insert a node with an existing `qualified_name` triggers the `ON CONFLICT(qualified_name) DO UPDATE` logic, converting potential duplicates into upsert operations.

## How qualified_name Is Constructed

The private helper function `_make_qualified` in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) builds the `qualified_name` string whenever a node is created or updated (lines 2159–2165). The algorithm follows a consistent pattern based on node type:

### File Nodes

For nodes representing files themselves, the `qualified_name` equals the normalized file path:

```

qualified_name = <file_path>

```

### Non-File Nodes

For classes, functions, methods, and other symbols, the format incorporates hierarchy:

```

qualified_name = <file_path>::[<parent_name>.]<identity_name>

```

The components work as follows:

- **`<file_path>`** – POSIX-normalized path relative to the repository root (via `normalize_file_path`)
- **`[<parent_name>.]`** – Optional enclosing scope (e.g., class name for methods)
- **`<identity_name>`** – Falls back to the node's `name` when not explicitly provided

## Upsert Behavior and Canonical Records

The `upsert_node` method in `GraphStore` leverages this design to maintain idempotency. Calling `upsert_node` with identical `qualified_name` components will always return the same node ID, regardless of how many times it is invoked.

```python
from code_review_graph.graph import GraphStore, NodeInfo

# Example: a top-level function in src/util.py

node = NodeInfo(
    kind="Function",
    name="do_work",
    file_path="src/util.py",               # stored POSIX path

    line_start=10,
    line_end=15,
    language="python",
    parent_name=None,                      # no enclosing class

    identity_name=None,                    # falls back to name

    is_test=False,
    extra={}
)

with GraphStore("graph.db") as store:
    node_id = store.upsert_node(node)      # → inserts or updates

    # The generated qualified_name is:

    # "src/util.py::do_work"

    retrieved = store.get_node("src/util.py::do_work")
    assert retrieved.id == node_id

```

The `qualified_name` generation is deterministic: the same source symbol will always produce the same identifier, enabling reliable lookups and cross-referencing.

## Handling Nested Symbols with Parent Names

When symbols exist within enclosing scopes—such as methods inside classes—the `parent_name` field captures this hierarchy in the identifier.

```python

# Example: a method inside a class

node = NodeInfo(
    kind="Method",
    name="run",
    file_path="src/app.py",
    line_start=45,
    line_end=50,
    language="python",
    parent_name="App",                     # enclosing class

    identity_name=None,
    is_test=False,
    extra={}
)

with GraphStore("graph.db") as store:
    qual = store.upsert_node(node)        # qualified_name = "src/app.py::App.run"

    print(store.get_node(qual).qualified_name)

```

This hierarchy encoding prevents collisions between methods with identical names in different classes within the same file.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) | Defines the SQLite schema with `qualified_name TEXT NOT NULL UNIQUE` and implements `_make_qualified` for identifier construction |
| [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py) | Provides `NodeInfo`, the dataclass used when creating or updating nodes |
| [`code_review_graph/constants.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/constants.py) | Supplies `normalize_file_path` for POSIX path normalization |

## Summary

- **Unique identification** relies on the `qualified_name` column with a `UNIQUE` constraint in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py)
- **Construction algorithm** in `_make_qualified` combines normalized file paths with optional parent and identity names
- **File nodes** use bare paths; **non-file nodes** use the `::` separator with hierarchical naming
- **Upsert semantics** guarantee canonical records through `ON CONFLICT(qualified_name) DO UPDATE`
- **Deterministic identifiers** enable reliable node lookup and graph traversal across analysis runs

## Frequently Asked Questions

### What happens if two nodes have the same qualified_name?

SQLite's `UNIQUE` constraint prevents insertion of duplicate `qualified_name` values. The `upsert_node` method catches this conflict and executes `DO UPDATE`, merging new information into the existing record while preserving the original node ID.

### Can qualified_name collisions occur across different languages?

No. The `qualified_name` includes the file path, and files from different languages typically reside in separate directories or use distinct extensions. Even if two languages define identical symbol names in similarly-named files, the full path—including extension—differentiates them.

### How does parent_name differ from identity_name?

`parent_name` captures the enclosing scope (e.g., the class containing a method), while `identity_name` serves as an override for the node's canonical name. When `identity_name` is `None`, the system falls back to the `name` field. Both feed into the hierarchical `qualified_name` construction: `parent_name` appears before the dot separator, and `identity_name` (or `name`) appears as the final segment.

### Where is the path normalization performed?

Path normalization occurs through `normalize_file_path`, defined in [`code_review_graph/constants.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/constants.py). This utility ensures Windows paths convert to POSIX format and eliminates redundant path components before the path enters the `qualified_name`, maintaining cross-platform consistency in node identification.