# How the Store Module Ensures Data Integrity for the Graph in Codebase-Memory-MCP

> Learn how the store module safeguards your graph by ensuring data integrity with SQLite constraints, ACID transactions, and runtime safety hooks. Prevent corruption and recover safely.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: internals
- Published: 2026-07-05

---

**The store module ensures data integrity for the graph through a layered defense of SQLite schema constraints, foreign-key cascades, ACID transactions, and runtime safety hooks that prevent corruption during both normal operations and crash recovery.**

The `store` module in the [DeusData/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp) repository serves as a thin yet robust wrapper around SQLite, designed to maintain a consistent code knowledge graph even under adverse conditions. By combining strict database-level constraints with careful transaction management and defensive programming practices, the module guarantees that nodes, edges, and project metadata remain valid and referentially intact across the entire graph lifecycle.

## Schema-Level Constraints

The foundation of data integrity begins with the SQL schema defined in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c). The module enforces structural correctness at the database level, preventing invalid data from ever being persisted.

### Primary Keys and Auto-Increment

Every entity in the graph receives a unique integer identifier through `INTEGER PRIMARY KEY AUTOINCREMENT`. In [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) at lines 222-236, the schema creation statements for both `nodes` and `edges` tables guarantee that each row has a non-null, unique `id` that SQLite manages automatically. This prevents identity collisions and ensures that every node and edge can be referenced unambiguously.

### Unique Constraints for Nodes and Edges

Duplicate entries are blocked through composite unique constraints. Nodes are unique per combination of `project` and `qualified_name`, ensuring that a single project cannot contain two definitions with the same fully-qualified path. Edges enforce uniqueness through the composite key `(source_id, target_id, type, local_name_gen)` defined at lines 244-265 in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c), which prevents duplicate relationships while allowing distinct import types between the same source and target nodes.

### Foreign-Key Cascades

Referential integrity is maintained through explicit foreign-key relationships. The `project` column in both `nodes` and `edges` tables references `projects(name)`, while `source_id` and `target_id` reference `nodes(id)`. These constraints use `ON DELETE CASCADE` (lines 226-235 in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c)), ensuring that deleting a project automatically removes its associated nodes and edges, preventing orphaned records that would violate graph consistency.

### Generated Columns

Deterministic indexing is enforced through generated columns. The schema defines `url_path_gen` and `local_name_gen` as computed values derived from the JSON `properties` column (lines 260-267). This ensures that indexes remain synchronized with the stored JSON data, preventing query plan inconsistencies that could arise from manual index maintenance.

## Runtime Safety Mechanisms

Beyond static schema definitions, the store module implements active protections during database operation to guard against both malicious manipulation and environmental hazards.

### SQLite Authorizer Callbacks

The module registers an authorizer callback (lines 601-607 in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c)) that intercepts SQL statements before execution. This callback rejects dangerous operations such as `ATTACH` and `DETACH`, preventing SQL injection attacks that could attempt to write arbitrary files or access external databases through the graph store interface.

### PRAGMA Configuration

Connection-level safety is enforced through strict SQLite pragma settings. During initialization (lines 569-585), the module explicitly enables `foreign_keys = ON` to enforce referential constraints, sets `temp_store = MEMORY` to prevent temporary tables from spilling to disk, and configures appropriate `synchronous` and `journal_mode` values based on the operational mode (in-memory, read-only, or normal). These settings guarantee ACID compliance and prevent accidental promotion of read-only databases to writable states.

### WAL Checkpointing

Crash recovery is handled through careful Write-Ahead Logging (WAL) management. Before closing any database connection, the module executes a checkpoint (lines 999-1002) to ensure all WAL data is flushed to the main database file. This prevents orphaned WAL files that could leave the database in an inconsistent state after a system crash.

### Integrity Check Helper

The `cbm_store_check_integrity()` function (lines 822-864) implements lightweight sanity checks that validate the `projects` table and verify that `root_path` values resemble valid filesystem paths. This runtime validation can detect corruption that schema constraints alone cannot catch, such as partially written rows or malformed path strings.

## Transactional Guarantees

All graph modifications occur within explicit transaction boundaries, ensuring atomicity and consistency across complex multi-table operations.

### Explicit Transaction API

The module exports `cbm_store_begin()`, `cbm_store_commit()`, and `cbm_store_rollback()` functions (lines 560-570) that wrap SQLite's `BEGIN IMMEDIATE`, `COMMIT`, and `ROLLBACK` statements. Using `BEGIN IMMEDIATE` rather than `BEGIN DEFERRED` prevents write-lock starvation and ensures that transactions fail fast if the database is unavailable, rather than waiting indefinitely.

### Bulk Write Mode

For high-throughput scenarios, `cbm_store_begin_bulk()` (lines 723-734) temporarily reconfigures the database for performance while maintaining safety. It disables `synchronous` mode and increases `cache_size` while preserving WAL mode, allowing rapid insertion of many nodes or edges. The counterpart `cbm_store_end_bulk()` restores normal pragmas, ensuring that crash protection is never permanently disabled.

### Prepared Statement Reuse

All write operations utilize cached prepared statements that are reset and rebound for each use (lines 889-904). This pattern eliminates SQL parsing overhead while guaranteeing that parameter bindings are always correct, preventing data leakage between operations and ensuring that bound values match the expected schema types.

## Defensive Design Patterns

The store module incorporates additional safeguards against edge cases including read-only filesystems and memory-mapped I/O errors.

### Read-Only Query Mode

The `cbm_store_open_path_query()` function (lines 720-754) implements a two-tier opening strategy. It first attempts a standard `READONLY` open; if this fails (for example, on a read-only filesystem), it falls back to an immutable `file:` URI that bypasses WAL processing entirely. This ensures that query operations can proceed safely even when the underlying storage cannot accept writes, preventing accidental modification attempts.

### Deterministic SQLite Functions

Custom functions registered with SQLite—such as `regexp`, `iregexp`, `cbm_cosine_i8`, and `cbm_camel_split`—are explicitly marked with the `SQLITE_DETERMINISTIC` flag (lines 635-647). This declaration guarantees that identical inputs always produce identical outputs, preventing non-deterministic side effects from contaminating query plans or index selections.

### Memory-Mapped I/O Limits

The module respects the `CBM_SQLITE_MMAP_SIZE` environment variable (lines 404-416) to cap memory-mapped I/O, defaulting to 64 MiB. This prevents SIGBUS crashes that could occur if the database file is truncated while memory-mapped. Setting the variable to negative values disables memory mapping entirely, providing a safe fallback for unreliable storage environments.

## Practical Code Examples

### Example 1: Atomic Node Insertion

```c
/* Open (or create) a per-project database */
cbm_store_t *s = cbm_store_open("myproject");
if (!s) { /* handle error */ }

/* Start a transaction */
if (cbm_store_begin(s) != CBM_STORE_OK) { /* handle error */ }

/* Upsert a node */
cbm_node_t node = {
    .project        = "myproject",
    .label          = "Function",
    .name           = "handleRequest",
    .qualified_name = "pkg.Service.handleRequest",
    .file_path      = "src/service.go",
    .start_line     = 42,
    .end_line       = 78,
    .properties_json = "{\"visibility\":\"public\"}"
};
int64_t node_id = cbm_store_upsert_node(s, &node);
if (node_id < 0) { /* handle error */ }

/* Commit the transaction */
if (cbm_store_commit(s) != CBM_STORE_OK) { /* handle error */ }

cbm_store_close(s);

```

This example demonstrates transaction isolation and the `UNIQUE(project, qualified_name)` constraint preventing duplicate nodes.

### Example 2: Safe Bulk Edge Insertion

```c
cbm_store_t *s = cbm_store_open("myproject");
cbm_store_begin_bulk(s);               // disables sync, enlarges cache

for (int i = 0; i < edge_count; ++i) {
    cbm_edge_t e = {
        .project         = "myproject",
        .source_id       = src_ids[i],
        .target_id       = tgt_ids[i],
        .type            = "CALLS",
        .properties_json = NULL      // defaults to '{}'
    };
    int64_t edge_id = cbm_store_insert_edge(s, &e);
    if (edge_id < 0) { /* handle individual failure */ }
}

cbm_store_end_bulk(s);                 // restores normal pragmas
cbm_store_close(s);

```

Bulk mode accelerates the write path while maintaining WAL safety; the edge `UNIQUE` constraint prevents duplicate relationships.

### Example 3: Post-Crash Integrity Verification

```c
cbm_store_t *s = cbm_store_open_path_query("/var/cache/cbm/myproject.db");
if (!s) { /* DB missing – re-index */ }

if (!cbm_store_check_integrity(s)) {
    fprintf(stderr, "Corrupt DB detected – re-index required\n");
    /* Recommended: delete the DB and rebuild from source */
}
cbm_store_close(s);

```

The integrity helper validates row counts and `root_path` formatting, flagging corruption before the application proceeds.

## Summary

The store module ensures data integrity for the graph through multiple reinforcing mechanisms:

- **Schema constraints** including primary keys, unique composites, and foreign-key cascades prevent invalid data at the database level
- **Runtime protections** such as the SQLite authorizer and pragma configurations block dangerous operations and enforce ACID compliance
- **Transaction management** via explicit `cbm_store_begin()` and `cbm_store_commit()` calls ensures atomic updates across nodes and edges
- **Bulk operations** maintain performance without sacrificing crash safety through temporary pragma adjustments
- **Integrity checks** and read-only fallback modes provide safe recovery paths when storage conditions degrade

## Frequently Asked Questions

### How does the store module prevent duplicate nodes in the graph?

The schema defines a unique constraint on the combination of `project` and `qualified_name` columns in the `nodes` table (lines 244-265 in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c)). When `cbm_store_upsert_node()` executes an `INSERT OR REPLACE` operation, SQLite enforces this constraint automatically, preventing the same qualified name from appearing twice within a single project while allowing updates to existing node properties.

### What happens if a write operation fails midway through a transaction?

Because all writes occur within explicit transactions managed by `cbm_store_begin()`, any failure before `cbm_store_commit()` executes leaves the database in its pre-transaction state. The module uses `BEGIN IMMEDIATE` semantics, which acquires the write lock immediately; if the application crashes or calls `cbm_store_rollback()`, SQLite automatically discards the pending changes, ensuring that partial graph updates never persist.

### How does the module detect database corruption?

The `cbm_store_check_integrity()` function (lines 822-864) performs lightweight validation by checking that the `projects` table contains reasonable row counts and that `root_path` values conform to expected filesystem path patterns. While this does not replace SQLite's internal `PRAGMA integrity_check`, it provides a fast application-level sanity check that can identify obvious corruption before expensive validation runs.

### Can the store operate safely on read-only filesystems?

Yes. The `cbm_store_open_path_query()` function implements a fallback mechanism (lines 720-754) that detects read-only filesystems and opens the database using an immutable `file:` URI with the `immutable=1` parameter. This mode bypasses WAL processing and write attempts entirely, allowing safe read-only access to graph data without risking I/O errors or corruption warnings.