How the Store Module Ensures Data Integrity for the Graph in Codebase-Memory-MCP
The store module ensures data integrity for the graph through a layered defense of SQLite schema constraints, foreign-key cascades, ACID transactions, and runtime safety hooks that prevent corruption during both normal operations and crash recovery.
The store module in the DeusData/codebase-memory-mcp repository serves as a thin yet robust wrapper around SQLite, designed to maintain a consistent code knowledge graph even under adverse conditions. By combining strict database-level constraints with careful transaction management and defensive programming practices, the module guarantees that nodes, edges, and project metadata remain valid and referentially intact across the entire graph lifecycle.
Schema-Level Constraints
The foundation of data integrity begins with the SQL schema defined in src/store/store.c. The module enforces structural correctness at the database level, preventing invalid data from ever being persisted.
Primary Keys and Auto-Increment
Every entity in the graph receives a unique integer identifier through INTEGER PRIMARY KEY AUTOINCREMENT. In src/store/store.c at lines 222-236, the schema creation statements for both nodes and edges tables guarantee that each row has a non-null, unique id that SQLite manages automatically. This prevents identity collisions and ensures that every node and edge can be referenced unambiguously.
Unique Constraints for Nodes and Edges
Duplicate entries are blocked through composite unique constraints. Nodes are unique per combination of project and qualified_name, ensuring that a single project cannot contain two definitions with the same fully-qualified path. Edges enforce uniqueness through the composite key (source_id, target_id, type, local_name_gen) defined at lines 244-265 in src/store/store.c, which prevents duplicate relationships while allowing distinct import types between the same source and target nodes.
Foreign-Key Cascades
Referential integrity is maintained through explicit foreign-key relationships. The project column in both nodes and edges tables references projects(name), while source_id and target_id reference nodes(id). These constraints use ON DELETE CASCADE (lines 226-235 in src/store/store.c), ensuring that deleting a project automatically removes its associated nodes and edges, preventing orphaned records that would violate graph consistency.
Generated Columns
Deterministic indexing is enforced through generated columns. The schema defines url_path_gen and local_name_gen as computed values derived from the JSON properties column (lines 260-267). This ensures that indexes remain synchronized with the stored JSON data, preventing query plan inconsistencies that could arise from manual index maintenance.
Runtime Safety Mechanisms
Beyond static schema definitions, the store module implements active protections during database operation to guard against both malicious manipulation and environmental hazards.
SQLite Authorizer Callbacks
The module registers an authorizer callback (lines 601-607 in src/store/store.c) that intercepts SQL statements before execution. This callback rejects dangerous operations such as ATTACH and DETACH, preventing SQL injection attacks that could attempt to write arbitrary files or access external databases through the graph store interface.
PRAGMA Configuration
Connection-level safety is enforced through strict SQLite pragma settings. During initialization (lines 569-585), the module explicitly enables foreign_keys = ON to enforce referential constraints, sets temp_store = MEMORY to prevent temporary tables from spilling to disk, and configures appropriate synchronous and journal_mode values based on the operational mode (in-memory, read-only, or normal). These settings guarantee ACID compliance and prevent accidental promotion of read-only databases to writable states.
WAL Checkpointing
Crash recovery is handled through careful Write-Ahead Logging (WAL) management. Before closing any database connection, the module executes a checkpoint (lines 999-1002) to ensure all WAL data is flushed to the main database file. This prevents orphaned WAL files that could leave the database in an inconsistent state after a system crash.
Integrity Check Helper
The cbm_store_check_integrity() function (lines 822-864) implements lightweight sanity checks that validate the projects table and verify that root_path values resemble valid filesystem paths. This runtime validation can detect corruption that schema constraints alone cannot catch, such as partially written rows or malformed path strings.
Transactional Guarantees
All graph modifications occur within explicit transaction boundaries, ensuring atomicity and consistency across complex multi-table operations.
Explicit Transaction API
The module exports cbm_store_begin(), cbm_store_commit(), and cbm_store_rollback() functions (lines 560-570) that wrap SQLite's BEGIN IMMEDIATE, COMMIT, and ROLLBACK statements. Using BEGIN IMMEDIATE rather than BEGIN DEFERRED prevents write-lock starvation and ensures that transactions fail fast if the database is unavailable, rather than waiting indefinitely.
Bulk Write Mode
For high-throughput scenarios, cbm_store_begin_bulk() (lines 723-734) temporarily reconfigures the database for performance while maintaining safety. It disables synchronous mode and increases cache_size while preserving WAL mode, allowing rapid insertion of many nodes or edges. The counterpart cbm_store_end_bulk() restores normal pragmas, ensuring that crash protection is never permanently disabled.
Prepared Statement Reuse
All write operations utilize cached prepared statements that are reset and rebound for each use (lines 889-904). This pattern eliminates SQL parsing overhead while guaranteeing that parameter bindings are always correct, preventing data leakage between operations and ensuring that bound values match the expected schema types.
Defensive Design Patterns
The store module incorporates additional safeguards against edge cases including read-only filesystems and memory-mapped I/O errors.
Read-Only Query Mode
The cbm_store_open_path_query() function (lines 720-754) implements a two-tier opening strategy. It first attempts a standard READONLY open; if this fails (for example, on a read-only filesystem), it falls back to an immutable file: URI that bypasses WAL processing entirely. This ensures that query operations can proceed safely even when the underlying storage cannot accept writes, preventing accidental modification attempts.
Deterministic SQLite Functions
Custom functions registered with SQLite—such as regexp, iregexp, cbm_cosine_i8, and cbm_camel_split—are explicitly marked with the SQLITE_DETERMINISTIC flag (lines 635-647). This declaration guarantees that identical inputs always produce identical outputs, preventing non-deterministic side effects from contaminating query plans or index selections.
Memory-Mapped I/O Limits
The module respects the CBM_SQLITE_MMAP_SIZE environment variable (lines 404-416) to cap memory-mapped I/O, defaulting to 64 MiB. This prevents SIGBUS crashes that could occur if the database file is truncated while memory-mapped. Setting the variable to negative values disables memory mapping entirely, providing a safe fallback for unreliable storage environments.
Practical Code Examples
Example 1: Atomic Node Insertion
/* Open (or create) a per-project database */
cbm_store_t *s = cbm_store_open("myproject");
if (!s) { /* handle error */ }
/* Start a transaction */
if (cbm_store_begin(s) != CBM_STORE_OK) { /* handle error */ }
/* Upsert a node */
cbm_node_t node = {
.project = "myproject",
.label = "Function",
.name = "handleRequest",
.qualified_name = "pkg.Service.handleRequest",
.file_path = "src/service.go",
.start_line = 42,
.end_line = 78,
.properties_json = "{\"visibility\":\"public\"}"
};
int64_t node_id = cbm_store_upsert_node(s, &node);
if (node_id < 0) { /* handle error */ }
/* Commit the transaction */
if (cbm_store_commit(s) != CBM_STORE_OK) { /* handle error */ }
cbm_store_close(s);
This example demonstrates transaction isolation and the UNIQUE(project, qualified_name) constraint preventing duplicate nodes.
Example 2: Safe Bulk Edge Insertion
cbm_store_t *s = cbm_store_open("myproject");
cbm_store_begin_bulk(s); // disables sync, enlarges cache
for (int i = 0; i < edge_count; ++i) {
cbm_edge_t e = {
.project = "myproject",
.source_id = src_ids[i],
.target_id = tgt_ids[i],
.type = "CALLS",
.properties_json = NULL // defaults to '{}'
};
int64_t edge_id = cbm_store_insert_edge(s, &e);
if (edge_id < 0) { /* handle individual failure */ }
}
cbm_store_end_bulk(s); // restores normal pragmas
cbm_store_close(s);
Bulk mode accelerates the write path while maintaining WAL safety; the edge UNIQUE constraint prevents duplicate relationships.
Example 3: Post-Crash Integrity Verification
cbm_store_t *s = cbm_store_open_path_query("/var/cache/cbm/myproject.db");
if (!s) { /* DB missing – re-index */ }
if (!cbm_store_check_integrity(s)) {
fprintf(stderr, "Corrupt DB detected – re-index required\n");
/* Recommended: delete the DB and rebuild from source */
}
cbm_store_close(s);
The integrity helper validates row counts and root_path formatting, flagging corruption before the application proceeds.
Summary
The store module ensures data integrity for the graph through multiple reinforcing mechanisms:
- Schema constraints including primary keys, unique composites, and foreign-key cascades prevent invalid data at the database level
- Runtime protections such as the SQLite authorizer and pragma configurations block dangerous operations and enforce ACID compliance
- Transaction management via explicit
cbm_store_begin()andcbm_store_commit()calls ensures atomic updates across nodes and edges - Bulk operations maintain performance without sacrificing crash safety through temporary pragma adjustments
- Integrity checks and read-only fallback modes provide safe recovery paths when storage conditions degrade
Frequently Asked Questions
How does the store module prevent duplicate nodes in the graph?
The schema defines a unique constraint on the combination of project and qualified_name columns in the nodes table (lines 244-265 in src/store/store.c). When cbm_store_upsert_node() executes an INSERT OR REPLACE operation, SQLite enforces this constraint automatically, preventing the same qualified name from appearing twice within a single project while allowing updates to existing node properties.
What happens if a write operation fails midway through a transaction?
Because all writes occur within explicit transactions managed by cbm_store_begin(), any failure before cbm_store_commit() executes leaves the database in its pre-transaction state. The module uses BEGIN IMMEDIATE semantics, which acquires the write lock immediately; if the application crashes or calls cbm_store_rollback(), SQLite automatically discards the pending changes, ensuring that partial graph updates never persist.
How does the module detect database corruption?
The cbm_store_check_integrity() function (lines 822-864) performs lightweight validation by checking that the projects table contains reasonable row counts and that root_path values conform to expected filesystem path patterns. While this does not replace SQLite's internal PRAGMA integrity_check, it provides a fast application-level sanity check that can identify obvious corruption before expensive validation runs.
Can the store operate safely on read-only filesystems?
Yes. The cbm_store_open_path_query() function implements a fallback mechanism (lines 720-754) that detects read-only filesystems and opens the database using an immutable file: URI with the immutable=1 parameter. This mode bypasses WAL processing and write attempts entirely, allowing safe read-only access to graph data without risking I/O errors or corruption warnings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →