How the Store Module in DeusData/codebase-memory-mcp Handles Graph Storage: SQLite-Backed Persistence and Traversal
The store module implements an opaque SQLite-backed graph database that manages knowledge-graph persistence through the cbm_store_t handle, offering atomic transactions, bulk indexing, BFS traversal, and vector similarity search without exposing raw SQL to callers.
The store module serves as the core persistence layer for DeusData/codebase-memory-mcp, a system designed to store and query code analysis results as graph structures. It abstracts SQLite operations behind a type-safe C API, enabling developers to persist entities like functions, classes, and their relationships while supporting complex graph algorithms directly against the stored data.
Core Architecture of the SQLite-Backed Graph Store
The architecture centers on an opaque handle pattern that encapsulates the database connection. Callers interact with cbm_store_t, which internally hides the sqlite3* pointer defined in src/store/store.h (lines 19-21), ensuring that all database operations flow through the module's controlled API surface.
The database schema is created lazily on first open via cbm_store_init_schema inside src/store/store.c. This schema defines tables for projects, nodes, edges, and file hashes, along with auxiliary indexes for vector data and graph traversal. Rather than using external SQL files, the schema is embedded in the C implementation, ensuring the binary is self-contained and schema version management is handled programmatically.
Thread safety requires that a single cbm_store_t instance never be used concurrently. Applications must either create one store per thread or implement external synchronization around store operations, as documented in src/store/store.h (lines 7-9).
Initializing and Managing Store Lifecycles
The module provides three distinct opening modes to accommodate different use cases. In-memory stores (cbm_store_open_memory) are useful for testing or temporary analysis, while file-backed stores (cbm_store_open_path) persist data to disk with WAL mode enabled. For read-only scenarios, cbm_store_open_path_query creates a store that queries an existing database without write capabilities.
Closing a store requires cbm_store_close, which properly finalizes statements, closes the SQLite connection, and frees the opaque handle. For long-running processes, cbm_store_checkpoint forces a WAL checkpoint and runs PRAGMA optimize to keep the database file compact, as implemented in src/store/store.h (lines 58-61).
cbm_store_t *store = cbm_store_open_path("my_project.db");
if (!store) { perror("open failed"); exit(1); }
// ... perform operations ...
cbm_store_checkpoint(store); // optional optimization
cbm_store_close(store);
CRUD Operations for Nodes and Edges
The API exposes upsert semantics for graph entities, automatically handling insert-or-update logic for nodes, edges, and file hashes. The cbm_store_upsert_node function takes a populated cbm_node_t structure and returns the stable database ID, while cbm_store_find_node_by_qn retrieves nodes by qualified name.
Query functions that return allocated arrays follow a consistent ownership pattern: the caller receives heap-allocated results that must be freed using the module's specific deallocation functions. For example, cbm_store_node_neighbor_names returns caller-allocated string arrays that require cbm_store_free_nodes to prevent memory leaks, as defined in src/store/store.h (lines 690-730).
cbm_node_t fn = {
.project = "my_project",
.label = "Function",
.name = "do_work",
.qualified_name = "my_pkg.do_work",
.file_path = "src/my_pkg.c",
.start_line = 12,
.end_line = 25,
.properties_json = "{\"static\":true}"
};
int64_t node_id = cbm_store_upsert_node(store, &fn);
printf("Inserted node id %lld\n", (long long)node_id);
Optimizing Ingestion with Bulk Operations
For indexing large codebases, the store module provides bulk write helpers that temporarily relax SQLite pragmas for maximum throughput while maintaining WAL safety. The pattern involves calling cbm_store_begin_bulk before a series of writes, performing operations, then calling cbm_store_end_bulk to restore normal durability settings, as specified in src/store/store.h (lines 42-51).
During bulk operations, the API also supports index dropping and recreation via cbm_store_drop_indexes and cbm_store_create_indexes. This technique avoids index maintenance overhead during heavy inserts, significantly improving ingestion speed for millions of nodes and edges.
cbm_store_begin_bulk(store); // relax pragmas for bulk writes
cbm_node_t nodes[3] = { /* ... */ }; // fill an array of nodes
int64_t ids[3];
cbm_store_upsert_node_batch(store, nodes, 3, ids);
cbm_store_end_bulk(store); // restore normal pragmas
Atomic Transaction Control
Outside of bulk mode, explicit transactions allow grouping multiple writes atomically. The functions cbm_store_begin, cbm_store_commit, and cbm_store_rollback provide standard ACID guarantees across src/store/store.h (lines 31-40), ensuring that complex updates to nodes and edges remain consistent even if the process encounters errors mid-operation.
Querying and Traversing the Graph
Beyond simple lookups, the store module implements graph-native query capabilities directly in C. The cbm_store_bfs function performs breadth-first traversal with edge-type filtering, accepting parameters for direction (inbound/outbound), edge types, maximum depth, and result limits, as declared in src/store/store.h (lines 108-112).
For general filtering, cbm_store_search provides a rich filter API over node properties and relationships without requiring manual SQL construction.
cbm_traverse_result_t result;
int rc = cbm_store_bfs(store, start_id, "outbound",
(const char *[]){"CALLS"}, 1,
5, // max depth
100, // max results
&result);
if (rc == CBM_STORE_OK) {
for (int i = 0; i < result.edge_count; ++i) {
printf("Edge %d: %s → %s (%s)\n",
i,
result.edges[i].source_id,
result.edges[i].target_id,
result.edges[i].type);
}
cbm_store_traverse_free(&result);
}
Vector Similarity Search
The module supports semantic search through cbm_store_vector_search, which computes cosine similarity over stored RI (Random Indexing) vectors. This allows finding semantically similar code entities based on keyword vectors rather than exact string matching, implemented in src/store/store.h (lines 70-86).
const char *kw[] = {"parse", "token"};
cbm_vector_result_t *vec_res = NULL;
int vec_cnt = 0;
cbm_store_vector_search(store, "my_project", kw, 2, 10,
&vec_res, &vec_cnt);
for (int i = 0; i < vec_cnt; ++i) {
printf("Score %.3f – %s (%s)\n",
vec_res[i].score,
vec_res[i].name,
vec_res[i].qualified_name);
}
cbm_store_free_vector_results(vec_res, vec_cnt);
Advanced Analytics and Community Detection
The store module extends beyond simple storage to support architecture extraction and impact analysis. The cbm_store_get_architecture function builds multi-aspect views by joining node and edge data, identifying languages, packages, entry points, routing tables, and hotspots within the codebase, as defined in src/store/store.h (lines 560-587).
For dependency analysis, cbm_hop_to_risk and cbm_build_impact_summary map graph traversal hops to risk levels, summarizing critical paths through the call graph (lines 518-525). These functions enable downstream tools to visualize not just static structure but dynamic behavioral impact.
Graph Algorithms Implementation
Native implementations of Leiden and Louvain community detection algorithms (cbm_leiden, cbm_louvain) operate directly on the SQLite-stored graph, returning community assignments for each node as specified in src/store/store.h (lines 640-658). These algorithms enable automatic modularization suggestions and architectural boundary detection without exporting data to external graph tools.
Summary
- The store module in
DeusData/codebase-memory-mcpprovides an opaque SQLite-backed graph database accessible through thecbm_store_thandle defined insrc/store/store.h. - CRUD operations use upsert semantics via functions like
cbm_store_upsert_node, with memory management responsibilities clearly delegated to the caller through dedicated free functions such ascbm_store_free_nodes. - Bulk operations (
cbm_store_begin_bulk,cbm_store_end_bulk) optimize large-scale ingestion by temporarily relaxing SQLite pragmas, while explicit transactions (cbm_store_begin,cbm_store_commit) ensure ACID compliance for critical updates. - Graph traversal capabilities include BFS (
cbm_store_bfs), neighbor lookup (cbm_store_node_neighbor_names), and vector similarity search (cbm_store_vector_search) for semantic code queries. - Advanced analytics such as community detection (
cbm_leiden,cbm_louvain) and architecture extraction (cbm_store_get_architecture) run natively against the stored graph without external dependencies.
Frequently Asked Questions
What is the primary purpose of the cbm_store_t opaque handle?
The cbm_store_t opaque handle serves as the primary interface to the SQLite database connection, intentionally hiding the underlying sqlite3* pointer from callers as defined in src/store/store.h (lines 19-21). This abstraction prevents direct SQL manipulation and ensures all database interactions occur through the validated API surface, enabling the implementation to manage connection state, schema versioning, and WAL configuration internally.
How does the store module handle concurrent access from multiple threads?
According to the source code in src/store/store.h (lines 7-9), a single cbm_store_t instance must not be used concurrently across threads. Applications requiring multi-threaded access must either create one store instance per thread or implement external synchronization mechanisms (such as mutexes) around all store operations to prevent race conditions and database corruption.
What is the difference between bulk operations and regular transactions?
Bulk operations (cbm_store_begin_bulk / cbm_store_end_bulk) are optimized for high-throughput ingestion of large codebases, temporarily relaxing SQLite pragmas (such as synchronous mode and journal settings) to maximize write speed while maintaining WAL safety. Regular transactions (cbm_store_begin / cbm_store_commit / cbm_store_rollback) provide standard ACID guarantees for general-purpose writes without the performance optimizations that sacrifice durability during the bulk window.
How does the store module support semantic search across codebases?
The module implements vector similarity search through cbm_store_vector_search, which computes cosine similarity over stored Random Indexing (RI) vectors associated with nodes. This allows developers to find semantically related functions or classes based on keyword vectors rather than exact string matches, enabling discovery of conceptually similar code across different naming conventions or file locations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →