Graph Storage Mechanism in the Store Module: Architecture and Components
The store module implements an opaque SQLite-backed graph database that exposes a clean C API while insulating the rest of the codebase from SQLite internals through the cbm_store_t handle, cached prepared statements, and structured CRUD operations for nodes, edges, and projects.
The codebase-memory-mcp repository provides a high-performance knowledge graph system for analyzing codebases. At its heart lies the graph storage mechanism implemented in the store module, which leverages SQLite as the underlying engine while presenting a simplified, opaque C interface for all graph operations.
Opaque Store Handle and Result Codes
The architecture centers on the cbm_store_t opaque struct defined in src/store/store.h (lines 19-20). This handle encapsulates the SQLite database connection, a cache of prepared statements, and an internal error buffer. By design, callers never interact with SQLite directly; all database operations route through the public API functions.
Operations return simple integer result codes defined in src/store/store.h (lines 23-26):
CBM_STORE_OK– Operation succeededCBM_STORE_ERR– General error occurredCBM_STORE_NOT_FOUND– Requested entity does not exist
Core Data Structures
The graph storage mechanism defines four primary entity structures in src/store/store.h:
cbm_node_t (lines 29-38) represents graph vertices such as functions, classes, or files. It contains fields for project affiliation, label type, name, qualified name, file location, and JSON properties.
cbm_edge_t (lines 41-48) models directed relationships between nodes with source and target IDs, edge type classifications (e.g., CALLS), and optional JSON properties.
cbm_project_t (lines 50-55) stores metadata for indexed projects including the project name, timestamp of last indexing, and root filesystem path.
cbm_file_hash_t (lines 56-62) maintains cryptographic hashes of source files to enable efficient change detection during incremental updates.
Schema Initialization and FTS5 Integration
The init_schema() function in src/store/store.c (lines 219-274) bootstrap the database by creating the core tables: projects, file_hashes, nodes, and edges. It also initializes a content-less FTS5 virtual table for full-text search across code entities, enabling fast text queries without external search engines. The function includes legacy schema compatibility checks to handle database migrations gracefully.
Prepared Statement Caching
To minimize parsing overhead, the store implements a prepared-statement cache inside the cbm_store_t struct. The helper prepare_cached() in src/store/store.c (lines 89-100) lazily prepares frequently used SQL statements (upserts, lookups, deletions) and reuses them for the store's lifetime. This approach dramatically reduces CPU overhead during high-volume graph operations.
CRUD Operations API
The graph storage mechanism exposes comprehensive CRUD functions organized by entity type:
Project Operations:
cbm_store_upsert_project()– Create or update project metadatacbm_store_get_project()– Retrieve specific project detailscbm_store_list_projects()– Enumerate all indexed projectscbm_store_delete_project()– Remove project and associated data
Node Operations:
cbm_store_upsert_node()– Insert or update nodes (implementation atsrc/store/store.clines 1888-1915)cbm_store_find_node_by_id()– Lookup by database IDcbm_store_find_node_by_qn()– Search by qualified namecbm_store_find_nodes_by_*family – Filtered queries by various attributes
Edge Operations:
cbm_store_insert_edge()– Create relationships between nodescbm_store_find_edges_by_*– Query edges by source, target, or typecbm_store_delete_edges_by_*– Remove edge subsets
File Hash Operations:
cbm_store_upsert_file_hash()– Store file checksumscbm_store_get_file_hashes()– Retrieve hash recordscbm_store_delete_file_hash()– Remove stale entries
Search and Graph Traversal
For complex queries, the module defines cbm_search_params_t and cbm_search_output_t structures (declared in src/store/store.h lines 401-404). These support filtered searches across node labels, names, file globs, edge types, and degree ranges using regular expressions.
The cbm_store_bfs() function (lines 410-416) performs breadth-first traversal of the graph, returning hop counts and edge paths. This enables dependency analysis and call-chain exploration without loading the entire graph into memory.
Bulk Operations and Performance Optimization
The store provides bulk-write optimization functions to accelerate large-scale indexing:
cbm_store_begin_bulk()– Temporarily relaxes SQLite pragmas (synchronous = OFF, increased cache size) and drops user indexescbm_store_end_bulk()– Restores normal operating parameterscbm_store_drop_indexes()andcbm_store_create_indexes()– Manual index management during data loading
These functions, implemented starting at src/store/store.c line 774, significantly improve throughput when ingesting large codebases.
Transaction and Integrity Management
Simple wrappers around SQLite transactions provide atomicity:
cbm_store_begin()– Starts immediate transaction (src/store/store.clines 960-962)cbm_store_commit()andcbm_store_rollback()– Finalize or abort operations
For maintenance, cbm_store_check_integrity() (lines 818-868) validates the projects table structure and root path formatting. cbm_store_checkpoint() forces a WAL checkpoint and runs PRAGMA optimize to reclaim space and improve query planning.
Lifecycle Management
Opening functions configure the database environment and return ready-to-use handles:
cbm_store_open_memory()– In-memory database for testingcbm_store_open_path()– File-based persistent storagecbm_store_open()– General entry point with configuration options
The internal store_open_internal() implementation (src/store/store.c lines 1089-1159) registers custom SQLite functions (including REGEXP, case-insensitive pattern matching, and cosine similarity for vector search), applies performance pragmas, and initializes the schema.
Closing via cbm_store_close() (lines 944-999) finalizes cached statements, checkpoints the WAL, and releases the SQLite handle.
Practical Usage Example
#include "store/store.h"
/* Open a persistent store for a specific project */
cbm_store_t *store = cbm_store_open("my_project");
if (!store) {
fprintf(stderr, "Failed to open store: %s\n", cbm_store_error(store));
return 1;
}
/* Define and upsert a function node */
cbm_node_t func = {
.project = "my_project",
.label = "Function",
.name = "do_work",
.qualified_name = "my_pkg.do_work",
.file_path = "src/do_work.c",
.start_line = 10,
.end_line = 26,
.properties_json = "{\"visibility\":\"public\"}"
};
int64_t node_id = cbm_store_upsert_node(store, &func);
if (node_id < 0) {
fprintf(stderr, "Node upsert failed: %s\n", cbm_store_error(store));
}
/* Search for functions matching a pattern */
cbm_search_params_t params = {
.project = "my_project",
.label = "Function",
.name_pattern = ".*work.*",
.case_sensitive = false,
.limit = 20
};
cbm_search_output_t result;
if (cbm_store_search(store, ¶ms, &result) == CBM_STORE_OK) {
for (int i = 0; i < result.count; i++) {
printf("Found: %s at %s:%d\n",
result.results[i].node.name,
result.results[i].node.file_path,
result.results[i].node.start_line);
}
cbm_store_search_free(&result);
}
/* Cleanup */
cbm_store_close(store);
Summary
- The graph storage mechanism relies on the opaque
cbm_store_thandle to encapsulate SQLite connections and statement caches, ensuring callers remain database-agnostic. - Four core structures (
cbm_node_t,cbm_edge_t,cbm_project_t,cbm_file_hash_t) model the code knowledge graph with JSON property support. - Schema initialization creates tables and an FTS5 full-text search virtual table for efficient text queries.
- Prepared statement caching via
prepare_cached()minimizes parsing overhead during repeated operations. - Comprehensive CRUD APIs cover projects, nodes, edges, and file hashes with batch operation support.
- Bulk-write optimization functions temporarily relax SQLite constraints to accelerate large ingestion tasks.
- BFS traversal and parameterized search enable complex graph analysis without loading entire datasets into memory.
Frequently Asked Questions
What is the purpose of the cbm_store_t opaque struct?
The cbm_store_t struct defined in src/store/store.h (lines 19-20) serves as the primary handle for all store operations. It encapsulates the SQLite database connection, a cache of prepared statements, and an error message buffer. This design abstracts SQLite internals from the rest of the codebase, allowing the graph storage mechanism to present a clean C API while maintaining flexibility to change underlying storage details without affecting callers.
How does the store module optimize query performance?
According to the codebase-memory-mcp source code, the store employs a prepared-statement cache managed by the prepare_cached() helper in src/store/store.c (lines 89-100). This function lazily prepares SQL statements for common operations (upserts, lookups, deletions) and reuses them throughout the store's lifetime, eliminating repetitive SQL parsing overhead. Additionally, bulk operations temporarily disable synchronous commits and drop indexes during large inserts, restoring them afterward for optimal query performance.
What data structures represent graph entities in the storage mechanism?
The graph storage mechanism defines four primary structures in src/store/store.h: cbm_node_t for vertices (functions, classes, files) with location and property metadata; cbm_edge_t for directed relationships with type classifications; cbm_project_t for project-level metadata including indexing timestamps; and cbm_file_hash_t for content-addressable storage of source file checksums. All structures support JSON properties for extensible metadata storage.
How does bulk write optimization work in the codebase-memory-mcp store?
The cbm_store_begin_bulk() function in src/store/store.c (lines 774-785) initiates a high-performance writing mode by setting SQLite pragmas such as synchronous = OFF and increasing cache sizes, while optionally dropping user-defined indexes. After bulk operations complete, cbm_store_end_bulk() restores normal settings and recreates indexes. This approach significantly reduces disk I/O and index maintenance overhead when ingesting large codebases, while cbm_store_drop_indexes() and cbm_store_create_indexes() provide manual control over the optimization process.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →