What Programming Language Is Used for the Store Module's Graph Storage in Codebase-Memory-MCP
The store module's graph storage is implemented in C, utilizing SQLite as the underlying database engine to persist the code-knowledge graph through relational tables and custom SQL functions.
The codebase-memory-mcp repository by DeusData provides a specialized storage layer for code knowledge graphs, where the core persistence logic resides in a high-performance C library. While the project includes TypeScript components for the user interface, the actual graph storage module—responsible for managing nodes, edges, projects, and file hashes—is written entirely in C and interacts with an embedded SQLite database.
The C-Based Architecture of the Store Module
The graph storage layer is architected as a C library that wraps SQLite operations, providing an opaque handle to manage database connections and transactions.
Opaque Handle and SQLite Connection
In src/store/store.c, the store's opaque handle cbm_store_t encapsulates an sqlite3 * database connection (see lines 1000‑1004). This design pattern abstracts the underlying SQLite implementation while exposing a clean C API declared in src/store/store.h. The handle manages the lifecycle of the database connection, ensuring proper resource management and transaction safety.
Relational Schema for Graph Data
Despite storing graph data, the implementation uses SQLite's relational model rather than a native graph database. The init_schema function (lines 221‑266 of store.c) defines the core tables:
- projects: Stores project metadata
- nodes: Represents code entities (functions, classes, etc.)
- edges: Maintains relationships between nodes
- file_hashes: Tracks content addressing for files
This schema allows the C implementation to leverage SQLite's indexing and query optimization while maintaining graph semantics through foreign key relationships.
CRUD Operations and Graph Traversal in C
All Create, Read, Update, and Delete (CRUD) operations are implemented as C functions that prepare and execute parameterized SQL statements.
Node Management Functions
The cbm_store_upsert_node function (lines 188‑197) handles insertion and updates of graph nodes, while cbm_store_find_node_by_qn (lines 891‑916) retrieves nodes by their qualified name. These functions demonstrate the thin wrapper pattern: they validate inputs, bind parameters to prepared statements, and handle SQLite execution results.
Custom SQLite Functions for Advanced Queries
The C implementation registers custom SQLite functions to support specialized graph operations (lines 635‑645):
regexp: Enables pattern matching for code searchescbm_cosine_i8: Calculates vector similarity for embedding comparisonscbm_camel_split: Parses camelCase identifiers for code analysis
These functions allow complex graph traversal and similarity search to be executed efficiently within the database engine rather than in application code.
Compression Support Modules
The store module includes optional compression layers to optimize storage of vector data, also implemented in C:
internal/cbm/zstd_store.candzstd_store.h: ZSTD compression supportinternal/cbm/lz4_store.candlz4_store.h: LZ4 compression support
These modules integrate with the main store to reduce disk footprint for large embedding vectors while maintaining query performance.
Working with the Store API
The following example demonstrates opening a store, inserting a node, querying by qualified name, and closing the connection:
/* Open (or create) a store for a project */
cbm_store_t *store = cbm_store_open("my_project");
if (!store) { fprintf(stderr, "failed to open store\n"); exit(1); }
/* Insert a node */
cbm_node_t node = {
.project = "my_project",
.label = "Function",
.name = "do_work",
.qualified_name = "my_pkg.do_work",
.file_path = "src/work.c",
.start_line = 10,
.end_line = 20,
.properties_json = "{\"doc\":\"does work\"}"
};
int64_t node_id = cbm_store_upsert_node(store, &node);
printf("Inserted node id: %lld\n", (long long)node_id);
/* Query a node by qualified name */
cbm_node_t fetched;
if (cbm_store_find_node_by_qn(store, "my_project", "my_pkg.do_work", &fetched) == CBM_STORE_OK) {
printf("Found node %s at %s:%d‑%d\n",
fetched.name, fetched.file_path, fetched.start_line, fetched.end_line);
}
/* Close the store */
cbm_store_close(store);
This code illustrates the C API's use of opaque pointers, struct initialization, and error handling patterns typical of the codebase-memory-mcp implementation.
Summary
- The store module's graph storage is written entirely in C, not Python or C++.
- It uses SQLite as the underlying storage format, with tables defined in
src/store/store.c. - The
cbm_store_topaque handle wrapssqlite3 *connections for database management. - CRUD operations like
cbm_store_upsert_nodeare C wrappers around prepared SQL statements. - Custom SQLite functions (
regexp,cbm_cosine_i8,cbm_camel_split) enable advanced graph traversal and similarity search. - Optional compression layers in
internal/cbm/provide additional C implementations for ZSTD and LZ4 algorithms.
Frequently Asked Questions
Is the graph storage implemented in Python or TypeScript?
No. While the codebase-memory-mcp repository includes a TypeScript frontend in the graph-ui/ directory, the actual graph storage layer is written in C. The TypeScript components communicate with the C store module through the MCP binary interface, but all persistence logic resides in src/store/store.c.
Why does the store module use SQLite instead of a dedicated graph database?
The implementation chooses SQLite for its embeddability, zero-configuration requirements, and transactional reliability. As implemented in src/store/store.c, the C code uses relational tables (projects, nodes, edges) with proper indexing to model graph relationships, allowing the system to run without external database dependencies while maintaining ACID compliance.
How does the C store module handle vector similarity search?
The C implementation registers custom SQLite functions such as cbm_cosine_i8 (lines 635‑645 of store.c) to compute cosine similarity between int8 vectors directly within SQL queries. This approach enables efficient similarity search without loading entire datasets into memory, leveraging SQLite's query optimizer for performance.
What compression options are available for the graph storage?
The store module supports optional compression through separate C modules: zstd_store.c for ZSTD compression and lz4_store.c for LZ4 compression. These implementations in internal/cbm/ provide optimized storage for vector embeddings and large property values while maintaining fast decompression speeds for query operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →