# How the Store Module in DeusData/codebase-memory-mcp Handles Graph Storage: SQLite-Backed Persistence and Traversal

> Explore how DeusData/codebase-memory-mcp's store module uses SQLite for efficient graph storage, enabling atomic transactions, indexing, BFS traversal, and vector search.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: internals
- Published: 2026-07-05

---

**The store module implements an opaque SQLite-backed graph database that manages knowledge-graph persistence through the `cbm_store_t` handle, offering atomic transactions, bulk indexing, BFS traversal, and vector similarity search without exposing raw SQL to callers.**

The **store module** serves as the core persistence layer for `DeusData/codebase-memory-mcp`, a system designed to store and query code analysis results as graph structures. It abstracts SQLite operations behind a type-safe C API, enabling developers to persist entities like functions, classes, and their relationships while supporting complex graph algorithms directly against the stored data.

## Core Architecture of the SQLite-Backed Graph Store

The architecture centers on an **opaque handle pattern** that encapsulates the database connection. Callers interact with `cbm_store_t`, which internally hides the `sqlite3*` pointer defined in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 19-21), ensuring that all database operations flow through the module's controlled API surface.

The database schema is created lazily on first open via `cbm_store_init_schema` inside [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c). This schema defines tables for **projects**, **nodes**, **edges**, and **file hashes**, along with auxiliary indexes for vector data and graph traversal. Rather than using external SQL files, the schema is embedded in the C implementation, ensuring the binary is self-contained and schema version management is handled programmatically.

Thread safety requires that a single `cbm_store_t` instance never be used concurrently. Applications must either create one store per thread or implement external synchronization around store operations, as documented in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 7-9).

## Initializing and Managing Store Lifecycles

The module provides three distinct opening modes to accommodate different use cases. **In-memory stores** (`cbm_store_open_memory`) are useful for testing or temporary analysis, while **file-backed stores** (`cbm_store_open_path`) persist data to disk with WAL mode enabled. For read-only scenarios, `cbm_store_open_path_query` creates a store that queries an existing database without write capabilities.

Closing a store requires `cbm_store_close`, which properly finalizes statements, closes the SQLite connection, and frees the opaque handle. For long-running processes, `cbm_store_checkpoint` forces a WAL checkpoint and runs `PRAGMA optimize` to keep the database file compact, as implemented in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 58-61).

```c
cbm_store_t *store = cbm_store_open_path("my_project.db");
if (!store) { perror("open failed"); exit(1); }

// ... perform operations ...

cbm_store_checkpoint(store);  // optional optimization
cbm_store_close(store);

```

## CRUD Operations for Nodes and Edges

The API exposes **upsert** semantics for graph entities, automatically handling insert-or-update logic for nodes, edges, and file hashes. The `cbm_store_upsert_node` function takes a populated `cbm_node_t` structure and returns the stable database ID, while `cbm_store_find_node_by_qn` retrieves nodes by qualified name.

Query functions that return allocated arrays follow a consistent ownership pattern: the caller receives heap-allocated results that must be freed using the module's specific deallocation functions. For example, `cbm_store_node_neighbor_names` returns caller-allocated string arrays that require `cbm_store_free_nodes` to prevent memory leaks, as defined in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 690-730).

```c
cbm_node_t fn = {
    .project = "my_project",
    .label   = "Function",
    .name    = "do_work",
    .qualified_name = "my_pkg.do_work",
    .file_path = "src/my_pkg.c",
    .start_line = 12,
    .end_line   = 25,
    .properties_json = "{\"static\":true}"
};

int64_t node_id = cbm_store_upsert_node(store, &fn);
printf("Inserted node id %lld\n", (long long)node_id);

```

## Optimizing Ingestion with Bulk Operations

For indexing large codebases, the store module provides **bulk write helpers** that temporarily relax SQLite pragmas for maximum throughput while maintaining WAL safety. The pattern involves calling `cbm_store_begin_bulk` before a series of writes, performing operations, then calling `cbm_store_end_bulk` to restore normal durability settings, as specified in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 42-51).

During bulk operations, the API also supports **index dropping and recreation** via `cbm_store_drop_indexes` and `cbm_store_create_indexes`. This technique avoids index maintenance overhead during heavy inserts, significantly improving ingestion speed for millions of nodes and edges.

```c
cbm_store_begin_bulk(store);                // relax pragmas for bulk writes
cbm_node_t nodes[3] = { /* ... */ };        // fill an array of nodes
int64_t ids[3];
cbm_store_upsert_node_batch(store, nodes, 3, ids);
cbm_store_end_bulk(store);                  // restore normal pragmas

```

### Atomic Transaction Control

Outside of bulk mode, explicit transactions allow grouping multiple writes atomically. The functions `cbm_store_begin`, `cbm_store_commit`, and `cbm_store_rollback` provide standard ACID guarantees across [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 31-40), ensuring that complex updates to nodes and edges remain consistent even if the process encounters errors mid-operation.

## Querying and Traversing the Graph

Beyond simple lookups, the store module implements **graph-native query capabilities** directly in C. The `cbm_store_bfs` function performs breadth-first traversal with edge-type filtering, accepting parameters for direction (inbound/outbound), edge types, maximum depth, and result limits, as declared in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 108-112).

For general filtering, `cbm_store_search` provides a rich filter API over node properties and relationships without requiring manual SQL construction.

```c
cbm_traverse_result_t result;
int rc = cbm_store_bfs(store, start_id, "outbound",
                       (const char *[]){"CALLS"}, 1,
                       5,    // max depth
                       100,  // max results
                       &result);

if (rc == CBM_STORE_OK) {
    for (int i = 0; i < result.edge_count; ++i) {
        printf("Edge %d: %s → %s (%s)\n",
               i,
               result.edges[i].source_id,
               result.edges[i].target_id,
               result.edges[i].type);
    }
    cbm_store_traverse_free(&result);
}

```

### Vector Similarity Search

The module supports **semantic search** through `cbm_store_vector_search`, which computes cosine similarity over stored RI (Random Indexing) vectors. This allows finding semantically similar code entities based on keyword vectors rather than exact string matching, implemented in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 70-86).

```c
const char *kw[] = {"parse", "token"};
cbm_vector_result_t *vec_res = NULL;
int vec_cnt = 0;

cbm_store_vector_search(store, "my_project", kw, 2, 10,
                        &vec_res, &vec_cnt);

for (int i = 0; i < vec_cnt; ++i) {
    printf("Score %.3f – %s (%s)\n",
           vec_res[i].score,
           vec_res[i].name,
           vec_res[i].qualified_name);
}
cbm_store_free_vector_results(vec_res, vec_cnt);

```

## Advanced Analytics and Community Detection

The store module extends beyond simple storage to support **architecture extraction** and **impact analysis**. The `cbm_store_get_architecture` function builds multi-aspect views by joining node and edge data, identifying languages, packages, entry points, routing tables, and hotspots within the codebase, as defined in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 560-587).

For dependency analysis, `cbm_hop_to_risk` and `cbm_build_impact_summary` map graph traversal hops to risk levels, summarizing critical paths through the call graph (lines 518-525). These functions enable downstream tools to visualize not just static structure but dynamic behavioral impact.

### Graph Algorithms Implementation

Native implementations of **Leiden and Louvain community detection algorithms** (`cbm_leiden`, `cbm_louvain`) operate directly on the SQLite-stored graph, returning community assignments for each node as specified in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 640-658). These algorithms enable automatic modularization suggestions and architectural boundary detection without exporting data to external graph tools.

## Summary

- The **store module** in `DeusData/codebase-memory-mcp` provides an **opaque SQLite-backed graph database** accessible through the `cbm_store_t` handle defined in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h).
- **CRUD operations** use upsert semantics via functions like `cbm_store_upsert_node`, with memory management responsibilities clearly delegated to the caller through dedicated free functions such as `cbm_store_free_nodes`.
- **Bulk operations** (`cbm_store_begin_bulk`, `cbm_store_end_bulk`) optimize large-scale ingestion by temporarily relaxing SQLite pragmas, while explicit transactions (`cbm_store_begin`, `cbm_store_commit`) ensure ACID compliance for critical updates.
- **Graph traversal** capabilities include BFS (`cbm_store_bfs`), neighbor lookup (`cbm_store_node_neighbor_names`), and **vector similarity search** (`cbm_store_vector_search`) for semantic code queries.
- **Advanced analytics** such as community detection (`cbm_leiden`, `cbm_louvain`) and architecture extraction (`cbm_store_get_architecture`) run natively against the stored graph without external dependencies.

## Frequently Asked Questions

### What is the primary purpose of the cbm_store_t opaque handle?

The `cbm_store_t` opaque handle serves as the primary interface to the SQLite database connection, intentionally hiding the underlying `sqlite3*` pointer from callers as defined in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 19-21). This abstraction prevents direct SQL manipulation and ensures all database interactions occur through the validated API surface, enabling the implementation to manage connection state, schema versioning, and WAL configuration internally.

### How does the store module handle concurrent access from multiple threads?

According to the source code in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 7-9), a single `cbm_store_t` instance must not be used concurrently across threads. Applications requiring multi-threaded access must either create one store instance per thread or implement external synchronization mechanisms (such as mutexes) around all store operations to prevent race conditions and database corruption.

### What is the difference between bulk operations and regular transactions?

**Bulk operations** (`cbm_store_begin_bulk` / `cbm_store_end_bulk`) are optimized for high-throughput ingestion of large codebases, temporarily relaxing SQLite pragmas (such as synchronous mode and journal settings) to maximize write speed while maintaining WAL safety. **Regular transactions** (`cbm_store_begin` / `cbm_store_commit` / `cbm_store_rollback`) provide standard ACID guarantees for general-purpose writes without the performance optimizations that sacrifice durability during the bulk window.

### How does the store module support semantic search across codebases?

The module implements **vector similarity search** through `cbm_store_vector_search`, which computes cosine similarity over stored Random Indexing (RI) vectors associated with nodes. This allows developers to find semantically related functions or classes based on keyword vectors rather than exact string matches, enabling discovery of conceptually similar code across different naming conventions or file locations.