# Graph Storage Mechanism in the Store Module: Architecture and Components

> Discover DeusData's graph storage mechanism: an opaque SQLite-backed graph database with a clean C API. Learn about its components and architecture in the store module.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: architecture
- Published: 2026-07-05

---

**The store module implements an opaque SQLite-backed graph database that exposes a clean C API while insulating the rest of the codebase from SQLite internals through the `cbm_store_t` handle, cached prepared statements, and structured CRUD operations for nodes, edges, and projects.**

The `codebase-memory-mcp` repository provides a high-performance knowledge graph system for analyzing codebases. At its heart lies the **graph storage mechanism** implemented in the store module, which leverages SQLite as the underlying engine while presenting a simplified, opaque C interface for all graph operations.

## Opaque Store Handle and Result Codes

The architecture centers on the **`cbm_store_t`** opaque struct defined in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 19-20). This handle encapsulates the SQLite database connection, a cache of prepared statements, and an internal error buffer. By design, callers never interact with SQLite directly; all database operations route through the public API functions.

Operations return simple integer **result codes** defined in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 23-26):
- `CBM_STORE_OK` – Operation succeeded
- `CBM_STORE_ERR` – General error occurred
- `CBM_STORE_NOT_FOUND` – Requested entity does not exist

## Core Data Structures

The graph storage mechanism defines four primary entity structures in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h):

**`cbm_node_t`** (lines 29-38) represents graph vertices such as functions, classes, or files. It contains fields for project affiliation, label type, name, qualified name, file location, and JSON properties.

**`cbm_edge_t`** (lines 41-48) models directed relationships between nodes with source and target IDs, edge type classifications (e.g., `CALLS`), and optional JSON properties.

**`cbm_project_t`** (lines 50-55) stores metadata for indexed projects including the project name, timestamp of last indexing, and root filesystem path.

**`cbm_file_hash_t`** (lines 56-62) maintains cryptographic hashes of source files to enable efficient change detection during incremental updates.

## Schema Initialization and FTS5 Integration

The **`init_schema()`** function in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) (lines 219-274) bootstrap the database by creating the core tables: `projects`, `file_hashes`, `nodes`, and `edges`. It also initializes a content-less **FTS5 virtual table** for full-text search across code entities, enabling fast text queries without external search engines. The function includes legacy schema compatibility checks to handle database migrations gracefully.

## Prepared Statement Caching

To minimize parsing overhead, the store implements a **prepared-statement cache** inside the `cbm_store_t` struct. The helper **`prepare_cached()`** in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) (lines 89-100) lazily prepares frequently used SQL statements (upserts, lookups, deletions) and reuses them for the store's lifetime. This approach dramatically reduces CPU overhead during high-volume graph operations.

## CRUD Operations API

The graph storage mechanism exposes comprehensive CRUD functions organized by entity type:

**Project Operations:**
- `cbm_store_upsert_project()` – Create or update project metadata
- `cbm_store_get_project()` – Retrieve specific project details
- `cbm_store_list_projects()` – Enumerate all indexed projects
- `cbm_store_delete_project()` – Remove project and associated data

**Node Operations:**
- `cbm_store_upsert_node()` – Insert or update nodes (implementation at [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) lines 1888-1915)
- `cbm_store_find_node_by_id()` – Lookup by database ID
- `cbm_store_find_node_by_qn()` – Search by qualified name
- `cbm_store_find_nodes_by_*` family – Filtered queries by various attributes

**Edge Operations:**
- `cbm_store_insert_edge()` – Create relationships between nodes
- `cbm_store_find_edges_by_*` – Query edges by source, target, or type
- `cbm_store_delete_edges_by_*` – Remove edge subsets

**File Hash Operations:**
- `cbm_store_upsert_file_hash()` – Store file checksums
- `cbm_store_get_file_hashes()` – Retrieve hash records
- `cbm_store_delete_file_hash()` – Remove stale entries

## Search and Graph Traversal

For complex queries, the module defines **`cbm_search_params_t`** and **`cbm_search_output_t`** structures (declared in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) lines 401-404). These support filtered searches across node labels, names, file globs, edge types, and degree ranges using regular expressions.

The **`cbm_store_bfs()`** function (lines 410-416) performs breadth-first traversal of the graph, returning hop counts and edge paths. This enables dependency analysis and call-chain exploration without loading the entire graph into memory.

## Bulk Operations and Performance Optimization

The store provides **bulk-write optimization** functions to accelerate large-scale indexing:

- `cbm_store_begin_bulk()` – Temporarily relaxes SQLite pragmas (`synchronous = OFF`, increased cache size) and drops user indexes
- `cbm_store_end_bulk()` – Restores normal operating parameters
- `cbm_store_drop_indexes()` and `cbm_store_create_indexes()` – Manual index management during data loading

These functions, implemented starting at [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) line 774, significantly improve throughput when ingesting large codebases.

## Transaction and Integrity Management

Simple wrappers around SQLite transactions provide atomicity:
- `cbm_store_begin()` – Starts immediate transaction ([`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) lines 960-962)
- `cbm_store_commit()` and `cbm_store_rollback()` – Finalize or abort operations

For maintenance, **`cbm_store_check_integrity()`** (lines 818-868) validates the `projects` table structure and root path formatting. **`cbm_store_checkpoint()`** forces a WAL checkpoint and runs `PRAGMA optimize` to reclaim space and improve query planning.

## Lifecycle Management

Opening functions configure the database environment and return ready-to-use handles:
- `cbm_store_open_memory()` – In-memory database for testing
- `cbm_store_open_path()` – File-based persistent storage
- `cbm_store_open()` – General entry point with configuration options

The internal **`store_open_internal()`** implementation ([`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) lines 1089-1159) registers custom SQLite functions (including REGEXP, case-insensitive pattern matching, and cosine similarity for vector search), applies performance pragmas, and initializes the schema.

Closing via **`cbm_store_close()`** (lines 944-999) finalizes cached statements, checkpoints the WAL, and releases the SQLite handle.

## Practical Usage Example

```c
#include "store/store.h"

/* Open a persistent store for a specific project */
cbm_store_t *store = cbm_store_open("my_project");
if (!store) {
    fprintf(stderr, "Failed to open store: %s\n", cbm_store_error(store));
    return 1;
}

/* Define and upsert a function node */
cbm_node_t func = {
    .project = "my_project",
    .label   = "Function",
    .name    = "do_work",
    .qualified_name = "my_pkg.do_work",
    .file_path = "src/do_work.c",
    .start_line = 10,
    .end_line   = 26,
    .properties_json = "{\"visibility\":\"public\"}"
};

int64_t node_id = cbm_store_upsert_node(store, &func);
if (node_id < 0) {
    fprintf(stderr, "Node upsert failed: %s\n", cbm_store_error(store));
}

/* Search for functions matching a pattern */
cbm_search_params_t params = {
    .project = "my_project",
    .label   = "Function",
    .name_pattern = ".*work.*",
    .case_sensitive = false,
    .limit = 20
};

cbm_search_output_t result;
if (cbm_store_search(store, &params, &result) == CBM_STORE_OK) {
    for (int i = 0; i < result.count; i++) {
        printf("Found: %s at %s:%d\n",
               result.results[i].node.name,
               result.results[i].node.file_path,
               result.results[i].node.start_line);
    }
    cbm_store_search_free(&result);
}

/* Cleanup */
cbm_store_close(store);

```

## Summary

- The **graph storage mechanism** relies on the opaque **`cbm_store_t`** handle to encapsulate SQLite connections and statement caches, ensuring callers remain database-agnostic.
- Four core structures (**`cbm_node_t`**, **`cbm_edge_t`**, **`cbm_project_t`**, **`cbm_file_hash_t`**) model the code knowledge graph with JSON property support.
- Schema initialization creates tables and an **FTS5 full-text search** virtual table for efficient text queries.
- **Prepared statement caching** via `prepare_cached()` minimizes parsing overhead during repeated operations.
- Comprehensive **CRUD APIs** cover projects, nodes, edges, and file hashes with batch operation support.
- **Bulk-write optimization** functions temporarily relax SQLite constraints to accelerate large ingestion tasks.
- **BFS traversal** and parameterized search enable complex graph analysis without loading entire datasets into memory.

## Frequently Asked Questions

### What is the purpose of the `cbm_store_t` opaque struct?

The **`cbm_store_t`** struct defined in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (lines 19-20) serves as the primary handle for all store operations. It encapsulates the SQLite database connection, a cache of prepared statements, and an error message buffer. This design abstracts SQLite internals from the rest of the codebase, allowing the graph storage mechanism to present a clean C API while maintaining flexibility to change underlying storage details without affecting callers.

### How does the store module optimize query performance?

According to the `codebase-memory-mcp` source code, the store employs a **prepared-statement cache** managed by the `prepare_cached()` helper in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) (lines 89-100). This function lazily prepares SQL statements for common operations (upserts, lookups, deletions) and reuses them throughout the store's lifetime, eliminating repetitive SQL parsing overhead. Additionally, bulk operations temporarily disable synchronous commits and drop indexes during large inserts, restoring them afterward for optimal query performance.

### What data structures represent graph entities in the storage mechanism?

The graph storage mechanism defines four primary structures in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h): **`cbm_node_t`** for vertices (functions, classes, files) with location and property metadata; **`cbm_edge_t`** for directed relationships with type classifications; **`cbm_project_t`** for project-level metadata including indexing timestamps; and **`cbm_file_hash_t`** for content-addressable storage of source file checksums. All structures support JSON properties for extensible metadata storage.

### How does bulk write optimization work in the codebase-memory-mcp store?

The **`cbm_store_begin_bulk()`** function in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) (lines 774-785) initiates a high-performance writing mode by setting SQLite pragmas such as `synchronous = OFF` and increasing cache sizes, while optionally dropping user-defined indexes. After bulk operations complete, **`cbm_store_end_bulk()`** restores normal settings and recreates indexes. This approach significantly reduces disk I/O and index maintenance overhead when ingesting large codebases, while **`cbm_store_drop_indexes()`** and **`cbm_store_create_indexes()`** provide manual control over the optimization process.