Graph Storage Mechanism in the Store Module: Architecture and Components

The store module implements an opaque SQLite-backed graph database that exposes a clean C API while insulating the rest of the codebase from SQLite internals through the cbm_store_t handle, cached prepared statements, and structured CRUD operations for nodes, edges, and projects.

The codebase-memory-mcp repository provides a high-performance knowledge graph system for analyzing codebases. At its heart lies the graph storage mechanism implemented in the store module, which leverages SQLite as the underlying engine while presenting a simplified, opaque C interface for all graph operations.

Opaque Store Handle and Result Codes

The architecture centers on the cbm_store_t opaque struct defined in src/store/store.h (lines 19-20). This handle encapsulates the SQLite database connection, a cache of prepared statements, and an internal error buffer. By design, callers never interact with SQLite directly; all database operations route through the public API functions.

Operations return simple integer result codes defined in src/store/store.h (lines 23-26):

  • CBM_STORE_OK – Operation succeeded
  • CBM_STORE_ERR – General error occurred
  • CBM_STORE_NOT_FOUND – Requested entity does not exist

Core Data Structures

The graph storage mechanism defines four primary entity structures in src/store/store.h:

cbm_node_t (lines 29-38) represents graph vertices such as functions, classes, or files. It contains fields for project affiliation, label type, name, qualified name, file location, and JSON properties.

cbm_edge_t (lines 41-48) models directed relationships between nodes with source and target IDs, edge type classifications (e.g., CALLS), and optional JSON properties.

cbm_project_t (lines 50-55) stores metadata for indexed projects including the project name, timestamp of last indexing, and root filesystem path.

cbm_file_hash_t (lines 56-62) maintains cryptographic hashes of source files to enable efficient change detection during incremental updates.

Schema Initialization and FTS5 Integration

The init_schema() function in src/store/store.c (lines 219-274) bootstrap the database by creating the core tables: projects, file_hashes, nodes, and edges. It also initializes a content-less FTS5 virtual table for full-text search across code entities, enabling fast text queries without external search engines. The function includes legacy schema compatibility checks to handle database migrations gracefully.

Prepared Statement Caching

To minimize parsing overhead, the store implements a prepared-statement cache inside the cbm_store_t struct. The helper prepare_cached() in src/store/store.c (lines 89-100) lazily prepares frequently used SQL statements (upserts, lookups, deletions) and reuses them for the store's lifetime. This approach dramatically reduces CPU overhead during high-volume graph operations.

CRUD Operations API

The graph storage mechanism exposes comprehensive CRUD functions organized by entity type:

Project Operations:

  • cbm_store_upsert_project() – Create or update project metadata
  • cbm_store_get_project() – Retrieve specific project details
  • cbm_store_list_projects() – Enumerate all indexed projects
  • cbm_store_delete_project() – Remove project and associated data

Node Operations:

  • cbm_store_upsert_node() – Insert or update nodes (implementation at src/store/store.c lines 1888-1915)
  • cbm_store_find_node_by_id() – Lookup by database ID
  • cbm_store_find_node_by_qn() – Search by qualified name
  • cbm_store_find_nodes_by_* family – Filtered queries by various attributes

Edge Operations:

  • cbm_store_insert_edge() – Create relationships between nodes
  • cbm_store_find_edges_by_* – Query edges by source, target, or type
  • cbm_store_delete_edges_by_* – Remove edge subsets

File Hash Operations:

  • cbm_store_upsert_file_hash() – Store file checksums
  • cbm_store_get_file_hashes() – Retrieve hash records
  • cbm_store_delete_file_hash() – Remove stale entries

Search and Graph Traversal

For complex queries, the module defines cbm_search_params_t and cbm_search_output_t structures (declared in src/store/store.h lines 401-404). These support filtered searches across node labels, names, file globs, edge types, and degree ranges using regular expressions.

The cbm_store_bfs() function (lines 410-416) performs breadth-first traversal of the graph, returning hop counts and edge paths. This enables dependency analysis and call-chain exploration without loading the entire graph into memory.

Bulk Operations and Performance Optimization

The store provides bulk-write optimization functions to accelerate large-scale indexing:

  • cbm_store_begin_bulk() – Temporarily relaxes SQLite pragmas (synchronous = OFF, increased cache size) and drops user indexes
  • cbm_store_end_bulk() – Restores normal operating parameters
  • cbm_store_drop_indexes() and cbm_store_create_indexes() – Manual index management during data loading

These functions, implemented starting at src/store/store.c line 774, significantly improve throughput when ingesting large codebases.

Transaction and Integrity Management

Simple wrappers around SQLite transactions provide atomicity:

  • cbm_store_begin() – Starts immediate transaction (src/store/store.c lines 960-962)
  • cbm_store_commit() and cbm_store_rollback() – Finalize or abort operations

For maintenance, cbm_store_check_integrity() (lines 818-868) validates the projects table structure and root path formatting. cbm_store_checkpoint() forces a WAL checkpoint and runs PRAGMA optimize to reclaim space and improve query planning.

Lifecycle Management

Opening functions configure the database environment and return ready-to-use handles:

  • cbm_store_open_memory() – In-memory database for testing
  • cbm_store_open_path() – File-based persistent storage
  • cbm_store_open() – General entry point with configuration options

The internal store_open_internal() implementation (src/store/store.c lines 1089-1159) registers custom SQLite functions (including REGEXP, case-insensitive pattern matching, and cosine similarity for vector search), applies performance pragmas, and initializes the schema.

Closing via cbm_store_close() (lines 944-999) finalizes cached statements, checkpoints the WAL, and releases the SQLite handle.

Practical Usage Example

#include "store/store.h"

/* Open a persistent store for a specific project */
cbm_store_t *store = cbm_store_open("my_project");
if (!store) {
    fprintf(stderr, "Failed to open store: %s\n", cbm_store_error(store));
    return 1;
}

/* Define and upsert a function node */
cbm_node_t func = {
    .project = "my_project",
    .label   = "Function",
    .name    = "do_work",
    .qualified_name = "my_pkg.do_work",
    .file_path = "src/do_work.c",
    .start_line = 10,
    .end_line   = 26,
    .properties_json = "{\"visibility\":\"public\"}"
};

int64_t node_id = cbm_store_upsert_node(store, &func);
if (node_id < 0) {
    fprintf(stderr, "Node upsert failed: %s\n", cbm_store_error(store));
}

/* Search for functions matching a pattern */
cbm_search_params_t params = {
    .project = "my_project",
    .label   = "Function",
    .name_pattern = ".*work.*",
    .case_sensitive = false,
    .limit = 20
};

cbm_search_output_t result;
if (cbm_store_search(store, &params, &result) == CBM_STORE_OK) {
    for (int i = 0; i < result.count; i++) {
        printf("Found: %s at %s:%d\n",
               result.results[i].node.name,
               result.results[i].node.file_path,
               result.results[i].node.start_line);
    }
    cbm_store_search_free(&result);
}

/* Cleanup */
cbm_store_close(store);

Summary

  • The graph storage mechanism relies on the opaque cbm_store_t handle to encapsulate SQLite connections and statement caches, ensuring callers remain database-agnostic.
  • Four core structures (cbm_node_t, cbm_edge_t, cbm_project_t, cbm_file_hash_t) model the code knowledge graph with JSON property support.
  • Schema initialization creates tables and an FTS5 full-text search virtual table for efficient text queries.
  • Prepared statement caching via prepare_cached() minimizes parsing overhead during repeated operations.
  • Comprehensive CRUD APIs cover projects, nodes, edges, and file hashes with batch operation support.
  • Bulk-write optimization functions temporarily relax SQLite constraints to accelerate large ingestion tasks.
  • BFS traversal and parameterized search enable complex graph analysis without loading entire datasets into memory.

Frequently Asked Questions

What is the purpose of the cbm_store_t opaque struct?

The cbm_store_t struct defined in src/store/store.h (lines 19-20) serves as the primary handle for all store operations. It encapsulates the SQLite database connection, a cache of prepared statements, and an error message buffer. This design abstracts SQLite internals from the rest of the codebase, allowing the graph storage mechanism to present a clean C API while maintaining flexibility to change underlying storage details without affecting callers.

How does the store module optimize query performance?

According to the codebase-memory-mcp source code, the store employs a prepared-statement cache managed by the prepare_cached() helper in src/store/store.c (lines 89-100). This function lazily prepares SQL statements for common operations (upserts, lookups, deletions) and reuses them throughout the store's lifetime, eliminating repetitive SQL parsing overhead. Additionally, bulk operations temporarily disable synchronous commits and drop indexes during large inserts, restoring them afterward for optimal query performance.

What data structures represent graph entities in the storage mechanism?

The graph storage mechanism defines four primary structures in src/store/store.h: cbm_node_t for vertices (functions, classes, files) with location and property metadata; cbm_edge_t for directed relationships with type classifications; cbm_project_t for project-level metadata including indexing timestamps; and cbm_file_hash_t for content-addressable storage of source file checksums. All structures support JSON properties for extensible metadata storage.

How does bulk write optimization work in the codebase-memory-mcp store?

The cbm_store_begin_bulk() function in src/store/store.c (lines 774-785) initiates a high-performance writing mode by setting SQLite pragmas such as synchronous = OFF and increasing cache sizes, while optionally dropping user-defined indexes. After bulk operations complete, cbm_store_end_bulk() restores normal settings and recreates indexes. This approach significantly reduces disk I/O and index maintenance overhead when ingesting large codebases, while cbm_store_drop_indexes() and cbm_store_create_indexes() provide manual control over the optimization process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →