Performance Considerations for Graph Storage in the Store Module: Optimizing SQLite in codebase-memory-mcp

The store module requires strict single-threaded handle usage, strategic bulk-write pragmas, and periodic WAL checkpointing to maintain high throughput when persisting code-knowledge graphs.

The store module in DeusData/codebase-memory-mcp provides an opaque SQLite-backed graph storage layer designed to persist code-knowledge relationships with minimal overhead. Understanding the performance considerations for the graph storage in the store module is essential when indexing large repositories or operating under concurrent workloads. This guide examines the architectural constraints, configuration options, and API patterns implemented in src/store/store.h and src/store/store.c that directly impact scalability.

Thread Safety and Connection Handling

SQLite connections are not thread-safe unless explicitly compiled with SQLITE_THREADSAFE=1, and the store module enforces strict single-threaded access per handle. As declared in src/store/store.h at lines 7-9, a single store handle must not be used concurrently across threads.

To achieve parallelism, create one store per thread using cbm_store_open_path() or cbm_store_open_memory(), or implement external synchronization with a mutex. Sharing a cbm_store_t * pointer between goroutines or pthreads causes lock contention, race conditions, and potential crashes.

/* Safe pattern: one handle per thread */
void *worker_thread(void *arg) {
    const char *project = (const char *)arg;
    cbm_store_t *store = cbm_store_open_path(project);
    if (!store) return NULL;
    
    /* Perform read-only queries or writes */
    cbm_store_close(store);
    return NULL;
}

Storage Backend Selection

The module supports two distinct storage modes with different performance characteristics:

  • cbm_store_open_memory (lines 98-102): Creates a transient :memory: database that bypasses filesystem I/O for maximum read/write speed. Data is lost when the handle closes, making this ideal for testing or temporary processing.
  • cbm_store_open_path (lines 98-102): Creates a persistent file-backed database that survives process termination and supports inter-process sharing at the cost of disk latency.

Choose in-memory storage for CI pipelines or ephemeral analysis, and file-backed storage for long-lived indexer daemons.

Bulk Write Optimizations

Large ingestion phases—such as indexing a new repository—benefit from three specific optimizations exposed in the API.

Pragma Tuning

The cbm_store_begin_bulk() function (lines 44-46) disables synchronous mode and enlarges the SQLite cache to minimize disk fsyncs during mass insertion. Conversely, cbm_store_end_bulk() (lines 48-49) restores normal durability pragmas after the bulk phase completes.

Always wrap bulk operations in a transaction using cbm_store_begin() (lines 31-40) and cbm_store_commit() to reduce per-statement overhead.

Index Management

User-defined indexes on project, label, and file_path columns accelerate queries but severely slow inserts. The store provides cbm_store_drop_indexes() (lines 51-55) to temporarily remove indexes during ingestion, followed by cbm_store_create_indexes() (lines 51-55) to rebuild them afterward. This pattern prevents index maintenance overhead during bulk loads.

WAL Checkpointing

Write-Ahead Logging (WAL) enables concurrent readers while a writer appends to the log, but the log file grows without bound. The cbm_store_checkpoint() function (lines 57-60) forces a PRAGMA wal_checkpoint and runs PRAGMA optimize to compact the database file and reclaim disk space.

Memory-Mapped I/O Configuration

The store supports memory-mapped I/O via the CBM_SQLITE_MMAP_SIZE environment variable, allowing reads to be served directly from the kernel's page cache. The cbm_store_resolve_mmap_size() function (lines 62-68) reads this variable and defaults to 64 MiB.

Setting this value too low reduces the benefit; setting it too high risks address space exhaustion on 32-bit platforms. Negative values clamp to 0, disabling mmap to avoid SIGBUS errors during concurrent truncation.

Transaction Granularity and Batch APIs

Small transactions (single INSERT per call) force expensive disk syncs. Group writes inside a single transaction boundary using cbm_store_begin(), cbm_store_commit(), and cbm_store_rollback() (lines 31-40).

For even greater efficiency, use batch CRUD APIs that leverage prepared statements internally:

  • cbm_store_upsert_node_batch() (lines 88-91): Bulk upsert nodes
  • cbm_store_insert_edge_batch() (lines 56-58): Bulk insert edges
  • cbm_store_upsert_file_hash_batch() (lines 93-95): Bulk upsert file hashes

These functions reduce parsing overhead and enable SQLite's native bulk-insert optimizations.

Query Performance and Pagination

Functions returning arrays (cbm_store_find_nodes_by_*, cbm_store_list_files, etc.) allocate memory for each result. For massive codebases, unbounded result sets consume excessive RAM. The cbm_search_params_t struct (lines 15-17) provides limit and offset fields to paginate results and keep working sets bounded.

Database Integrity and Maintenance

Corruption can silently degrade performance by creating dead pages. The cbm_store_check_integrity() function (lines 14-18) returns false on detectable corruption and should be invoked after unexpected shutdowns before heavy operations.

Summary

  • Never share a cbm_store_t * between threads; create one handle per thread or use external locking.
  • Bulk ingestion follows this sequence: cbm_store_begin_bulk() → cbm_store_drop_indexes() → batch upserts → cbm_store_create_indexes() → cbm_store_end_bulk() → cbm_store_commit().
  • Configure CBM_SQLITE_MMAP_SIZE to match available RAM for large databases; the default 64 MiB suits most workloads.
  • Checkpoint periodically using cbm_store_checkpoint() to prevent unbounded WAL growth.
  • Paginate large queries using limit and offset in cbm_search_params_t to control memory usage.

Frequently Asked Questions

Can I share a single store handle between multiple threads?

No. As implemented in src/store/store.h, store handles are not thread-safe. You must create one handle per thread using cbm_store_open_path() or cbm_store_open_memory(), or protect shared access with a mutex. Concurrent access to a single cbm_store_t * causes race conditions and crashes.

How do I optimize bulk imports for large repositories?

Wrap the import in cbm_store_begin_bulk() to disable synchronous mode and enlarge the cache, then call cbm_store_drop_indexes() to remove indexes temporarily. Use batch APIs like cbm_store_upsert_node_batch() for the actual insertion. Afterward, restore indexes with cbm_store_create_indexes() and finalize with cbm_store_end_bulk() followed by cbm_store_commit().

What is the impact of WAL mode on graph storage performance?

WAL mode enables concurrent readers while writes proceed, improving throughput under mixed workloads. However, the WAL file grows continuously until checkpointed. call cbm_store_checkpoint() periodically to rewrite the database file and prevent excessive disk usage from an ever-growing log.

When should I use in-memory storage versus file-backed storage?

Use cbm_store_open_memory() for ephemeral testing, CI pipelines, or temporary analysis where persistence is unnecessary. Use cbm_store_open_path() for production indexers, long-running daemons, or when data must survive process restarts and be shared across processes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →