How Graph Data Is Persisted and Retrieved in Codebase-Memory-MCP

Codebase-Memory-MCP persists graph data by dumping an in-memory SQLite database to disk, compressing it with Zstandard, and retrieving it by decompressing the artifact on startup to enable fast initialization without re-parsing.

The DeusData/codebase-memory-mcp repository implements a Model Context Protocol (MCP) server that constructs a knowledge graph from source code repositories. Understanding how this tool persists and retrieves graph data is essential for teams sharing graph artifacts across development environments and optimizing server startup times.

Architecture Overview

The persistence strategy follows a memory-first, disk-backup pattern. During indexing, the graph lives entirely in an in-memory SQLite connection for maximum performance. Once parsing completes, the system performs a durable backup to disk, compresses the result, and stores it in a team-shareable location.

This approach balances ACID compliance with portability, allowing the compressed graph.db.zst file to be committed to version control or shared via cloud storage.

In-Memory Graph Construction

When you initiate repository indexing, the pipeline activates Tree-sitter parsers to extract definitions, function calls, imports, and other code relationships. Rather than writing incrementally to disk, the system accumulates all nodes and edges in an in-memory SQLite database.

This phase utilizes WAL (Write-Ahead Log) mode on the SQLite connection, ensuring that even during the in-memory phase, the database maintains ACID-safe transaction properties. The low-level inserts are handled by the storage layer, potentially implemented in src/store/sqlite_writer.c if present in the build, though the primary orchestration occurs in src/store/store.c.

Dumping and Compressing the Database

The store_dump Function

After the full repository pass completes, the system triggers the persistence routine. In src/store/store.c, the function store_dump executes an sqlite3_backup operation to copy the entire in-memory database to a file-backed SQLite database named graph.db.

This single backup operation ensures consistency, capturing the complete graph state at the moment of indexing completion.

Compression and Artifact Storage

Following the dump, the system applies two optimization steps:

  1. Compaction: Executes VACUUM INTO to strip auxiliary indexes and defragment the database, minimizing file size.
  2. Compression: Applies Zstandard (zstd) compression to produce graph.db.zst.

The resulting artifact is placed in the .codebase-memory/ directory adjacent to your source tree. According to the Team-Shared Graph Artifact section in README.md, this compressed file can be committed to version control or shared among team members, eliminating the need for each developer to re-index large codebases.

Retrieving the Graph on Startup

When the MCP server launches and detects an existing graph.db.zst file in the project’s .codebase-memory/ directory, it automatically decompresses the artifact and restores the SQLite file before any incremental indexing begins.

This retrieval mechanism allows the server to bootstrap from durable storage in seconds rather than re-parsing the entire codebase. The decompressed database is opened in WAL mode, establishing a persistent connection that remains active for the server’s lifetime.

Querying the Persisted Graph

Once loaded, all MCP tools execute SQL queries directly against the persisted SQLite file. The graph is always read from durable storage, with WAL ensuring that any modifications made by incremental indexing flush to disk automatically.

Common operations include:


# Index a repository and create the persistence artifact

codebase-memory-mcp cli index_repository '{"repo_path":"/path/to/project"}'

# Verify the graph schema against the persisted database

codebase-memory-mcp cli get_graph_schema

# Execute Cypher-like queries against stored graph data

codebase-memory-mcp cli query_graph '{"query":"MATCH (f:Function)-[:CALLS]->(g) WHERE f.name=\"main\" RETURN g.name"}'

# Remove persisted data for a specific project

codebase-memory-mcp cli delete_project '{"project":"myproject"}'

Configuration and Storage Locations

By default, the SQLite database and compressed artifacts are stored under the user’s cache directory at ~/.cache/codebase-memory-mcp/. You can override this location by setting the CBM_CACHE_DIR environment variable, as documented in docs/CONFIGURATION.md.

The separation between the global cache directory (for temporary files) and the project-local .codebase-memory/ directory (for shareable artifacts) allows flexibility in deployment strategies.

Summary

  • In-memory indexing builds the graph in a high-performance SQLite connection before persisting to disk.
  • The store_dump function in src/store/store.c uses sqlite3_backup to create a durable graph.db file.
  • Compression via Zstandard creates graph.db.zst in .codebase-memory/, producing a team-shareable artifact.
  • Retrieval decompresses the artifact on startup, enabling fast initialization without re-parsing source files.
  • WAL mode ensures ACID compliance and automatic durability for all graph modifications during runtime.

Frequently Asked Questions

What file format does Codebase-Memory-MCP use to store graph data?

The system uses SQLite as the underlying database format, compressed with Zstandard (zstd) to create a .zst artifact. This combines the query flexibility of SQL with efficient storage and fast decompression rates.

Where is the graph database stored by default?

The default location is ~/.cache/codebase-memory-mcp/ for the active SQLite file, while the compressed graph.db.zst artifact is stored in the .codebase-memory/ directory within your project root. You can customize the cache location using the CBM_CACHE_DIR environment variable.

How does the system ensure data durability during indexing?

The system operates in WAL (Write-Ahead Log) mode, which provides ACID-safe durability guarantees. When the in-memory database is dumped to disk via sqlite3_backup in the store_dump function, the operation creates a complete, consistent snapshot of the graph at that moment.

Can team members share the persisted graph without re-indexing?

Yes. The graph.db.zst file in the .codebase-memory/ directory is designed as a team-shared graph artifact. Developers can commit this compressed file to version control or distribute it via cloud storage, allowing other team members to skip the initial indexing phase entirely by decompressing the artifact on their first run.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →