How Graph Data Is Persisted and Retrieved in Codebase-Memory-MCP
Codebase-Memory-MCP persists graph data by dumping an in-memory SQLite database to disk, compressing it with Zstandard, and retrieving it by decompressing the artifact on startup to enable fast initialization without re-parsing.
The DeusData/codebase-memory-mcp repository implements a Model Context Protocol (MCP) server that constructs a knowledge graph from source code repositories. Understanding how this tool persists and retrieves graph data is essential for teams sharing graph artifacts across development environments and optimizing server startup times.
Architecture Overview
The persistence strategy follows a memory-first, disk-backup pattern. During indexing, the graph lives entirely in an in-memory SQLite connection for maximum performance. Once parsing completes, the system performs a durable backup to disk, compresses the result, and stores it in a team-shareable location.
This approach balances ACID compliance with portability, allowing the compressed graph.db.zst file to be committed to version control or shared via cloud storage.
In-Memory Graph Construction
When you initiate repository indexing, the pipeline activates Tree-sitter parsers to extract definitions, function calls, imports, and other code relationships. Rather than writing incrementally to disk, the system accumulates all nodes and edges in an in-memory SQLite database.
This phase utilizes WAL (Write-Ahead Log) mode on the SQLite connection, ensuring that even during the in-memory phase, the database maintains ACID-safe transaction properties. The low-level inserts are handled by the storage layer, potentially implemented in src/store/sqlite_writer.c if present in the build, though the primary orchestration occurs in src/store/store.c.
Dumping and Compressing the Database
The store_dump Function
After the full repository pass completes, the system triggers the persistence routine. In src/store/store.c, the function store_dump executes an sqlite3_backup operation to copy the entire in-memory database to a file-backed SQLite database named graph.db.
This single backup operation ensures consistency, capturing the complete graph state at the moment of indexing completion.
Compression and Artifact Storage
Following the dump, the system applies two optimization steps:
- Compaction: Executes
VACUUM INTOto strip auxiliary indexes and defragment the database, minimizing file size. - Compression: Applies Zstandard (zstd) compression to produce
graph.db.zst.
The resulting artifact is placed in the .codebase-memory/ directory adjacent to your source tree. According to the Team-Shared Graph Artifact section in README.md, this compressed file can be committed to version control or shared among team members, eliminating the need for each developer to re-index large codebases.
Retrieving the Graph on Startup
When the MCP server launches and detects an existing graph.db.zst file in the project’s .codebase-memory/ directory, it automatically decompresses the artifact and restores the SQLite file before any incremental indexing begins.
This retrieval mechanism allows the server to bootstrap from durable storage in seconds rather than re-parsing the entire codebase. The decompressed database is opened in WAL mode, establishing a persistent connection that remains active for the server’s lifetime.
Querying the Persisted Graph
Once loaded, all MCP tools execute SQL queries directly against the persisted SQLite file. The graph is always read from durable storage, with WAL ensuring that any modifications made by incremental indexing flush to disk automatically.
Common operations include:
# Index a repository and create the persistence artifact
codebase-memory-mcp cli index_repository '{"repo_path":"/path/to/project"}'
# Verify the graph schema against the persisted database
codebase-memory-mcp cli get_graph_schema
# Execute Cypher-like queries against stored graph data
codebase-memory-mcp cli query_graph '{"query":"MATCH (f:Function)-[:CALLS]->(g) WHERE f.name=\"main\" RETURN g.name"}'
# Remove persisted data for a specific project
codebase-memory-mcp cli delete_project '{"project":"myproject"}'
Configuration and Storage Locations
By default, the SQLite database and compressed artifacts are stored under the user’s cache directory at ~/.cache/codebase-memory-mcp/. You can override this location by setting the CBM_CACHE_DIR environment variable, as documented in docs/CONFIGURATION.md.
The separation between the global cache directory (for temporary files) and the project-local .codebase-memory/ directory (for shareable artifacts) allows flexibility in deployment strategies.
Summary
- In-memory indexing builds the graph in a high-performance SQLite connection before persisting to disk.
- The
store_dumpfunction insrc/store/store.cusessqlite3_backupto create a durablegraph.dbfile. - Compression via Zstandard creates
graph.db.zstin.codebase-memory/, producing a team-shareable artifact. - Retrieval decompresses the artifact on startup, enabling fast initialization without re-parsing source files.
- WAL mode ensures ACID compliance and automatic durability for all graph modifications during runtime.
Frequently Asked Questions
What file format does Codebase-Memory-MCP use to store graph data?
The system uses SQLite as the underlying database format, compressed with Zstandard (zstd) to create a .zst artifact. This combines the query flexibility of SQL with efficient storage and fast decompression rates.
Where is the graph database stored by default?
The default location is ~/.cache/codebase-memory-mcp/ for the active SQLite file, while the compressed graph.db.zst artifact is stored in the .codebase-memory/ directory within your project root. You can customize the cache location using the CBM_CACHE_DIR environment variable.
How does the system ensure data durability during indexing?
The system operates in WAL (Write-Ahead Log) mode, which provides ACID-safe durability guarantees. When the in-memory database is dumped to disk via sqlite3_backup in the store_dump function, the operation creates a complete, consistent snapshot of the graph at that moment.
Can team members share the persisted graph without re-indexing?
Yes. The graph.db.zst file in the .codebase-memory/ directory is designed as a team-shared graph artifact. Developers can commit this compressed file to version control or distribute it via cloud storage, allowing other team members to skip the initial indexing phase entirely by decompressing the artifact on their first run.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →