# How Graph Data Is Persisted and Retrieved in Codebase-Memory-MCP

> Learn how Codebase-Memory-MCP persists graph data via SQLite disk dumps compressed with Zstandard for rapid retrieval. See how it decompressess on startup for fast initialization.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: internals
- Published: 2026-07-05

---

**Codebase-Memory-MCP persists graph data by dumping an in-memory SQLite database to disk, compressing it with Zstandard, and retrieving it by decompressing the artifact on startup to enable fast initialization without re-parsing.**

The `DeusData/codebase-memory-mcp` repository implements a Model Context Protocol (MCP) server that constructs a knowledge graph from source code repositories. Understanding how this tool persists and retrieves graph data is essential for teams sharing graph artifacts across development environments and optimizing server startup times.

## Architecture Overview

The persistence strategy follows a **memory-first, disk-backup** pattern. During indexing, the graph lives entirely in an in-memory SQLite connection for maximum performance. Once parsing completes, the system performs a durable backup to disk, compresses the result, and stores it in a team-shareable location.

This approach balances **ACID compliance** with **portability**, allowing the compressed `graph.db.zst` file to be committed to version control or shared via cloud storage.

## In-Memory Graph Construction

When you initiate repository indexing, the pipeline activates Tree-sitter parsers to extract definitions, function calls, imports, and other code relationships. Rather than writing incrementally to disk, the system accumulates all nodes and edges in an in-memory SQLite database.

This phase utilizes **WAL (Write-Ahead Log) mode** on the SQLite connection, ensuring that even during the in-memory phase, the database maintains ACID-safe transaction properties. The low-level inserts are handled by the storage layer, potentially implemented in [`src/store/sqlite_writer.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/sqlite_writer.c) if present in the build, though the primary orchestration occurs in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c).

## Dumping and Compressing the Database

### The store_dump Function

After the full repository pass completes, the system triggers the persistence routine. In [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c), the function `store_dump` executes an `sqlite3_backup` operation to copy the entire in-memory database to a file-backed SQLite database named `graph.db`.

This single backup operation ensures consistency, capturing the complete graph state at the moment of indexing completion.

### Compression and Artifact Storage

Following the dump, the system applies two optimization steps:

1. **Compaction**: Executes `VACUUM INTO` to strip auxiliary indexes and defragment the database, minimizing file size.
2. **Compression**: Applies **Zstandard (zstd)** compression to produce `graph.db.zst`.

The resulting artifact is placed in the `.codebase-memory/` directory adjacent to your source tree. According to the **Team-Shared Graph Artifact** section in [`README.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/README.md), this compressed file can be committed to version control or shared among team members, eliminating the need for each developer to re-index large codebases.

## Retrieving the Graph on Startup

When the MCP server launches and detects an existing `graph.db.zst` file in the project’s `.codebase-memory/` directory, it automatically decompresses the artifact and restores the SQLite file before any incremental indexing begins.

This retrieval mechanism allows the server to bootstrap from durable storage in seconds rather than re-parsing the entire codebase. The decompressed database is opened in WAL mode, establishing a persistent connection that remains active for the server’s lifetime.

## Querying the Persisted Graph

Once loaded, all MCP tools execute SQL queries directly against the persisted SQLite file. The graph is always read from durable storage, with WAL ensuring that any modifications made by incremental indexing flush to disk automatically.

Common operations include:

```bash

# Index a repository and create the persistence artifact

codebase-memory-mcp cli index_repository '{"repo_path":"/path/to/project"}'

# Verify the graph schema against the persisted database

codebase-memory-mcp cli get_graph_schema

# Execute Cypher-like queries against stored graph data

codebase-memory-mcp cli query_graph '{"query":"MATCH (f:Function)-[:CALLS]->(g) WHERE f.name=\"main\" RETURN g.name"}'

# Remove persisted data for a specific project

codebase-memory-mcp cli delete_project '{"project":"myproject"}'

```

## Configuration and Storage Locations

By default, the SQLite database and compressed artifacts are stored under the user’s cache directory at `~/.cache/codebase-memory-mcp/`. You can override this location by setting the `CBM_CACHE_DIR` environment variable, as documented in [`docs/CONFIGURATION.md`](https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/CONFIGURATION.md).

The separation between the global cache directory (for temporary files) and the project-local `.codebase-memory/` directory (for shareable artifacts) allows flexibility in deployment strategies.

## Summary

- **In-memory indexing** builds the graph in a high-performance SQLite connection before persisting to disk.
- **The `store_dump` function** in [`src/store/store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.c) uses `sqlite3_backup` to create a durable `graph.db` file.
- **Compression** via Zstandard creates `graph.db.zst` in `.codebase-memory/`, producing a team-shareable artifact.
- **Retrieval** decompresses the artifact on startup, enabling fast initialization without re-parsing source files.
- **WAL mode** ensures ACID compliance and automatic durability for all graph modifications during runtime.

## Frequently Asked Questions

### What file format does Codebase-Memory-MCP use to store graph data?

The system uses **SQLite** as the underlying database format, compressed with **Zstandard (zstd)** to create a `.zst` artifact. This combines the query flexibility of SQL with efficient storage and fast decompression rates.

### Where is the graph database stored by default?

The default location is `~/.cache/codebase-memory-mcp/` for the active SQLite file, while the compressed `graph.db.zst` artifact is stored in the `.codebase-memory/` directory within your project root. You can customize the cache location using the `CBM_CACHE_DIR` environment variable.

### How does the system ensure data durability during indexing?

The system operates in **WAL (Write-Ahead Log) mode**, which provides ACID-safe durability guarantees. When the in-memory database is dumped to disk via `sqlite3_backup` in the `store_dump` function, the operation creates a complete, consistent snapshot of the graph at that moment.

### Can team members share the persisted graph without re-indexing?

Yes. The `graph.db.zst` file in the `.codebase-memory/` directory is designed as a **team-shared graph artifact**. Developers can commit this compressed file to version control or distribute it via cloud storage, allowing other team members to skip the initial indexing phase entirely by decompressing the artifact on their first run.