Team-Shared Graph Artifact Format and Bootstrap Process for codebase-memory-mcp
The team-shared graph artifact in codebase-memory-mcp is a zstd-compressed SQLite database stored at .codebase-memory/graph.db.zst that enables instant team onboarding by automatically decompressing and importing the snapshot before running incremental indexing only on changed files.
The DeusData/codebase-memory-mcp repository implements a sophisticated knowledge-graph persistence mechanism designed to eliminate cold-start indexing costs for development teams. Understanding the team-shared graph artifact format and bootstrap process is essential for teams looking to share pre-built code intelligence across multiple developer environments without requiring full re-indexing on every clone.
Artifact Format and Storage Location
The artifact lives at .codebase-memory/graph.db.zst within your repository root, sitting alongside your source tree. This file represents a zstd-compressed snapshot of the entire knowledge graph, wrapping a standard SQLite database (internally named graph.db) in a high-efficiency compression stream.
Zstd-Compressed SQLite Structure
The format combines SQLite durability with zstd compression speed to create a portable, versionable database. When decompressed, the artifact reveals a SQLite file containing node tables (e.g., Function, Class, Resource) and edge tables (CALLS, IMPORTS, SEMANTICALLY_RELATED). This structure allows the system to maintain referential integrity while keeping the artifact size small enough for Git storage.
Schema Overview
According to the repository's README (lines 203-210), the database schema stores semantic relationships between code entities. Nodes represent definable symbols like functions and classes, while edges capture relationships including call graphs, import dependencies, and semantic correlations. The zstd wrapper ensures that decompression remains fast enough to serve as a bootstrap mechanism without adding significant latency to the indexing workflow.
Bootstrap Process and Incremental Indexing
When a developer runs the index_repository command, the binary performs a detection sequence to determine whether to bootstrap from the artifact or start from scratch.
Detection and Decompression
The tool first checks for the existence of a local graph.db in ~/.cache/codebase-memory-mcp/. If no local database exists but .codebase-memory/graph.db.zst is present, the system automatically decompresses the zstd stream to reconstruct the SQLite database. This operation occurs transparently before any indexing begins, as implemented in internal/cbm/cbm.c.
Import and Indexing Pipeline
After decompression, the tool imports the artifact and immediately transitions to incremental indexing rather than full repository analysis. The pipeline performs tree-sitter parsing, Hybrid LSP type resolution, and edge creation only for files that have changed since the artifact was created. This approach dramatically reduces startup costs, as the bulk of the codebase knowledge is already present in the imported snapshot.
Configuration and Git Integration
The bootstrap process includes Git-aware configuration to prevent merge conflicts. On first export, the tool creates a .gitattributes entry specifying merge=ours for the artifact file. This ensures that when multiple developers modify the graph independently, Git preserves the local version rather than attempting a binary merge, eliminating collision scenarios for the zstd-compressed database.
Implementation Details
The core logic resides in internal/cbm/cbm.c with API declarations in internal/cbm/cbm.h. These files implement the index_repository and get_graph_schema functions that orchestrate the bootstrap sequence. The install.sh script ensures proper placement of the artifact during tool installation, while the CLI commands defined in the header file provide the interface for repository analysis.
# Clone a repository containing the graph artifact
git clone https://github.com/DeusData/codebase-memory-mcp.git
cd codebase-memory-mcp
# Bootstrap automatically: decompresses artifact then indexes only changes
codebase-memory-mcp index_repository
# Verify the imported graph structure
codebase-memory-mcp get_graph_schema
# Explicit project targeting with optional name parameter
codebase-memory-mcp index_repository --project=my-project
Summary
- The team-shared graph artifact uses zstd-compressed SQLite stored at
.codebase-memory/graph.db.zstto persist knowledge graphs. - The bootstrap process automatically decompresses and imports the artifact when no local
graph.dbexists in the cache directory. - Incremental indexing runs only on changed files after import, avoiding full re-index costs.
- Automatic
.gitattributesconfiguration withmerge=oursprevents merge conflicts for the binary artifact. - Implementation in
internal/cbm/cbm.chandles the decompression, import, and incremental analysis pipeline.
Frequently Asked Questions
What is the exact file format of the team-shared graph artifact?
The artifact is a zstd-compressed SQLite database containing tables for nodes (such as Function, Class, and Resource) and edges (including CALLS, IMPORTS, and SEMANTICALLY_RELATED). The zstd compression allows for efficient storage in Git while maintaining fast decompression speeds during the bootstrap process.
How does the bootstrap process avoid re-indexing the entire codebase?
The index_repository command checks for an existing local graph.db in ~/.cache/codebase-memory-mcp/. When the artifact is present but no local database exists, it decompresses .codebase-memory/graph.db.zst to import the full graph, then executes the incremental indexing pipeline (tree-sitter parsing through Hybrid LSP resolution) only on files modified since the artifact was created.
Where is the graph artifact stored and how is it version controlled?
The artifact resides at .codebase-memory/graph.db.zst in your repository root. The tool automatically configures .gitattributes with a merge=ours rule during the first export, ensuring Git handles the binary file correctly and prevents merge conflicts when multiple developers update the graph independently.
Which source files implement the bootstrap and import logic?
The core bootstrap logic is implemented in internal/cbm/cbm.c, with API declarations in internal/cbm/cbm.h. These files define the index_repository function that handles artifact detection, zstd decompression, SQLite import, and the subsequent incremental indexing workflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →