How the Team‑Shared Graph Artifact with Zstd Compression Works in Codebase Memory
The team-shared graph artifact packages a complete SQLite knowledge graph as a Zstandard-compressed .zst file with a JSON sidecar, enabling version-controlled, instant synchronization across development teams via Git.
Codebase Memory MCP solves the cold-start problem of knowledge graphs by serializing the entire database into a team-shared graph artifact with zstd compression that lives inside your repository. According to the DeusData/codebase-memory-mcp source code, this implementation uses thin C wrappers around the Zstandard library to compress the SQLite store while preserving metadata integrity and preventing merge conflicts.
Export Pipeline: Creating the Compressed Artifact
Validation and Directory Setup
The process begins in src/pipeline/artifact.c where cbm_artifact_export validates inputs and creates the .codebase-memory/ directory with mode 0755 via prepare_artifact_dir.
Database Preparation and Index Stripping
Depending on the quality flag—CBM_ARTIFACT_FAST or CBM_ARTIFACT_BEST—the system either reads the DB directly or creates a stripped copy. The prepare_stripped_db function (lines 335-370) uses VACUUM INTO to create a temporary copy, then executes DROP INDEX on all user-created indexes followed by another VACUUM to maximize compression ratio.
Zstandard Compression and Atomic Writes
The actual compression happens in cbm_zstd_compress (defined in internal/cbm/zstd_store.c), a thin wrapper around ZSTD_compress using level 3 for fast exports or level 9 for best quality. The buffer size is pre-computed with cbm_zstd_compress_bound. The compressed data is written atomically using write_file_atomic (lines 59-109), which writes to a temporary file and renames it to avoid partial artifacts.
Metadata Generation and Git Integration
After compression, write_metadata (lines 330-372) generates a JSON sidecar containing:
schema_versionand Git commit hash- Node and edge counts
- Original and compressed size metrics
- Compression level used
Finally, ensure_gitattributes creates a .gitattributes file with merge=ours binary to prevent Git from attempting to merge the binary artifact during conflicts.
Import Pipeline: Reconstructing the Graph Securely
Metadata Validation and Size Verification
The import process in cbm_artifact_import first reads artifact.json to check schema_version compatibility. Crucially, it ignores the original_size field in metadata for memory allocation decisions.
Secure Decompression with Frame Header Validation
Instead of trusting the JSON metadata, the code calls cbm_zstd_frame_content_size (lines 124-133 in artifact.c) to extract the true decompressed size from the Zstandard frame header. This value is validated against ART_MAX_DECOMPRESSED_BYTES (64 GB) and cross-checked with the stored original_size before calling cbm_zstd_decompress.
Integrity Verification and Cache Population
The decompressed database is written atomically to the cache path, then opened with cbm_store_check_integrity to verify SQLite integrity. The temporary file is renamed to the final location only after successful validation.
Zstandard Wrapper Implementation
The compression layer resides in internal/cbm/zstd_store.c and zstd_store.h:
cbm_zstd_compress– WrapsZSTD_compresswith configurable levels (3 for fast, 9 for best).cbm_zstd_decompress– WrapsZSTD_decompressand returns the decompressed byte count.cbm_zstd_frame_content_size– Reads the content-size field from the ZSTD frame header for secure buffer allocation.cbm_zstd_compress_bound– Returns the maximum possible compressed size for buffer allocation.
Practical Code Examples
Exporting a Graph with Fast Compression
#include "src/pipeline/artifact.h"
int rc = cbm_artifact_export(
"/path/to/project.db", // db_path
"/path/to/git/repo", // repo_path
"my-project", // project_name
CBM_ARTIFACT_FAST // quality flag (level 3)
);
if (rc != 0) {
fprintf(stderr, "Export failed: %s\n",
cbm_artifact_export_last_error());
}
Importing a Shared Artifact
int rc = cbm_artifact_import(
"/path/to/git/repo", // repo_path
"/home/user/.cache/codebase-memory/my-project.db" // cache_db_path
);
if (rc != 0) {
fprintf(stderr, "Import failed\n");
}
Retrieving the Source Commit Hash
char *commit = cbm_artifact_commit("/path/to/git/repo");
if (commit) {
printf("Artifact built from commit %s\n", commit);
free(commit);
}
Summary
- The team-shared graph artifact with zstd compression consists of a
.zstfile and JSON metadata stored in.codebase-memory/for version-controlled distribution. - Export strips database indexes when using
CBM_ARTIFACT_BESTto achieve higher compression ratios, whileCBM_ARTIFACT_FASTprioritizes speed. - All file writes use atomic rename operations to prevent corruption, and Git attributes are configured to treat artifacts as binary.
- Import derives decompression buffer sizes from the Zstandard frame header rather than JSON metadata, mitigating memory exhaustion attacks.
- The wrapper functions in
internal/cbm/zstd_store.cprovide a thin, secure interface to the Zstandard library with configurable compression levels.
Frequently Asked Questions
How does the team-shared graph artifact prevent merge conflicts in Git?
The system automatically generates a .gitattributes file inside .codebase-memory/ with the directive merge=ours binary. This instructs Git to keep the local version during merges and prevents diff algorithms from processing the binary Zstandard file, eliminating merge conflicts on the artifact itself.
What is the difference between CBM_ARTIFACT_FAST and CBM_ARTIFACT_BEST?
CBM_ARTIFACT_FAST reads the SQLite database directly and compresses with Zstandard level 3, prioritizing export speed. CBM_ARTIFACT_BEST first creates a stripped copy using VACUUM INTO, removes all user-created indexes via DROP INDEX, runs VACUUM again, and compresses with level 9, yielding significantly smaller files at the cost of processing time.
Why does the import process ignore the original_size field in the JSON metadata?
The import routine in src/pipeline/artifact.c extracts the true decompressed size from the Zstandard frame header using cbm_zstd_frame_content_size rather than trusting the JSON metadata. This security measure prevents buffer overflow or memory exhaustion attacks that could occur if a malicious actor tampered with the original_size field to request excessive memory allocation.
What happens if the compressed artifact is corrupted during transfer?
During import, after decompression, the code calls cbm_store_check_integrity on the resulting SQLite database. If the integrity check fails—which would detect corruption from transfer or storage errors—the import aborts before the file is moved to the final cache location, ensuring only valid databases are used for querying.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →