# How the Team‑Shared Graph Artifact with Zstd Compression Works in Codebase Memory

> Discover how the team-shared graph artifact with Zstd compression works in Codebase Memory for instant, version-controlled synchronization across development teams using Git.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: deep-dive
- Published: 2026-07-08

---

**The team-shared graph artifact packages a complete SQLite knowledge graph as a Zstandard-compressed `.zst` file with a JSON sidecar, enabling version-controlled, instant synchronization across development teams via Git.**

Codebase Memory MCP solves the cold-start problem of knowledge graphs by serializing the entire database into a **team-shared graph artifact with zstd compression** that lives inside your repository. According to the DeusData/codebase-memory-mcp source code, this implementation uses thin C wrappers around the Zstandard library to compress the SQLite store while preserving metadata integrity and preventing merge conflicts.

## Export Pipeline: Creating the Compressed Artifact

### Validation and Directory Setup

The process begins in [`src/pipeline/artifact.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/artifact.c) where `cbm_artifact_export` validates inputs and creates the `.codebase-memory/` directory with mode `0755` via `prepare_artifact_dir`.

### Database Preparation and Index Stripping

Depending on the quality flag—`CBM_ARTIFACT_FAST` or `CBM_ARTIFACT_BEST`—the system either reads the DB directly or creates a stripped copy. The `prepare_stripped_db` function (lines 335-370) uses `VACUUM INTO` to create a temporary copy, then executes `DROP INDEX` on all user-created indexes followed by another `VACUUM` to maximize compression ratio.

### Zstandard Compression and Atomic Writes

The actual compression happens in `cbm_zstd_compress` (defined in [`internal/cbm/zstd_store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/zstd_store.c)), a thin wrapper around `ZSTD_compress` using level 3 for fast exports or level 9 for best quality. The buffer size is pre-computed with `cbm_zstd_compress_bound`. The compressed data is written atomically using `write_file_atomic` (lines 59-109), which writes to a temporary file and renames it to avoid partial artifacts.

### Metadata Generation and Git Integration

After compression, `write_metadata` (lines 330-372) generates a JSON sidecar containing:

- `schema_version` and Git commit hash
- Node and edge counts
- Original and compressed size metrics
- Compression level used

Finally, `ensure_gitattributes` creates a `.gitattributes` file with `merge=ours binary` to prevent Git from attempting to merge the binary artifact during conflicts.

## Import Pipeline: Reconstructing the Graph Securely

### Metadata Validation and Size Verification

The import process in `cbm_artifact_import` first reads [`artifact.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/artifact.json) to check `schema_version` compatibility. Crucially, it ignores the `original_size` field in metadata for memory allocation decisions.

### Secure Decompression with Frame Header Validation

Instead of trusting the JSON metadata, the code calls `cbm_zstd_frame_content_size` (lines 124-133 in [`artifact.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/artifact.c)) to extract the true decompressed size from the Zstandard frame header. This value is validated against `ART_MAX_DECOMPRESSED_BYTES` (64 GB) and cross-checked with the stored `original_size` before calling `cbm_zstd_decompress`.

### Integrity Verification and Cache Population

The decompressed database is written atomically to the cache path, then opened with `cbm_store_check_integrity` to verify SQLite integrity. The temporary file is renamed to the final location only after successful validation.

## Zstandard Wrapper Implementation

The compression layer resides in [`internal/cbm/zstd_store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/zstd_store.c) and [`zstd_store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/zstd_store.h):

- **`cbm_zstd_compress`** – Wraps `ZSTD_compress` with configurable levels (3 for fast, 9 for best).
- **`cbm_zstd_decompress`** – Wraps `ZSTD_decompress` and returns the decompressed byte count.
- **`cbm_zstd_frame_content_size`** – Reads the content-size field from the ZSTD frame header for secure buffer allocation.
- **`cbm_zstd_compress_bound`** – Returns the maximum possible compressed size for buffer allocation.

## Practical Code Examples

### Exporting a Graph with Fast Compression

```c
#include "src/pipeline/artifact.h"

int rc = cbm_artifact_export(
    "/path/to/project.db",           // db_path
    "/path/to/git/repo",             // repo_path  
    "my-project",                    // project_name
    CBM_ARTIFACT_FAST               // quality flag (level 3)
);

if (rc != 0) {
    fprintf(stderr, "Export failed: %s\n", 
            cbm_artifact_export_last_error());
}

```

### Importing a Shared Artifact

```c
int rc = cbm_artifact_import(
    "/path/to/git/repo",             // repo_path
    "/home/user/.cache/codebase-memory/my-project.db"  // cache_db_path
);

if (rc != 0) {
    fprintf(stderr, "Import failed\n");
}

```

### Retrieving the Source Commit Hash

```c
char *commit = cbm_artifact_commit("/path/to/git/repo");
if (commit) {
    printf("Artifact built from commit %s\n", commit);
    free(commit);
}

```

## Summary

- The **team-shared graph artifact with zstd compression** consists of a `.zst` file and JSON metadata stored in `.codebase-memory/` for version-controlled distribution.
- Export strips database indexes when using `CBM_ARTIFACT_BEST` to achieve higher compression ratios, while `CBM_ARTIFACT_FAST` prioritizes speed.
- All file writes use atomic rename operations to prevent corruption, and Git attributes are configured to treat artifacts as binary.
- Import derives decompression buffer sizes from the Zstandard frame header rather than JSON metadata, mitigating memory exhaustion attacks.
- The wrapper functions in [`internal/cbm/zstd_store.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/zstd_store.c) provide a thin, secure interface to the Zstandard library with configurable compression levels.

## Frequently Asked Questions

### How does the team-shared graph artifact prevent merge conflicts in Git?

The system automatically generates a `.gitattributes` file inside `.codebase-memory/` with the directive `merge=ours binary`. This instructs Git to keep the local version during merges and prevents diff algorithms from processing the binary Zstandard file, eliminating merge conflicts on the artifact itself.

### What is the difference between CBM_ARTIFACT_FAST and CBM_ARTIFACT_BEST?

`CBM_ARTIFACT_FAST` reads the SQLite database directly and compresses with Zstandard level 3, prioritizing export speed. `CBM_ARTIFACT_BEST` first creates a stripped copy using `VACUUM INTO`, removes all user-created indexes via `DROP INDEX`, runs `VACUUM` again, and compresses with level 9, yielding significantly smaller files at the cost of processing time.

### Why does the import process ignore the original_size field in the JSON metadata?

The import routine in [`src/pipeline/artifact.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/artifact.c) extracts the true decompressed size from the Zstandard frame header using `cbm_zstd_frame_content_size` rather than trusting the JSON metadata. This security measure prevents buffer overflow or memory exhaustion attacks that could occur if a malicious actor tampered with the `original_size` field to request excessive memory allocation.

### What happens if the compressed artifact is corrupted during transfer?

During import, after decompression, the code calls `cbm_store_check_integrity` on the resulting SQLite database. If the integrity check fails—which would detect corruption from transfer or storage errors—the import aborts before the file is moved to the final cache location, ensuring only valid databases are used for querying.