Team Artifact Export Import Process in Codebase Memory MCP: A Complete Technical Guide

The team artifact export import process in Codebase Memory MCP enables developers to share compressed SQLite graph database snapshots via atomic file operations, storing metadata in JSON and using Zstandard compression to ensure reproducible repository indexing across teams.

The DeusData/codebase-memory-mcp repository implements a persistent artifact system that allows teams to distribute pre-indexed codebase knowledge. This mechanism captures the entire graph database state, including discovery and extraction results, into portable files that teammates can import without re-running the full indexing pipeline. The implementation resides primarily in src/pipeline/artifact.c and integrates with the pipeline orchestration logic in src/pipeline/pipeline.c.

What Is the Team Artifact Feature?

The persistent artifact feature is the mechanism by which Codebase Memory MCP shares fully-compressed snapshots of indexed repositories. When enabled, the system exports the SQLite graph database (graph.db) into a portable format after the indexing pipeline completes, allowing other team members to import the exact same graph state without executing the discovery and extraction phases locally.

Where Artifacts Are Stored

Storage Layout and File Types

Artifacts reside under <repo>/.codebase-memory/ as two distinct files:

  • artifact.zst: A Zstandard-compressed copy of the SQLite graph database.
  • artifact.json: Human-readable metadata containing schema version, Git commit hash, node/edge counts, and UTC timestamps.

Both files are written atomically using temporary files followed by rename operations to prevent partial writes during system crashes or power failures. According to the source in artifact.c【artifact.c†L52-L62】, the implementation uses write_file_atomic which performs a rename on Unix systems or MoveFileEx on Windows to ensure data consistency.

Export Process: How Codebase Memory MCP Creates Artifacts

Pipeline Integration

The export process begins when persistence is enabled on the pipeline. The client code sets p->persistence = true via cbm_pipeline_set_persistence (defined in pipeline.h). After the graph dump succeeds, the function dump_and_persist_hashes in src/pipeline/pipeline.c invokes cbm_artifact_export using the best available compression quality【pipeline.c†L889-L902】.

Step-by-Step Export Implementation

The cbm_artifact_export function in src/pipeline/artifact.c executes the following steps:

  1. Argument Validation: Checks for null pointers and returns uniform errors via artifact_export_fail, which records messages in a thread-local buffer (g_export_error) under the artifact.export namespace【artifact.c†L68-L84】.

  2. Directory Creation: Calls cbm_mkdir_p to ensure .codebase-memory/ exists before writing.

  3. Database Preparation: For "best" quality exports, prepare_stripped_db drops indexes and vacuums the database into a temporary copy before compression to minimize size.

  4. Compression: Uses cbm_zstd_compress with level 3 (fast) or level 9 (best) depending on the quality parameter.

  5. Atomic File Write: Writes to artifact.zst via write_file_atomic, creating a temporary file first then renaming it to prevent corruption.

  6. Metadata Generation: write_metadata executes git rev-parse HEAD (shell-safe) to capture the commit hash, records a UTC timestamp, and writes artifact.json.

  7. Git Attributes: ensure_gitattributes adds merge=ours to prevent merge conflicts on the compressed binary file.

Import Process: How Codebase Memory MCP Restores Artifacts

Import Validation and Steps

The cbm_artifact_import(repo_path, cache_db_path) function (invoked by mcp --import) restores artifacts through the following sequence:

  1. Metadata Verification: read_metadata_version validates schema compatibility, aborting if the stored version exceeds the binary's current version【artifact.c†L90-L96】.

  2. Load Compressed Data: read_file_alloc loads artifact.zst into memory.

  3. Decompression: cbm_zstd_decompress allocates the destination buffer using the original size recorded in metadata.

  4. Atomic Database Write: Writes the decompressed SQLite file to <cache_db_path>.import_tmp then renames it atomically.

  5. Integrity Check: cbm_store_check_integrity opens the database to verify indexes and FTS5 tables are sane.

  6. Cleanup: Removes WAL/SHM leftovers to ensure a clean final database state.

Any failure logs detailed errors under the artifact.import namespace.

Safety Mechanisms and Atomic Operations

Shell Argument Validation: The function cbm_artifact_repo_path_is_shell_safe mirrors Git command hardening, rejecting paths containing quotes, semicolons, backticks, or Windows-specific meta-characters【artifact.c†L16-L27】.

Atomic File Operations: Both export and import use write_file_atomic to avoid leaving half-written files on crash or power loss, implementing the write-to-temp-then-rename pattern observed in artifact.c【artifact.c†L52-L63】.

Schema Versioning: The JSON metadata includes a schema_version field. The import process aborts if the stored version is newer than the running binary's version, protecting against forward-incompatible artifacts.

Merge Conflict Prevention: The system automatically creates .gitattributes entries with merge=ours for artifact files, ensuring that distributed artifacts never conflict during Git merges.

Code Examples

Enable persistence and run the full pipeline:

cbm_pipeline_t *p = cbm_pipeline_new(repo_path, NULL, CBM_MODE_FULL);
if (!p) { /* handle error */ }
cbm_pipeline_set_persistence(p, true);          /* turn on artifact export */
int rc = cbm_pipeline_run(p);                   /* runs discovery, extraction, etc. */
if (rc != 0) {
    const char *err = cbm_artifact_export_last_error();
    fprintf(stderr, "pipeline failed: %s\n", err ? err : "unknown");
}
cbm_pipeline_free(p);

Perform a direct export without running the pipeline:

int rc = cbm_artifact_export(db_path, repo_path, "my‑project", CBM_ARTIFACT_BEST);
if (rc != 0) {
    fprintf(stderr, "export error: %s\n", cbm_artifact_export_last_error());
}

Import an artifact into a fresh cache location:

int rc = cbm_artifact_import(repo_path, "/tmp/my‑project.db");
if (rc != 0) {
    fprintf(stderr, "import failed\n");
}

Check for existing artifacts before re-exporting:

if (cbm_artifact_exists(repo_path)) {
    printf("artifact already present – skip export\n");
}

Summary

  • Atomic Operations: Both export and import rely on write_file_atomic in src/pipeline/artifact.c to prevent data corruption during system failures.
  • Compression Strategy: Uses Zstandard (ZSTD) with configurable levels 3 (fast) and 9 (best) to balance size and speed.
  • Schema Protection: Version metadata prevents incompatible artifact imports, ensuring binary compatibility across team members.
  • Security: Shell-safe path validation prevents injection attacks when executing git rev-parse HEAD and other commands.
  • Git Integration: Automatic .gitattributes configuration eliminates merge conflicts on binary artifact files.

Frequently Asked Questions

What files constitute a team artifact in Codebase Memory MCP?

A team artifact consists of two files stored in the .codebase-memory/ directory: artifact.zst (the Zstandard-compressed SQLite database) and artifact.json (metadata including schema version, Git commit hash, and node counts). Together these files capture the complete indexed state of a repository.

How does the export process ensure data integrity during system crashes?

The export process uses atomic file writes via write_file_atomic, which writes to a temporary file first and then performs a rename operation (or MoveFileEx on Windows). This ensures that artifact.zst and artifact.json never exist in a partially-written state, as the rename operation is atomic at the filesystem level.

Can I import an artifact created with a newer version of the MCP binary?

No, the import process explicitly blocks forward-incompatible artifacts. The read_metadata_version function in src/pipeline/artifact.c compares the stored schema version against the binary's current version and aborts if the stored version is newer, preventing undefined behavior from schema mismatches.

What compression algorithm does the team artifact process use?

The system uses Zstandard (ZSTD) compression via cbm_zstd_compress, supporting two quality levels: level 3 for fast compression and level 9 for best compression. The "best" quality mode also drops database indexes and vacuums the SQLite file before compression to maximize space savings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →