How to Share Team Graph Artifacts Across Repositories with codebase-memory-mcp

Teams can share pre-indexed code knowledge graphs by committing a ZSTD-compressed SQLite snapshot (.codebase-memory/graph.db.zst) to version control, allowing teammates to bootstrap instantly without re-indexing.

codebase-memory-mcp is a static, single-binary knowledge-graph engine that enables teams to share team graph artifacts across repositories via MCP (Model Context Protocol). The tool indexes your codebase into a rich graph of symbols and exposes it through 14 JSON-RPC tools, eliminating the need for teammates to re-index large codebases on every clone.

Understanding the Graph Artifact Format

The Team-Shared Graph Artifact is a compressed SQLite database that persists the complete symbol graph of your repository. When you run codebase-memory-mcp index_repository, the tool writes a snapshot to .codebase-memory/graph.db.zst that can be checked into Git.

The artifact uses ZSTD compression to achieve an 8–13:1 size reduction over raw SQLite. It implements a two-tiered storage strategy: a "best" high-ratio version for full indexing portability and a "fast" low-ratio version for incremental watcher updates. To prevent merge conflicts, the tool automatically configures .gitattributes with merge=ours for the binary file.

Repository Architecture and Key Components

The DeusData/codebase-memory-mcp repository organizes functionality into three distinct layers:

  • Discovery & Indexing (src/discover/, src/pipeline/) – Walks the file tree respecting .gitignore and .cbmignore, parses files using vendored Tree-Sitter grammars, and resolves imports via the Hybrid LSP engine.
  • Graph Store & Query Engine (src/store/, src/cypher/) – Maintains an in-memory SQLite instance that dumps to the compressed artifact format. Parses Cypher-like read-only queries through a custom lexer.
  • MCP Server & UI (src/mcp/, src/ui/) – Exposes 14 tools including search_graph, trace_path, and get_architecture via JSON-RPC, plus an optional embedded 3-D visualization server.

The Hybrid LSP layer runs inside the static binary to provide type-aware resolution for Python, TypeScript/JSX, PHP, C#, Go, C/C++, Java, Kotlin, and Rust—no external language servers required.

Workflow for Sharing Graph Artifacts Across Teams

Implement this four-step workflow to distribute pre-computed graphs:

  1. Index and Export – A maintainer runs the indexing command to generate .codebase-memory/graph.db.zst.
  2. Commit to Repository – The compressed artifact is added to version control with appropriate Git attributes.
  3. Teammate Bootstrap – When developers clone the repository, the binary detects the artifact, inflates it, and applies only incremental diffs for files changed since the snapshot.
  4. Continuous Updates – A background watcher monitors Git changes and writes low-ratio incremental updates to keep the shared artifact fresh.

This workflow allows a full index of the Linux kernel (28 M LOC) to complete in approximately 3 minutes, while typical projects under 100 K LOC index in seconds.

CLI Commands for Graph Management

Execute these commands to manage and query shared graph artifacts:


# Index the current repository and create the shared artifact

codebase-memory-mcp cli index_repository '{"repo_path": "."}'

# Output: .codebase-memory/graph.db.zst

# Commit the artifact for team sharing

git add .codebase-memory/graph.db.zst
git commit -m "Add shared codegraph snapshot"

# Search for function symbols matching a pattern

codebase-memory-mcp cli search_graph '{
  "label": "Function",
  "name_pattern": ".*Handler.*",
  "limit": 10
}'

# Trace call paths for a specific function

codebase-memory-mcp cli trace_path '{
  "function_name": "Search",
  "direction": "both",
  "depth": 3
}'

# Execute Cypher-like queries against the graph

codebase-memory-mcp cli query_graph '{
  "query": "MATCH (f:Function)-[:CALLS]->(g) WHERE f.name = \"ProcessOrder\" RETURN g.name"
}'

Implementation Details and Source Files

The core functionality resides in specific source directories:

  • src/main.c – Entry point handling both CLI and MCP server modes.
  • src/mcp/ – Implements the 14 MCP tools including search_graph and trace_path.
  • src/store/ – SQLite-backed graph storage with dump/restore logic for the ZSTD format.
  • src/pipeline/ – Multi-pass indexing pipeline integrating Tree-Sitter parsing with Hybrid LSP resolution.
  • src/ui/ – Embedded HTTP server serving the 3-D visualization interface.

The graph artifact persistence logic in src/store/ handles the two-tier compression strategy and automatic .gitattributes configuration to ensure merge-friendly behavior.

Summary

  • ZSTD-compressed SQLite snapshots (.codebase-memory/graph.db.zst) enable portable, pre-indexed knowledge graphs.
  • Two-tier storage provides both high-compression archives and fast incremental updates.
  • Automatic Git configuration prevents binary merge conflicts via .gitattributes rules.
  • Hybrid LSP integration delivers type-aware symbol resolution for 9 languages without external dependencies.
  • 14 MCP tools expose graph querying capabilities to agents like Claude, Gemini, and VS Code.

Frequently Asked Questions

How do teammates avoid re-indexing when they clone a repository?

When a teammate clones a repository containing a committed graph.db.zst artifact, the codebase-memory-mcp binary automatically detects the snapshot, decompresses it, and loads the graph into memory. It then performs an incremental scan only for files changed since the snapshot was created, skipping the full re-index process and saving significant time.

What prevents merge conflicts in the binary graph file?

The tool automatically adds a merge=ours entry to .gitattributes for the .codebase-memory/graph.db.zst file. This Git configuration ensures that when concurrent edits occur, Git always keeps the local version during merges, preventing binary conflict resolution headaches while allowing the background watcher to regenerate the artifact locally.

Can the graph artifact be used across different repositories?

Yes, the same graph format supports cross-repository intelligence. While each repository typically maintains its own artifact, the standardized ZSTD-compressed SQLite format enables linking symbols across services and generating fleet-wide architecture views when multiple repositories are indexed.

What is the performance impact of using pre-indexed graphs?

Structural queries against the pre-indexed graph consume approximately 3,400 tokens compared to 412,000 tokens for naive file-by-file analysis, representing a 99% reduction in LLM context window usage. Query response times remain under 1 millisecond for symbol lookups, while the initial bootstrap from a shared artifact reduces indexing time from minutes to seconds.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →