How to Use the Codebase-Memory-MCP Graph Database: A Complete Guide

Codebase-Memory-MCP builds a persistent knowledge graph from your source code and exposes 14 MCP tools—including search_graph, trace_path, and query_graph—to query the graph via JSON-RPC commands.

Codebase-Memory-MCP (CBM) is an open-source tool developed by DeusData that transforms any repository into a queryable knowledge graph. By parsing your codebase with tree-sitter grammars and enriching the AST with Hybrid LSP type resolution, it creates an immutable SQLite database at ~/.cache/codebase-memory-mcp/ that represents functions, classes, imports, and call relationships as nodes and edges.

Indexing Your Repository

Before querying, you must index your codebase. This process parses every file using vendored tree-sitter grammars, extracts definitions and symbols, and refines the raw AST with type-aware edges in internal/cbm/cbm.c.

Run the indexing command once per project:

codebase-memory-mcp cli index_repository '{"repo_path":"/path/to/your/project"}'

This creates a SQLite database at ~/.cache/codebase-memory-mcp/<project>.db and registers the project in the internal catalog. The storage layer, implemented in store/store.c, writes the graph as immutable nodes and edges. After initial creation, only incremental updates from the file watcher modify the database.

Querying the Graph Database

CBM exposes 14 MCP tools that return JSON-RPC payloads on stdout. These tools enable structured search, path tracing, and arbitrary Cypher queries against the graph.

Inspecting the Schema with get_graph_schema

Always inspect the schema first to discover available node labels and edge types:

codebase-memory-mcp cli get_graph_schema

The output lists labels such as Function, Class, and Route, alongside edge types like CALLS, IMPORTS, and HTTP_CALLS.

Structured Search with search_graph

Use search_graph for filtered lookups by label, name regex, degree limits, or file scope:

codebase-memory-mcp cli search_graph '{"label":"Function","name_pattern":"^process_.*$","limit":10}'

Results include qualified names (project.module.file.function) that serve as inputs for other tools.

Path Tracing with trace_path

Trace call relationships breadth-first using trace_path. Specify direction as inbound for callers or outbound for callees:


# Find functions that call process_order

codebase-memory-mcp cli trace_path '{"function_name":"process_order","direction":"inbound","depth":5}'

# Find functions called by process_order

codebase-memory-mcp cli trace_path '{"function_name":"process_order","direction":"outbound","depth":3}'

Custom Cypher Queries with query_graph

For complex analysis, use query_graph with a read-only OpenCypher subset:

codebase-memory-mcp cli query_graph '{
  "query": "MATCH (f:Function)-[:CALLS]->(g) WHERE f.name = \"process_order\" RETURN g.name"
}'

The supported Cypher syntax allows MATCH, WHERE, and RETURN clauses to traverse arbitrary paths.

Practical Code Examples

These ready-to-run commands demonstrate common workflows. All examples assume the binary is on your PATH.

Index and Verify Schema


# Create the graph database

codebase-memory-mcp cli index_repository '{"repo_path":"/home/user/myapp"}'

# Verify available labels and properties

codebase-memory-mcp cli get_graph_schema | jq .

Map HTTP Routes to Handler Functions

codebase-memory-mcp cli search_graph '{
  "label":"Route",
  "limit":100
}' | jq -r '.results[].qualified_name' |
while read route; do
  echo "Route: $route"
  codebase-memory-mcp cli query_graph "{
    \"query\": \"MATCH (r:Route {qualified_name: '$route'})-[:HANDLES]->(f:Function) RETURN f.qualified_name\"
  }" | jq -r '.results[].f.qualified_name'
done

Detect Dead Code

Find functions with no incoming CALLS edges:

codebase-memory-mcp cli query_graph '{
  "query": "MATCH (f:Function) WHERE NOT EXISTS { (f)<-[:CALLS]-() } RETURN f.qualified_name"
}' | jq -r '.results[].f.qualified_name'

Trace Third-Party Dependencies

Identify all functions calling a specific package:

codebase-memory-mcp cli query_graph '{
  "query": "MATCH (f:Function)-[:CALLS]->(p:Package {name:\"requests\"}) RETURN f.qualified_name"
}' | jq -r '.results[].f.qualified_name'

Launch the 3-D Visualizer

If you installed the ui variant, start the interactive visualizer:

codebase-memory-mcp --ui=true --port=9749 &
open http://localhost:9749

The UI connects to the SQLite backend in store/store.c and renders the graph interactively.

Advanced Features

Hybrid LSP provides type-aware call resolution during indexing. It is built-in for supported languages like Python, Go, and TypeScript, and refines edges in internal/cbm/extract_calls.c.

Cross-repo edges (CROSS_* types) connect symbols across multiple repositories. Index multiple projects under the same cache directory to enable automatic linking.

Auto-indexing watches for git changes and re-indexes automatically. Enable it with:

codebase-memory-mcp config set auto_index true

Semantic search uses bundled Nomic embeddings via the semantic_query tool for natural language code search.

Community detection runs the Louvain algorithm through the detect_communities tool (implemented in store/community.c) to identify functional modules within the graph.

Summary

  • Indexing creates an immutable SQLite database at ~/.cache/codebase-memory-mcp/ using tree-sitter parsing and Hybrid LSP resolution as defined in internal/cbm/cbm.c.
  • Querying uses 14 MCP tools including search_graph for filtered search, trace_path for call graph traversal, and query_graph for OpenCypher queries.
  • Storage is handled in store/store.c, providing persistent nodes and edges that support complex graph algorithms.
  • Visualization is available via the UI variant on port 9749 for interactive exploration.

Frequently Asked Questions

What file formats does Codebase-Memory-MCP support?

The tool uses vendored tree-sitter grammars to parse any language supported by tree-sitter, including Python, Go, TypeScript, JavaScript, Rust, and C. The Hybrid LSP layer provides enhanced type resolution for specific languages as documented in the repository's Hybrid LSP section.

Where is the graph database stored?

The SQLite database is stored at ~/.cache/codebase-memory-mcp/<project>.db on your local filesystem. This location is immutable after indexing, except for incremental updates triggered by the file watcher when auto_index is enabled.

Can I query the database without using the MCP tools?

While the MCP tools are the primary interface, the underlying storage is a standard SQLite file. However, CBM is designed to be queried through its JSON-RPC interface using tools like query_graph, which enforces read-only access and validates Cypher syntax against the supported subset.

How does the graph handle cross-repository dependencies?

When you index multiple repositories under the same cache directory, CBM automatically creates CROSS_* edge types linking symbols across projects. This enables tracing calls from your application code into library dependencies or across microservice boundaries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →