How to Use the Codebase-Memory-MCP Graph Database: A Complete Guide
Codebase-Memory-MCP builds a persistent knowledge graph from your source code and exposes 14 MCP tools—including search_graph, trace_path, and query_graph—to query the graph via JSON-RPC commands.
Codebase-Memory-MCP (CBM) is an open-source tool developed by DeusData that transforms any repository into a queryable knowledge graph. By parsing your codebase with tree-sitter grammars and enriching the AST with Hybrid LSP type resolution, it creates an immutable SQLite database at ~/.cache/codebase-memory-mcp/ that represents functions, classes, imports, and call relationships as nodes and edges.
Indexing Your Repository
Before querying, you must index your codebase. This process parses every file using vendored tree-sitter grammars, extracts definitions and symbols, and refines the raw AST with type-aware edges in internal/cbm/cbm.c.
Run the indexing command once per project:
codebase-memory-mcp cli index_repository '{"repo_path":"/path/to/your/project"}'
This creates a SQLite database at ~/.cache/codebase-memory-mcp/<project>.db and registers the project in the internal catalog. The storage layer, implemented in store/store.c, writes the graph as immutable nodes and edges. After initial creation, only incremental updates from the file watcher modify the database.
Querying the Graph Database
CBM exposes 14 MCP tools that return JSON-RPC payloads on stdout. These tools enable structured search, path tracing, and arbitrary Cypher queries against the graph.
Inspecting the Schema with get_graph_schema
Always inspect the schema first to discover available node labels and edge types:
codebase-memory-mcp cli get_graph_schema
The output lists labels such as Function, Class, and Route, alongside edge types like CALLS, IMPORTS, and HTTP_CALLS.
Structured Search with search_graph
Use search_graph for filtered lookups by label, name regex, degree limits, or file scope:
codebase-memory-mcp cli search_graph '{"label":"Function","name_pattern":"^process_.*$","limit":10}'
Results include qualified names (project.module.file.function) that serve as inputs for other tools.
Path Tracing with trace_path
Trace call relationships breadth-first using trace_path. Specify direction as inbound for callers or outbound for callees:
# Find functions that call process_order
codebase-memory-mcp cli trace_path '{"function_name":"process_order","direction":"inbound","depth":5}'
# Find functions called by process_order
codebase-memory-mcp cli trace_path '{"function_name":"process_order","direction":"outbound","depth":3}'
Custom Cypher Queries with query_graph
For complex analysis, use query_graph with a read-only OpenCypher subset:
codebase-memory-mcp cli query_graph '{
"query": "MATCH (f:Function)-[:CALLS]->(g) WHERE f.name = \"process_order\" RETURN g.name"
}'
The supported Cypher syntax allows MATCH, WHERE, and RETURN clauses to traverse arbitrary paths.
Practical Code Examples
These ready-to-run commands demonstrate common workflows. All examples assume the binary is on your PATH.
Index and Verify Schema
# Create the graph database
codebase-memory-mcp cli index_repository '{"repo_path":"/home/user/myapp"}'
# Verify available labels and properties
codebase-memory-mcp cli get_graph_schema | jq .
Map HTTP Routes to Handler Functions
codebase-memory-mcp cli search_graph '{
"label":"Route",
"limit":100
}' | jq -r '.results[].qualified_name' |
while read route; do
echo "Route: $route"
codebase-memory-mcp cli query_graph "{
\"query\": \"MATCH (r:Route {qualified_name: '$route'})-[:HANDLES]->(f:Function) RETURN f.qualified_name\"
}" | jq -r '.results[].f.qualified_name'
done
Detect Dead Code
Find functions with no incoming CALLS edges:
codebase-memory-mcp cli query_graph '{
"query": "MATCH (f:Function) WHERE NOT EXISTS { (f)<-[:CALLS]-() } RETURN f.qualified_name"
}' | jq -r '.results[].f.qualified_name'
Trace Third-Party Dependencies
Identify all functions calling a specific package:
codebase-memory-mcp cli query_graph '{
"query": "MATCH (f:Function)-[:CALLS]->(p:Package {name:\"requests\"}) RETURN f.qualified_name"
}' | jq -r '.results[].f.qualified_name'
Launch the 3-D Visualizer
If you installed the ui variant, start the interactive visualizer:
codebase-memory-mcp --ui=true --port=9749 &
open http://localhost:9749
The UI connects to the SQLite backend in store/store.c and renders the graph interactively.
Advanced Features
Hybrid LSP provides type-aware call resolution during indexing. It is built-in for supported languages like Python, Go, and TypeScript, and refines edges in internal/cbm/extract_calls.c.
Cross-repo edges (CROSS_* types) connect symbols across multiple repositories. Index multiple projects under the same cache directory to enable automatic linking.
Auto-indexing watches for git changes and re-indexes automatically. Enable it with:
codebase-memory-mcp config set auto_index true
Semantic search uses bundled Nomic embeddings via the semantic_query tool for natural language code search.
Community detection runs the Louvain algorithm through the detect_communities tool (implemented in store/community.c) to identify functional modules within the graph.
Summary
- Indexing creates an immutable SQLite database at
~/.cache/codebase-memory-mcp/using tree-sitter parsing and Hybrid LSP resolution as defined ininternal/cbm/cbm.c. - Querying uses 14 MCP tools including
search_graphfor filtered search,trace_pathfor call graph traversal, andquery_graphfor OpenCypher queries. - Storage is handled in
store/store.c, providing persistent nodes and edges that support complex graph algorithms. - Visualization is available via the UI variant on port 9749 for interactive exploration.
Frequently Asked Questions
What file formats does Codebase-Memory-MCP support?
The tool uses vendored tree-sitter grammars to parse any language supported by tree-sitter, including Python, Go, TypeScript, JavaScript, Rust, and C. The Hybrid LSP layer provides enhanced type resolution for specific languages as documented in the repository's Hybrid LSP section.
Where is the graph database stored?
The SQLite database is stored at ~/.cache/codebase-memory-mcp/<project>.db on your local filesystem. This location is immutable after indexing, except for incremental updates triggered by the file watcher when auto_index is enabled.
Can I query the database without using the MCP tools?
While the MCP tools are the primary interface, the underlying storage is a standard SQLite file. However, CBM is designed to be queried through its JSON-RPC interface using tools like query_graph, which enforces read-only access and validates Cypher syntax against the supported subset.
How does the graph handle cross-repository dependencies?
When you index multiple repositories under the same cache directory, CBM automatically creates CROSS_* edge types linking symbols across projects. This enables tracing calls from your application code into library dependencies or across microservice boundaries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →