Performance Benchmarks for Graph Operations in Codebase Memory MCP: Sub‑Second Speed at 50K+ Nodes

The Codebase Memory MCP server executes graph indexing, Cypher traversals, and deep call‑path tracing in sub‑second times for repositories approaching 50,000 nodes, scaling linearly to handle 196,022 edges without timeouts.

The DeusData/codebase-memory-mcp repository ships a comprehensive benchmark suite that measures the speed and scalability of core graph operations across dozens of real‑world codebases. These performance benchmarks for graph operations demonstrate that even heavily‑connected symbol graphs—such as the Linux kernel subset with approximately 20,000 nodes—complete complex multi‑hop queries without hitting time‑outs. All measurements capture resource counters (RSS, CPU time) and execute on an Apple M3 Pro (macOS Darwin 25.3.0).

Benchmark Methodology and Hardware Configuration

The evaluation methodology is documented in docs/EVALUATION_PLAN.md, which defines how the system measures latency and throughput for indexing, search, and traversal workloads. Tests run on an Apple M3 Pro (macOS Darwin 25.3.0) and record resident set size (RSS) and CPU time for every operation. The Language Benchmark reports in docs/BENCHMARK.md provide representative results across Python, Go, Rust, and C/C++ codebases, ensuring the graph engine handles diverse language semantics without degradation.

Graph Indexing Performance and Scale Metrics

The indexing phase constructs the full graph representation from source code, creating nodes for symbols and edges for relationships like CALLS or IMPORTS. According to the benchmark data in docs/BENCHMARK.md, the system successfully indexed the Django repository—generating 49,398 nodes and 196,022 edges—without any indexing failures. Other massive codebases, including Laravel (≈38,000 nodes) and neovim (≈24,000 nodes), processed with no performance issues, demonstrating that internal/cbm/graph.go implements graph data structures capable of handling enterprise‑scale dependency graphs.

Cypher Query Execution Speed

Query performance stays sub‑second for typical traversals, even on densely connected graphs. Query 9 measures a CALLS edge‑level traversal using Cypher syntax:

codebase-memory-mcp query_graph --project myproj --query '{"query":"MATCH (a)-[r:CALLS]->(b) RETURN a.name,b.name LIMIT 20"}'

This query pattern completed successfully against a Linux kernel driver subset containing roughly 20,000 nodes, returning full call graphs without triggering the default 30‑second timeout. The core traversal logic in internal/cbm/graph.go ensures that edge‑level lookups scale linearly with graph size.

Deep Call-Path Tracing Benchmarks

Multi‑hop call‑path tracing exercises the graph engine’s ability to resolve long dependency chains. The benchmark suite tests inbound and outbound directions at depth 5, measuring how the system handles heavily‑connected symbols.

A stress test on the kernel function i40e_probe produced a result set of 129,026 characters (approximately 130 KB) containing the complete call chain. This deep trace finished without timeout, confirming that the trace_call_path command—mapped through pkg/pypi/src/codebase_memory_mcp/_cli.py—can return massive result sets while maintaining responsiveness.

codebase-memory-mcp trace_call_path --project myproj --symbol i40e_probe --direction outbound --max-depth 5

Stress Testing on Production Codebases

To validate extreme scale, the benchmark includes a kernel driver subset comprising approximately 19,000 nodes and 67,000 edges. All queries in the suite finished successfully, with a 3‑hop CALLS chain producing 10 valid chains without timeouts. Heavy‑weight extraction queries, such as those in scripts/extract_nomic_vectors.py that generate embeddings for every symbol in the graph, further confirm that the engine handles read‑intensive workloads without memory pressure or latency spikes.

Reproducing Benchmarks Locally

You can replicate these performance benchmarks for graph operations using the following CLI invocations. Each command emits a .ndjson trajectory file for inspection and can be timed with standard shell utilities:


# Index a repository (creates nodes and edges)

time codebase-memory-mcp index /path/to/repo --project myproj

# Schema inspection (counts nodes and relationships)

time codebase-memory-mcp query_graph --project myproj --query '{"query":"CALL db.schema"}'

# Symbol discovery with regex filtering

time codebase-memory-mcp search_graph --project myproj --label Function --filter "name =~ '(?i)render'"

For repositories under 50,000 nodes, these operations complete in a few hundred milliseconds. The Python entry point in pkg/pypi/src/codebase_memory_mcp/_cli.py marshals these commands to the Go‑based graph engine implemented in internal/cbm/graph.go, ensuring consistent performance across language bindings.

Summary

  • Sub‑second latency for typical Cypher queries and symbol lookups on graphs under 50,000 nodes.
  • Linear scalability demonstrated up to 196,022 edges (Django benchmark) without indexing failures.
  • Deep tracing resilience at depth 5 produces 130 KB result sets well under the 30‑second timeout threshold.
  • Stress‑test validation on Linux kernel subsets (≈20,000 nodes) and massive Python codebases confirms production readiness.
  • Hardware baseline established on Apple M3 Pro with resource monitoring via RSS and CPU time counters.

Frequently Asked Questions

What hardware configuration was used for the official performance benchmarks?

The benchmark suite executed on an Apple M3 Pro running macOS Darwin 25.3.0. The test harness monitors resident set size (RSS) and CPU time for every graph operation to ensure accurate latency and memory footprint measurements.

How large a codebase can the graph engine index before performance degrades?

The system successfully indexed Django with 49,398 nodes and 196,022 edges, and processed Laravel (≈38,000 nodes) and neovim (≈24,000 nodes) without timeouts or memory issues. Performance remains linear up to tens of thousands of nodes, though the benchmark demonstrates stability even at hundreds of thousands of edges.

Where is the core graph traversal logic implemented?

The graph data structures and traversal algorithms reside in internal/cbm/graph.go, written in Go. The Python CLI entry point at pkg/pypi/src/codebase_memory_mcp/_cli.py marshals user commands to this core engine, which handles Cypher query execution and call‑path tracing.

What is the timeout threshold for complex graph queries?

The default timeout is 30 seconds. Even deep operations—such as 3‑hop CALLS chains on kernel driver subsets or depth‑5 inbound traces producing 130 KB results—complete well under this limit, typically in sub‑second time for moderately sized repositories.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →