Performance Benchmarks for code-review-graph Graph Building

The code-review-graph repository provides a built-in benchmarking suite in code_review_graph/eval/benchmarks/ that quantifies graph construction throughput via the build_performance benchmark, measuring flow detection latency, community detection time, and average search query speed.

The code-review-graph project ships with dedicated performance benchmarks for code-review-graph graph building to ensure the knowledge graph construction pipeline scales efficiently with codebase size. These benchmarks capture precise timing data for the three-stage transformation process that converts raw source code into a queryable graph structure. Developers can execute these measurements through the evaluation runner to generate reproducible performance reports and detect latency regressions.

Core Graph Building Benchmarks

The build_performance Implementation

The primary benchmark for graph construction is build_performance, implemented in code_review_graph/eval/benchmarks/build_performance.py (lines 12‑60). This benchmark instruments the three critical phases of graph generation:

  • Flow detection: Measures the time required to traverse the store and discover call-flow edges using trace_flows(store), followed by persistence via store_flows(store, flows). Timing is captured using time.perf_counter().
  • Community detection: Times the clustering phase that identifies related node groups through detect_communities(store), persisting results with store_communities(store, comms).
  • Search latency: Executes a configurable set of search queries (default: first 10) against store.search_nodes(sq["query"], limit=20) and calculates average query duration.

Metrics and Return Values

The build_performance benchmark returns a dictionary containing quantitative performance indicators:

  • flow_detection_seconds: Time spent identifying and storing code flows
  • community_detection_seconds: Duration of community clustering operations
  • search_avg_ms: Average query response time in milliseconds
  • nodes_per_second: Processing throughput indicating scalability
  • stats.files_count, stats.total_nodes, stats.total_edges: Repository scale metrics
  • config["name"]: Identifier for the evaluated repository

These metrics enable direct comparison of graph construction efficiency across different repository sizes and hardware configurations.

Running the Benchmark Suite

Execute all graph building benchmarks using the evaluation runner:

python -m code_review_graph.eval.runner \
    --repo-root /path/to/your/repo \
    --store /path/to/graph/store \
    --benchmarks build_performance,token_efficiency,search_quality

The runner (code_review_graph/eval/runner.py) orchestrates the process by:

  1. Loading the graph store via store.get_stats() to capture baseline statistics
  2. Executing selected benchmarks (e.g., build_performance.run(...))
  3. Aggregating results into a list of dictionaries
  4. Passing data to code_review_graph/eval/reporter.py for markdown report generation

Complementary Performance Evaluations

In addition to core construction timing, the repository includes benchmarks that indirectly measure graph-building efficiency:

Token Efficiency Analysis

The token_efficiency benchmark in code_review_graph/eval/benchmarks/token_benchmark.py measures total token consumption for agent workflows versus naive file-reading baselines. Key outputs include the token-reduction ratio and total tokens saved, quantifying how graph construction impacts downstream AI agent costs.

Search Quality Metrics

The search_quality benchmark in code_review_graph/eval/benchmarks/search_quality.py evaluates retrieval accuracy using Mean Reciprocal Rank (MRR). While focused on query results, this metric validates that the graph construction process produces a search index that improves information retrieval accuracy.

Summary

  • The build_performance benchmark in code_review_graph/eval/benchmarks/build_performance.py provides the primary measurement for performance benchmarks for code-review-graph graph building.
  • Three critical phases are instrumented: flow detection (trace_flows), community detection (detect_communities), and search latency (store.search_nodes).
  • Key metrics include nodes_per_second, flow_detection_seconds, and search_avg_ms, enabling scalability analysis.
  • The evaluation runner (code_review_graph/eval/runner.py) coordinates benchmark execution and report generation via code_review_graph/eval/reporter.py.
  • Complementary benchmarks measure token efficiency and search quality (MRR) to validate the cost-effectiveness and accuracy of the constructed graph.

Frequently Asked Questions

Which file contains the main graph building performance benchmark?

The primary benchmark resides in code_review_graph/eval/benchmarks/build_performance.py, specifically implementing the build_performance function between lines 12 and 60. This file contains the timing logic for flow detection, community detection, and search latency measurements using Python's time.perf_counter().

How is the benchmark execution orchestrated?

Benchmark execution is coordinated by code_review_graph/eval/runner.py, which loads the graph store, invokes each selected benchmark's run() method, and aggregates results. The runner accepts command-line arguments specifying the repository root, store path, and comma-separated benchmark names, then delegates final reporting to code_review_graph/eval/reporter.py.

What does the nodes_per_second metric indicate?

The nodes_per_second metric measures graph construction throughput by dividing the total node count (stats.total_nodes) by the combined processing time. This indicator helps developers assess scalability and predict construction time for larger codebases, with higher values indicating more efficient graph building performance.

How do token efficiency benchmarks relate to graph construction?

The token_efficiency benchmark in code_review_graph/eval/benchmarks/token_benchmark.py quantifies the token cost savings achieved by using the constructed graph versus reading entire files. This validates that the graph building overhead is justified by reduced downstream agent token consumption, providing a cost-benefit analysis of the construction pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →