# Performance Benchmarks for code-review-graph Graph Building

> Discover code-review-graph graph building performance benchmarks. Measure throughput, flow detection, community detection, and query speed with our built-in suite.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: performance
- Published: 2026-08-13

---

**The code-review-graph repository provides a built-in benchmarking suite in `code_review_graph/eval/benchmarks/` that quantifies graph construction throughput via the `build_performance` benchmark, measuring flow detection latency, community detection time, and average search query speed.**

The `code-review-graph` project ships with dedicated **performance benchmarks for code-review-graph graph building** to ensure the knowledge graph construction pipeline scales efficiently with codebase size. These benchmarks capture precise timing data for the three-stage transformation process that converts raw source code into a queryable graph structure. Developers can execute these measurements through the evaluation runner to generate reproducible performance reports and detect latency regressions.

## Core Graph Building Benchmarks

### The build_performance Implementation

The primary benchmark for graph construction is **`build_performance`**, implemented in [`code_review_graph/eval/benchmarks/build_performance.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/build_performance.py) (lines 12‑60). This benchmark instruments the three critical phases of graph generation:

- **Flow detection**: Measures the time required to traverse the store and discover call-flow edges using `trace_flows(store)`, followed by persistence via `store_flows(store, flows)`. Timing is captured using `time.perf_counter()`.
- **Community detection**: Times the clustering phase that identifies related node groups through `detect_communities(store)`, persisting results with `store_communities(store, comms)`.
- **Search latency**: Executes a configurable set of search queries (default: first 10) against `store.search_nodes(sq["query"], limit=20)` and calculates average query duration.

### Metrics and Return Values

The `build_performance` benchmark returns a dictionary containing quantitative performance indicators:

- `flow_detection_seconds`: Time spent identifying and storing code flows
- `community_detection_seconds`: Duration of community clustering operations
- `search_avg_ms`: Average query response time in milliseconds
- `nodes_per_second`: Processing throughput indicating scalability
- `stats.files_count`, `stats.total_nodes`, `stats.total_edges`: Repository scale metrics
- `config["name"]`: Identifier for the evaluated repository

These metrics enable direct comparison of graph construction efficiency across different repository sizes and hardware configurations.

## Running the Benchmark Suite

Execute all graph building benchmarks using the evaluation runner:

```bash
python -m code_review_graph.eval.runner \
    --repo-root /path/to/your/repo \
    --store /path/to/graph/store \
    --benchmarks build_performance,token_efficiency,search_quality

```

The runner ([`code_review_graph/eval/runner.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/runner.py)) orchestrates the process by:

1. Loading the graph store via `store.get_stats()` to capture baseline statistics
2. Executing selected benchmarks (e.g., `build_performance.run(...)`)
3. Aggregating results into a list of dictionaries
4. Passing data to [`code_review_graph/eval/reporter.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/reporter.py) for markdown report generation

## Complementary Performance Evaluations

In addition to core construction timing, the repository includes benchmarks that indirectly measure graph-building efficiency:

### Token Efficiency Analysis

The **`token_efficiency`** benchmark in [`code_review_graph/eval/benchmarks/token_benchmark.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_benchmark.py) measures total token consumption for agent workflows versus naive file-reading baselines. Key outputs include the token-reduction ratio and total tokens saved, quantifying how graph construction impacts downstream AI agent costs.

### Search Quality Metrics

The **`search_quality`** benchmark in [`code_review_graph/eval/benchmarks/search_quality.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/search_quality.py) evaluates retrieval accuracy using Mean Reciprocal Rank (MRR). While focused on query results, this metric validates that the graph construction process produces a search index that improves information retrieval accuracy.

## Summary

- The `build_performance` benchmark in [`code_review_graph/eval/benchmarks/build_performance.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/build_performance.py) provides the primary measurement for **performance benchmarks for code-review-graph graph building**.
- Three critical phases are instrumented: **flow detection** (`trace_flows`), **community detection** (`detect_communities`), and **search latency** (`store.search_nodes`).
- Key metrics include `nodes_per_second`, `flow_detection_seconds`, and `search_avg_ms`, enabling scalability analysis.
- The evaluation runner ([`code_review_graph/eval/runner.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/runner.py)) coordinates benchmark execution and report generation via [`code_review_graph/eval/reporter.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/reporter.py).
- Complementary benchmarks measure **token efficiency** and **search quality (MRR)** to validate the cost-effectiveness and accuracy of the constructed graph.

## Frequently Asked Questions

### Which file contains the main graph building performance benchmark?

The primary benchmark resides in [`code_review_graph/eval/benchmarks/build_performance.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/build_performance.py), specifically implementing the `build_performance` function between lines 12 and 60. This file contains the timing logic for flow detection, community detection, and search latency measurements using Python's `time.perf_counter()`.

### How is the benchmark execution orchestrated?

Benchmark execution is coordinated by [`code_review_graph/eval/runner.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/runner.py), which loads the graph store, invokes each selected benchmark's `run()` method, and aggregates results. The runner accepts command-line arguments specifying the repository root, store path, and comma-separated benchmark names, then delegates final reporting to [`code_review_graph/eval/reporter.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/reporter.py).

### What does the nodes_per_second metric indicate?

The `nodes_per_second` metric measures graph construction throughput by dividing the total node count (`stats.total_nodes`) by the combined processing time. This indicator helps developers assess scalability and predict construction time for larger codebases, with higher values indicating more efficient graph building performance.

### How do token efficiency benchmarks relate to graph construction?

The `token_efficiency` benchmark in [`code_review_graph/eval/benchmarks/token_benchmark.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_benchmark.py) quantifies the token cost savings achieved by using the constructed graph versus reading entire files. This validates that the graph building overhead is justified by reduced downstream agent token consumption, providing a cost-benefit analysis of the construction pipeline.