# How to Run Evaluation Benchmarks for Token Efficiency and Impact Accuracy in code-review-graph

> Learn to run evaluation benchmarks for token efficiency and impact accuracy in code-review-graph. Install via pip and use the eval CLI or programmatic import for quick analysis.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: performance
- Published: 2026-08-14

---

**To run evaluation benchmarks for token efficiency and impact accuracy in code-review-graph, install the package via pip and invoke the `eval` CLI sub-command with `--benchmarks token_efficiency,impact_accuracy`, or import `run_all_benchmarks` from `code_review_graph.eval.runner` for programmatic execution.**

The `code-review-graph` repository provides a built-in evaluation framework that quantifies how effectively the graph-based review context reduces token consumption compared to naive approaches, and how accurately its change-impact predictions match actual file co-change patterns. This guide walks through the exact commands, Python APIs, and source file locations needed to execute these benchmarks on any local git repository.

## Installation and Repository Setup

Before running benchmarks, you must install the package and prepare a target git repository that contains commit history for analysis.

### Clone and Install the Package

The evaluation framework ships with the core library but requires optional dependencies for data processing and reporting. Install from source with development dependencies:

```bash
git clone https://github.com/tirth8205/code-review-graph.git
cd code-review-graph
pip install -e .

```

Ensure you are running **Python 3.9 or newer**. The evaluation modules rely on `pandas`, `tabulate`, and `tiktoken`, which are declared as optional dependencies in [`pyproject.toml`](https://github.com/tirth8205/code-review-graph/blob/main/pyproject.toml) and installed alongside the package.

### Prepare a Repository Snapshot

Benchmarks execute against a local git repository with existing commit history. For quick validation, create a minimal demo repository:

```bash
mkdir demo-repo && cd demo-repo
git init
echo "print('hello')" > hello.py
git add hello.py && git commit -m "initial"
echo "print('world')" >> hello.py
git commit -am "update"
cd ..

```

Any repository path can be targeted via the `--repo` flag, but the directory must contain a valid `.git` folder with at least two commits to generate meaningful diffs.

## Running Benchmarks via the CLI

The primary entry point for evaluation is the `eval` sub-command defined in [`code_review_graph/cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/cli.py). This command orchestrates the benchmark discovery, execution, and report generation.

### CLI Flags and Parameters

The `eval` command accepts several configuration options:

- `--repo <path>`: Path to the target git repository
- `--benchmarks <list>`: Comma-separated list including `token_efficiency` and/or `impact_accuracy`
- `--output <dir>`: Destination directory for CSV results and markdown reports (defaults to `eval_results`)
- `--config <json>`: Optional JSON file overriding benchmark settings such as test commit ranges

### Execute Both Benchmarks

Run the complete evaluation suite on your prepared repository:

```bash
python -m code_review_graph eval \
    --repo ./demo-repo \
    --benchmarks token_efficiency,impact_accuracy \
    --output ./eval-results

```

During execution, the runner performs the following for each test commit:

1. Collects changed files using `git diff --name-only`
2. Counts **naive tokens** (full file contents) and **standard tokens** (raw `git diff` output)
3. Invokes `code_review_graph.tools.get_review_context` to capture **graph-based token counts**
4. Compares ground-truth co-change sets against graph predictions in both `MODE_GRAPH_DERIVED` and `MODE_CO_CHANGE` configurations

The system writes `token_efficiency_*.csv` and `impact_accuracy_*.csv` to the output directory, plus a [`REPORT.md`](https://github.com/tirth8205/code-review-graph/blob/main/REPORT.md) generated by `code_review_graph.eval.reporter` containing median ratios and error statistics.

## Running Benchmarks Programmatically

For CI/CD integration or custom analysis pipelines, bypass the CLI and call the runner directly from Python.

### Using the Benchmark Runner

Import `run_all_benchmarks` from `code_review_graph.eval.runner` to execute all registered benchmarks:

```python
from pathlib import Path
from code_review_graph.eval.runner import run_all_benchmarks

repo_root = Path("./demo-repo")
results = run_all_benchmarks(repo_root=str(repo_root), base="HEAD~1")

```

This function dynamically loads benchmark modules from `code_review_graph/eval/benchmarks/`, including [`token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/token_efficiency.py) and [`impact_accuracy.py`](https://github.com/tirth8205/code-review-graph/blob/main/impact_accuracy.py), and returns a flat list of result dictionaries.

### Generating Reports

Pass the results list to the reporter to obtain markdown output:

```python
from code_review_graph.eval.reporter import generate_report

markdown_report = generate_report(results)
print(markdown_report)

```

## Core Benchmark Components

Understanding the source file structure helps when customizing evaluation logic or debugging failures.

| Component | Source File | Purpose |
|-----------|-------------|---------|
| **Token Efficiency** | [`code_review_graph/eval/benchmarks/token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py) | Measures naive vs. standard vs. graph-based token consumption; computes median and mean efficiency ratios |
| **Impact Accuracy** | [`code_review_graph/eval/benchmarks/impact_accuracy.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/impact_accuracy.py) | Calculates precision, recall, and MRR for graph-derived and co-change prediction modes |
| **Runner** | [`code_review_graph/eval/runner.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/runner.py) | Orchestrates benchmark discovery, configuration loading, CSV serialization, and handles missing semantic indexes |
| **Reporter** | [`code_review_graph/eval/reporter.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/reporter.py) | Aggregates CSV rows into the final markdown report with statistical summaries |
| **CLI Entry** | [`code_review_graph/cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/cli.py) | Parses command-line arguments and forwards them to the runner |

## Handling Edge Cases and Failures

The evaluation framework includes specific logic for handling imperfect data conditions.

### Error Row Exclusion

If `get_review_context` raises an exception during a commit evaluation, the runner marks that row with `status="error"` and **excludes it from aggregate calculations**. This prevents skewed results that previously occurred when error cases defaulted to `graph_tokens=0`, artificially inflating efficiency ratios.

### Single-File Commit Handling

In [`impact_accuracy.py`](https://github.com/tirth8205/code-review-graph/blob/main/impact_accuracy.py), the **co-change ground-truth mode** (`MODE_CO_CHANGE`) is automatically skipped for single-file commits because no co-change signal exists to compare against. The graph-derived mode (`MODE_GRAPH_DERIVED`) continues to evaluate these commits against the predicted impact set.

### Semantic Search Prerequisites

Benchmarks involving semantic search quality require a pre-built vector index. The runner emits a warning and skips rows requiring the index if it is not present, but continues executing token efficiency and impact accuracy benchmarks that rely solely on graph topology.

## Summary

- **Install** `code-review-graph` from source with `pip install -e .` to access the evaluation framework
- **Run benchmarks via CLI** using `python -m code_review_graph eval --benchmarks token_efficiency,impact_accuracy --repo <path>`
- **Call programmatically** by importing `run_all_benchmarks` from `code_review_graph.eval.runner`
- **Locate source logic** in `code_review_graph/eval/benchmarks/` for token counting and accuracy metrics
- **Review outputs** in the specified output directory: CSV files for raw data and [`REPORT.md`](https://github.com/tirth8205/code-review-graph/blob/main/REPORT.md) for aggregated statistics
- **Handle errors** gracefully: failed tool calls are logged with `status="error"` and excluded from final ratios

## Frequently Asked Questions

### What Python version is required to run the benchmarks?

The evaluation framework requires **Python 3.9 or newer**. This ensures compatibility with the type hints used in [`code_review_graph/eval/runner.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/runner.py) and the dependencies specified in [`pyproject.toml`](https://github.com/tirth8205/code-review-graph/blob/main/pyproject.toml), including modern `pandas` and `tiktoken` versions.

### How does the token efficiency benchmark calculate savings?

The benchmark defined in [`code_review_graph/eval/benchmarks/token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py) compares three token counts: naive (full file contents), standard (raw `git diff` output), and graph-based (output from `get_review_context`). It calculates ratios between these values and reports median and mean savings across the test commit set, excluding any commits where the graph tool failed.

### What metrics does the impact accuracy benchmark report?

According to [`code_review_graph/eval/benchmarks/impact_accuracy.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/impact_accuracy.py), the benchmark reports **precision**, **recall**, and **Mean Reciprocal Rank (MRR)** for both graph-derived predictions and co-change baseline predictions. These metrics compare the predicted set of impacted files against the ground-truth files that actually changed together in the commit history.

### Can I run benchmarks on repositories without a semantic search index?

Yes. Token efficiency and impact accuracy benchmarks function without a semantic search index because they rely on git history and graph topology alone. However, if you include semantic search benchmarks in your `--benchmarks` list, the runner will skip those specific evaluations and emit a warning, while continuing to process the other benchmarks.