How to Run Evaluation Benchmarks for Token Efficiency and Impact Accuracy in code-review-graph

To run evaluation benchmarks for token efficiency and impact accuracy in code-review-graph, install the package via pip and invoke the eval CLI sub-command with --benchmarks token_efficiency,impact_accuracy, or import run_all_benchmarks from code_review_graph.eval.runner for programmatic execution.

The code-review-graph repository provides a built-in evaluation framework that quantifies how effectively the graph-based review context reduces token consumption compared to naive approaches, and how accurately its change-impact predictions match actual file co-change patterns. This guide walks through the exact commands, Python APIs, and source file locations needed to execute these benchmarks on any local git repository.

Installation and Repository Setup

Before running benchmarks, you must install the package and prepare a target git repository that contains commit history for analysis.

Clone and Install the Package

The evaluation framework ships with the core library but requires optional dependencies for data processing and reporting. Install from source with development dependencies:

git clone https://github.com/tirth8205/code-review-graph.git
cd code-review-graph
pip install -e .

Ensure you are running Python 3.9 or newer. The evaluation modules rely on pandas, tabulate, and tiktoken, which are declared as optional dependencies in pyproject.toml and installed alongside the package.

Prepare a Repository Snapshot

Benchmarks execute against a local git repository with existing commit history. For quick validation, create a minimal demo repository:

mkdir demo-repo && cd demo-repo
git init
echo "print('hello')" > hello.py
git add hello.py && git commit -m "initial"
echo "print('world')" >> hello.py
git commit -am "update"
cd ..

Any repository path can be targeted via the --repo flag, but the directory must contain a valid .git folder with at least two commits to generate meaningful diffs.

Running Benchmarks via the CLI

The primary entry point for evaluation is the eval sub-command defined in code_review_graph/cli.py. This command orchestrates the benchmark discovery, execution, and report generation.

CLI Flags and Parameters

The eval command accepts several configuration options:

  • --repo <path>: Path to the target git repository
  • --benchmarks <list>: Comma-separated list including token_efficiency and/or impact_accuracy
  • --output <dir>: Destination directory for CSV results and markdown reports (defaults to eval_results)
  • --config <json>: Optional JSON file overriding benchmark settings such as test commit ranges

Execute Both Benchmarks

Run the complete evaluation suite on your prepared repository:

python -m code_review_graph eval \
    --repo ./demo-repo \
    --benchmarks token_efficiency,impact_accuracy \
    --output ./eval-results

During execution, the runner performs the following for each test commit:

  1. Collects changed files using git diff --name-only
  2. Counts naive tokens (full file contents) and standard tokens (raw git diff output)
  3. Invokes code_review_graph.tools.get_review_context to capture graph-based token counts
  4. Compares ground-truth co-change sets against graph predictions in both MODE_GRAPH_DERIVED and MODE_CO_CHANGE configurations

The system writes token_efficiency_*.csv and impact_accuracy_*.csv to the output directory, plus a REPORT.md generated by code_review_graph.eval.reporter containing median ratios and error statistics.

Running Benchmarks Programmatically

For CI/CD integration or custom analysis pipelines, bypass the CLI and call the runner directly from Python.

Using the Benchmark Runner

Import run_all_benchmarks from code_review_graph.eval.runner to execute all registered benchmarks:

from pathlib import Path
from code_review_graph.eval.runner import run_all_benchmarks

repo_root = Path("./demo-repo")
results = run_all_benchmarks(repo_root=str(repo_root), base="HEAD~1")

This function dynamically loads benchmark modules from code_review_graph/eval/benchmarks/, including token_efficiency.py and impact_accuracy.py, and returns a flat list of result dictionaries.

Generating Reports

Pass the results list to the reporter to obtain markdown output:

from code_review_graph.eval.reporter import generate_report

markdown_report = generate_report(results)
print(markdown_report)

Core Benchmark Components

Understanding the source file structure helps when customizing evaluation logic or debugging failures.

Component Source File Purpose
Token Efficiency code_review_graph/eval/benchmarks/token_efficiency.py Measures naive vs. standard vs. graph-based token consumption; computes median and mean efficiency ratios
Impact Accuracy code_review_graph/eval/benchmarks/impact_accuracy.py Calculates precision, recall, and MRR for graph-derived and co-change prediction modes
Runner code_review_graph/eval/runner.py Orchestrates benchmark discovery, configuration loading, CSV serialization, and handles missing semantic indexes
Reporter code_review_graph/eval/reporter.py Aggregates CSV rows into the final markdown report with statistical summaries
CLI Entry code_review_graph/cli.py Parses command-line arguments and forwards them to the runner

Handling Edge Cases and Failures

The evaluation framework includes specific logic for handling imperfect data conditions.

Error Row Exclusion

If get_review_context raises an exception during a commit evaluation, the runner marks that row with status="error" and excludes it from aggregate calculations. This prevents skewed results that previously occurred when error cases defaulted to graph_tokens=0, artificially inflating efficiency ratios.

Single-File Commit Handling

In impact_accuracy.py, the co-change ground-truth mode (MODE_CO_CHANGE) is automatically skipped for single-file commits because no co-change signal exists to compare against. The graph-derived mode (MODE_GRAPH_DERIVED) continues to evaluate these commits against the predicted impact set.

Semantic Search Prerequisites

Benchmarks involving semantic search quality require a pre-built vector index. The runner emits a warning and skips rows requiring the index if it is not present, but continues executing token efficiency and impact accuracy benchmarks that rely solely on graph topology.

Summary

  • Install code-review-graph from source with pip install -e . to access the evaluation framework
  • Run benchmarks via CLI using python -m code_review_graph eval --benchmarks token_efficiency,impact_accuracy --repo <path>
  • Call programmatically by importing run_all_benchmarks from code_review_graph.eval.runner
  • Locate source logic in code_review_graph/eval/benchmarks/ for token counting and accuracy metrics
  • Review outputs in the specified output directory: CSV files for raw data and REPORT.md for aggregated statistics
  • Handle errors gracefully: failed tool calls are logged with status="error" and excluded from final ratios

Frequently Asked Questions

What Python version is required to run the benchmarks?

The evaluation framework requires Python 3.9 or newer. This ensures compatibility with the type hints used in code_review_graph/eval/runner.py and the dependencies specified in pyproject.toml, including modern pandas and tiktoken versions.

How does the token efficiency benchmark calculate savings?

The benchmark defined in code_review_graph/eval/benchmarks/token_efficiency.py compares three token counts: naive (full file contents), standard (raw git diff output), and graph-based (output from get_review_context). It calculates ratios between these values and reports median and mean savings across the test commit set, excluding any commits where the graph tool failed.

What metrics does the impact accuracy benchmark report?

According to code_review_graph/eval/benchmarks/impact_accuracy.py, the benchmark reports precision, recall, and Mean Reciprocal Rank (MRR) for both graph-derived predictions and co-change baseline predictions. These metrics compare the predicted set of impacted files against the ground-truth files that actually changed together in the commit history.

Can I run benchmarks on repositories without a semantic search index?

Yes. Token efficiency and impact accuracy benchmarks function without a semantic search index because they rely on git history and graph topology alone. However, if you include semantic search benchmarks in your --benchmarks list, the runner will skip those specific evaluations and emit a warning, while continuing to process the other benchmarks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →