How to Run Evaluation Benchmarks for Token Efficiency and Impact Accuracy in code-review-graph
To run evaluation benchmarks for token efficiency and impact accuracy in code-review-graph, install the package via pip and invoke the eval CLI sub-command with --benchmarks token_efficiency,impact_accuracy, or import run_all_benchmarks from code_review_graph.eval.runner for programmatic execution.
The code-review-graph repository provides a built-in evaluation framework that quantifies how effectively the graph-based review context reduces token consumption compared to naive approaches, and how accurately its change-impact predictions match actual file co-change patterns. This guide walks through the exact commands, Python APIs, and source file locations needed to execute these benchmarks on any local git repository.
Installation and Repository Setup
Before running benchmarks, you must install the package and prepare a target git repository that contains commit history for analysis.
Clone and Install the Package
The evaluation framework ships with the core library but requires optional dependencies for data processing and reporting. Install from source with development dependencies:
git clone https://github.com/tirth8205/code-review-graph.git
cd code-review-graph
pip install -e .
Ensure you are running Python 3.9 or newer. The evaluation modules rely on pandas, tabulate, and tiktoken, which are declared as optional dependencies in pyproject.toml and installed alongside the package.
Prepare a Repository Snapshot
Benchmarks execute against a local git repository with existing commit history. For quick validation, create a minimal demo repository:
mkdir demo-repo && cd demo-repo
git init
echo "print('hello')" > hello.py
git add hello.py && git commit -m "initial"
echo "print('world')" >> hello.py
git commit -am "update"
cd ..
Any repository path can be targeted via the --repo flag, but the directory must contain a valid .git folder with at least two commits to generate meaningful diffs.
Running Benchmarks via the CLI
The primary entry point for evaluation is the eval sub-command defined in code_review_graph/cli.py. This command orchestrates the benchmark discovery, execution, and report generation.
CLI Flags and Parameters
The eval command accepts several configuration options:
--repo <path>: Path to the target git repository--benchmarks <list>: Comma-separated list includingtoken_efficiencyand/orimpact_accuracy--output <dir>: Destination directory for CSV results and markdown reports (defaults toeval_results)--config <json>: Optional JSON file overriding benchmark settings such as test commit ranges
Execute Both Benchmarks
Run the complete evaluation suite on your prepared repository:
python -m code_review_graph eval \
--repo ./demo-repo \
--benchmarks token_efficiency,impact_accuracy \
--output ./eval-results
During execution, the runner performs the following for each test commit:
- Collects changed files using
git diff --name-only - Counts naive tokens (full file contents) and standard tokens (raw
git diffoutput) - Invokes
code_review_graph.tools.get_review_contextto capture graph-based token counts - Compares ground-truth co-change sets against graph predictions in both
MODE_GRAPH_DERIVEDandMODE_CO_CHANGEconfigurations
The system writes token_efficiency_*.csv and impact_accuracy_*.csv to the output directory, plus a REPORT.md generated by code_review_graph.eval.reporter containing median ratios and error statistics.
Running Benchmarks Programmatically
For CI/CD integration or custom analysis pipelines, bypass the CLI and call the runner directly from Python.
Using the Benchmark Runner
Import run_all_benchmarks from code_review_graph.eval.runner to execute all registered benchmarks:
from pathlib import Path
from code_review_graph.eval.runner import run_all_benchmarks
repo_root = Path("./demo-repo")
results = run_all_benchmarks(repo_root=str(repo_root), base="HEAD~1")
This function dynamically loads benchmark modules from code_review_graph/eval/benchmarks/, including token_efficiency.py and impact_accuracy.py, and returns a flat list of result dictionaries.
Generating Reports
Pass the results list to the reporter to obtain markdown output:
from code_review_graph.eval.reporter import generate_report
markdown_report = generate_report(results)
print(markdown_report)
Core Benchmark Components
Understanding the source file structure helps when customizing evaluation logic or debugging failures.
| Component | Source File | Purpose |
|---|---|---|
| Token Efficiency | code_review_graph/eval/benchmarks/token_efficiency.py |
Measures naive vs. standard vs. graph-based token consumption; computes median and mean efficiency ratios |
| Impact Accuracy | code_review_graph/eval/benchmarks/impact_accuracy.py |
Calculates precision, recall, and MRR for graph-derived and co-change prediction modes |
| Runner | code_review_graph/eval/runner.py |
Orchestrates benchmark discovery, configuration loading, CSV serialization, and handles missing semantic indexes |
| Reporter | code_review_graph/eval/reporter.py |
Aggregates CSV rows into the final markdown report with statistical summaries |
| CLI Entry | code_review_graph/cli.py |
Parses command-line arguments and forwards them to the runner |
Handling Edge Cases and Failures
The evaluation framework includes specific logic for handling imperfect data conditions.
Error Row Exclusion
If get_review_context raises an exception during a commit evaluation, the runner marks that row with status="error" and excludes it from aggregate calculations. This prevents skewed results that previously occurred when error cases defaulted to graph_tokens=0, artificially inflating efficiency ratios.
Single-File Commit Handling
In impact_accuracy.py, the co-change ground-truth mode (MODE_CO_CHANGE) is automatically skipped for single-file commits because no co-change signal exists to compare against. The graph-derived mode (MODE_GRAPH_DERIVED) continues to evaluate these commits against the predicted impact set.
Semantic Search Prerequisites
Benchmarks involving semantic search quality require a pre-built vector index. The runner emits a warning and skips rows requiring the index if it is not present, but continues executing token efficiency and impact accuracy benchmarks that rely solely on graph topology.
Summary
- Install
code-review-graphfrom source withpip install -e .to access the evaluation framework - Run benchmarks via CLI using
python -m code_review_graph eval --benchmarks token_efficiency,impact_accuracy --repo <path> - Call programmatically by importing
run_all_benchmarksfromcode_review_graph.eval.runner - Locate source logic in
code_review_graph/eval/benchmarks/for token counting and accuracy metrics - Review outputs in the specified output directory: CSV files for raw data and
REPORT.mdfor aggregated statistics - Handle errors gracefully: failed tool calls are logged with
status="error"and excluded from final ratios
Frequently Asked Questions
What Python version is required to run the benchmarks?
The evaluation framework requires Python 3.9 or newer. This ensures compatibility with the type hints used in code_review_graph/eval/runner.py and the dependencies specified in pyproject.toml, including modern pandas and tiktoken versions.
How does the token efficiency benchmark calculate savings?
The benchmark defined in code_review_graph/eval/benchmarks/token_efficiency.py compares three token counts: naive (full file contents), standard (raw git diff output), and graph-based (output from get_review_context). It calculates ratios between these values and reports median and mean savings across the test commit set, excluding any commits where the graph tool failed.
What metrics does the impact accuracy benchmark report?
According to code_review_graph/eval/benchmarks/impact_accuracy.py, the benchmark reports precision, recall, and Mean Reciprocal Rank (MRR) for both graph-derived predictions and co-change baseline predictions. These metrics compare the predicted set of impacted files against the ground-truth files that actually changed together in the commit history.
Can I run benchmarks on repositories without a semantic search index?
Yes. Token efficiency and impact accuracy benchmarks function without a semantic search index because they rely on git history and graph topology alone. However, if you include semantic search benchmarks in your --benchmarks list, the runner will skip those specific evaluations and emit a warning, while continuing to process the other benchmarks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →