How the Token Benchmark in code-review-graph Measures and Reports Efficiency Gains
The token benchmark compares LLM token consumption across three analysis strategies—naïve, standard, and graph-optimised—and reports an efficiency ratio showing how many times fewer tokens the graph-based pipeline requires.
This evaluation system is part of the tirth8205/code-review-graph repository, which optimises code review workflows by representing codebases as graphs. The token benchmark quantifies the practical cost savings of this approach for teams using LLM-based analysis tools.
What the Token Benchmark Measures
The benchmark executes three distinct analysis strategies on the same repository and records token usage for each.
The Three Analysis Strategies
| Strategy | Description | Token Measurement |
|---|---|---|
| Naïve | Full LLM-driven analysis without graph-based pruning | Every LLM call is recorded; total is stored as naive_tokens |
| Standard | Normal pipeline with graph and caching, but no optimisation passes | Token spend is summed as standard_tokens |
| Graph-Optimised | Full pipeline with call-graph pruning, dependency-aware caching, and incremental recomputation | Measured total stored as graph_tokens |
Core Metric: The Efficiency Ratio
The benchmark calculates efficiency using compute_token_efficiency(raw, graph) from code_review_graph/eval/scorer.py. The formula is straightforward:
ratio = raw_tokens / graph_tokens
A ratio greater than 1 indicates multiplicative token savings. For example, a ratio of 4.0 means the graph-optimised approach used 75% fewer tokens than the naïve baseline.
How the Token Benchmark Reports Results
Per-Repository CSV Output
After execution, token_benchmark.run() generates a CSV file with a naming pattern like test_token_efficiency_2026-01-01.csv. Each row contains:
repo— repository nametokens— naïve token count (raw baseline)ratio— efficiency ratio (raw / graph)naive_tokens,standard_tokens,graph_tokens— detailed breakdown by strategy
Aggregation and Summarisation
The code_review_graph/eval/benchmarks.py module provides aggregation helpers that consolidate multiple runs. This enables comparison across repositories or analysis configurations.
Code Example: Running the Token Benchmark
from code_review_graph.eval.token_benchmark import run, aggregate
from code_review_graph.eval.scorer import compute_token_efficiency
# Execute benchmark on a local repository
results = run(repo_path="/path/to/repo", store=my_store, config=my_cfg)
# Individual result structure
# {
# "repo": "my-repo",
# "naive_tokens": 1200,
# "standard_tokens": 800,
# "graph_tokens": 300,
# "ratio": 4.0
# }
print(results[0])
# Aggregate multiple runs into summary CSV
summary = aggregate(results)
summary.to_csv("token_efficiency_summary.csv")
Key Source Files
| File | Purpose |
|---|---|
code_review_graph/eval/token_benchmark.py |
Implements run() to execute the three analysis modes and record token usage |
code_review_graph/eval/scorer.py |
Provides compute_token_efficiency() for ratio calculation |
code_review_graph/eval/benchmarks.py |
Aggregates per-repo results into formatted summaries |
tests/test_eval.py |
Validates benchmark logic with unit tests (e.g., verifying that raw_tokens=10000 and graph_tokens=3000 produce expected ratio of ~3.33) |
Performance Characteristics of the Measurement System
The benchmark is designed for reproducibility and comparability. By fixing the repository state and configuration across all three strategy runs, it isolates the token efficiency impact of graph-based optimisations. The measurement captures actual LLM API calls, making it suitable for cost estimation in production deployments.
Teams can use these measurements to:
- Project API cost reductions when adopting
code-review-graph - Identify repositories where graph optimisations yield the highest returns
- Compare efficiency across different LLM providers or model versions
Summary
- The token benchmark compares naïve, standard, and graph-optimised analysis strategies on identical codebases
- Efficiency gains are reported as a ratio of raw tokens to graph tokens via
compute_token_efficiency() - Results are persisted in CSV format with granular token breakdowns
- The
aggregate()function inbenchmarks.pyenables multi-repository analysis - All measurement logic is validated in
tests/test_eval.pyaccording to thetirth8205/code-review-graphsource code
Frequently Asked Questions
How is the efficiency ratio calculated in the token benchmark?
The ratio is computed by dividing naïve token consumption by graph-optimised token consumption: ratio = naive_tokens / graph_tokens. This value is produced by compute_token_efficiency() in code_review_graph/eval/scorer.py. A ratio of 4.0 indicates the graph-optimised pipeline required one-quarter the tokens of the baseline.
What does the token benchmark CSV output include?
Each CSV row contains the repository name, overall naïve token count, efficiency ratio, and three detailed columns: naive_tokens, standard_tokens, and graph_tokens. This structure enables both high-level summaries and granular analysis of where token savings occur.
Where is the token benchmark implementation located?
The primary implementation resides in code_review_graph/eval/token_benchmark.py, with supporting aggregation logic in code_review_graph/eval/benchmarks.py and scoring utilities in code_review_graph/eval/scorer.py. Unit tests validating the measurement accuracy are found in tests/test_eval.py.
Can I run the token benchmark on multiple repositories?
Yes. Call run() for each repository to collect individual results, then pass the combined list to aggregate() from code_review_graph/eval/benchmarks.py. This produces a consolidated summary suitable for cross-repository efficiency comparisons.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →