How the Token Benchmark in code-review-graph Measures and Reports Efficiency Gains

The token benchmark compares LLM token consumption across three analysis strategies—naïve, standard, and graph-optimised—and reports an efficiency ratio showing how many times fewer tokens the graph-based pipeline requires.

This evaluation system is part of the tirth8205/code-review-graph repository, which optimises code review workflows by representing codebases as graphs. The token benchmark quantifies the practical cost savings of this approach for teams using LLM-based analysis tools.

What the Token Benchmark Measures

The benchmark executes three distinct analysis strategies on the same repository and records token usage for each.

The Three Analysis Strategies

Strategy Description Token Measurement
Naïve Full LLM-driven analysis without graph-based pruning Every LLM call is recorded; total is stored as naive_tokens
Standard Normal pipeline with graph and caching, but no optimisation passes Token spend is summed as standard_tokens
Graph-Optimised Full pipeline with call-graph pruning, dependency-aware caching, and incremental recomputation Measured total stored as graph_tokens

Core Metric: The Efficiency Ratio

The benchmark calculates efficiency using compute_token_efficiency(raw, graph) from code_review_graph/eval/scorer.py. The formula is straightforward:

ratio = raw_tokens / graph_tokens

A ratio greater than 1 indicates multiplicative token savings. For example, a ratio of 4.0 means the graph-optimised approach used 75% fewer tokens than the naïve baseline.

How the Token Benchmark Reports Results

Per-Repository CSV Output

After execution, token_benchmark.run() generates a CSV file with a naming pattern like test_token_efficiency_2026-01-01.csv. Each row contains:

  • repo — repository name
  • tokens — naïve token count (raw baseline)
  • ratio — efficiency ratio (raw / graph)
  • naive_tokens, standard_tokens, graph_tokens — detailed breakdown by strategy

Aggregation and Summarisation

The code_review_graph/eval/benchmarks.py module provides aggregation helpers that consolidate multiple runs. This enables comparison across repositories or analysis configurations.

Code Example: Running the Token Benchmark

from code_review_graph.eval.token_benchmark import run, aggregate
from code_review_graph.eval.scorer import compute_token_efficiency

# Execute benchmark on a local repository

results = run(repo_path="/path/to/repo", store=my_store, config=my_cfg)

# Individual result structure

# {

#   "repo": "my-repo",

#   "naive_tokens": 1200,

#   "standard_tokens": 800,

#   "graph_tokens": 300,

#   "ratio": 4.0

# }

print(results[0])

# Aggregate multiple runs into summary CSV

summary = aggregate(results)
summary.to_csv("token_efficiency_summary.csv")

Key Source Files

File Purpose
code_review_graph/eval/token_benchmark.py Implements run() to execute the three analysis modes and record token usage
code_review_graph/eval/scorer.py Provides compute_token_efficiency() for ratio calculation
code_review_graph/eval/benchmarks.py Aggregates per-repo results into formatted summaries
tests/test_eval.py Validates benchmark logic with unit tests (e.g., verifying that raw_tokens=10000 and graph_tokens=3000 produce expected ratio of ~3.33)

Performance Characteristics of the Measurement System

The benchmark is designed for reproducibility and comparability. By fixing the repository state and configuration across all three strategy runs, it isolates the token efficiency impact of graph-based optimisations. The measurement captures actual LLM API calls, making it suitable for cost estimation in production deployments.

Teams can use these measurements to:

  • Project API cost reductions when adopting code-review-graph
  • Identify repositories where graph optimisations yield the highest returns
  • Compare efficiency across different LLM providers or model versions

Summary

  • The token benchmark compares naïve, standard, and graph-optimised analysis strategies on identical codebases
  • Efficiency gains are reported as a ratio of raw tokens to graph tokens via compute_token_efficiency()
  • Results are persisted in CSV format with granular token breakdowns
  • The aggregate() function in benchmarks.py enables multi-repository analysis
  • All measurement logic is validated in tests/test_eval.py according to the tirth8205/code-review-graph source code

Frequently Asked Questions

How is the efficiency ratio calculated in the token benchmark?

The ratio is computed by dividing naïve token consumption by graph-optimised token consumption: ratio = naive_tokens / graph_tokens. This value is produced by compute_token_efficiency() in code_review_graph/eval/scorer.py. A ratio of 4.0 indicates the graph-optimised pipeline required one-quarter the tokens of the baseline.

What does the token benchmark CSV output include?

Each CSV row contains the repository name, overall naïve token count, efficiency ratio, and three detailed columns: naive_tokens, standard_tokens, and graph_tokens. This structure enables both high-level summaries and granular analysis of where token savings occur.

Where is the token benchmark implementation located?

The primary implementation resides in code_review_graph/eval/token_benchmark.py, with supporting aggregation logic in code_review_graph/eval/benchmarks.py and scoring utilities in code_review_graph/eval/scorer.py. Unit tests validating the measurement accuracy are found in tests/test_eval.py.

Can I run the token benchmark on multiple repositories?

Yes. Call run() for each repository to collect individual results, then pass the combined list to aggregate() from code_review_graph/eval/benchmarks.py. This produces a consolidated summary suitable for cross-repository efficiency comparisons.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →