# How the Token Benchmark in code-review-graph Measures and Reports Efficiency Gains

> Discover how the token benchmark in code-review-graph quantifies efficiency gains. Learn how graph-optimized pipelines use significantly fewer LLM tokens for code analysis.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: performance
- Published: 2026-08-15

---

**The token benchmark compares LLM token consumption across three analysis strategies—naïve, standard, and graph-optimised—and reports an efficiency ratio showing how many times fewer tokens the graph-based pipeline requires.**

This evaluation system is part of the `tirth8205/code-review-graph` repository, which optimises code review workflows by representing codebases as graphs. The token benchmark quantifies the practical cost savings of this approach for teams using LLM-based analysis tools.

## What the Token Benchmark Measures

The benchmark executes three distinct analysis strategies on the same repository and records token usage for each.

### The Three Analysis Strategies

| Strategy | Description | Token Measurement |
|----------|-------------|-----------------|
| **Naïve** | Full LLM-driven analysis without graph-based pruning | Every LLM call is recorded; total is stored as `naive_tokens` |
| **Standard** | Normal pipeline with graph and caching, but no optimisation passes | Token spend is summed as `standard_tokens` |
| **Graph-Optimised** | Full pipeline with call-graph pruning, dependency-aware caching, and incremental recomputation | Measured total stored as `graph_tokens` |

### Core Metric: The Efficiency Ratio

The benchmark calculates efficiency using `compute_token_efficiency(raw, graph)` from [`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py). The formula is straightforward:

```python
ratio = raw_tokens / graph_tokens

```

A ratio greater than 1 indicates multiplicative token savings. For example, a ratio of `4.0` means the graph-optimised approach used 75% fewer tokens than the naïve baseline.

## How the Token Benchmark Reports Results

### Per-Repository CSV Output

After execution, `token_benchmark.run()` generates a CSV file with a naming pattern like `test_token_efficiency_2026-01-01.csv`. Each row contains:

- `repo` — repository name
- `tokens` — naïve token count (raw baseline)
- `ratio` — efficiency ratio (`raw / graph`)
- `naive_tokens`, `standard_tokens`, `graph_tokens` — detailed breakdown by strategy

### Aggregation and Summarisation

The [`code_review_graph/eval/benchmarks.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks.py) module provides aggregation helpers that consolidate multiple runs. This enables comparison across repositories or analysis configurations.

## Code Example: Running the Token Benchmark

```python
from code_review_graph.eval.token_benchmark import run, aggregate
from code_review_graph.eval.scorer import compute_token_efficiency

# Execute benchmark on a local repository

results = run(repo_path="/path/to/repo", store=my_store, config=my_cfg)

# Individual result structure

# {

#   "repo": "my-repo",

#   "naive_tokens": 1200,

#   "standard_tokens": 800,

#   "graph_tokens": 300,

#   "ratio": 4.0

# }

print(results[0])

# Aggregate multiple runs into summary CSV

summary = aggregate(results)
summary.to_csv("token_efficiency_summary.csv")

```

## Key Source Files

| File | Purpose |
|------|---------|
| [`code_review_graph/eval/token_benchmark.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/token_benchmark.py) | Implements `run()` to execute the three analysis modes and record token usage |
| [`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py) | Provides `compute_token_efficiency()` for ratio calculation |
| [`code_review_graph/eval/benchmarks.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks.py) | Aggregates per-repo results into formatted summaries |
| [`tests/test_eval.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py) | Validates benchmark logic with unit tests (e.g., verifying that `raw_tokens=10000` and `graph_tokens=3000` produce expected ratio of ~3.33) |

## Performance Characteristics of the Measurement System

The benchmark is designed for reproducibility and comparability. By fixing the repository state and configuration across all three strategy runs, it isolates the token efficiency impact of graph-based optimisations. The measurement captures actual LLM API calls, making it suitable for cost estimation in production deployments.

Teams can use these measurements to:
- Project API cost reductions when adopting `code-review-graph`
- Identify repositories where graph optimisations yield the highest returns
- Compare efficiency across different LLM providers or model versions

## Summary

- The token benchmark compares **naïve**, **standard**, and **graph-optimised** analysis strategies on identical codebases
- Efficiency gains are reported as a **ratio of raw tokens to graph tokens** via `compute_token_efficiency()`
- Results are persisted in **CSV format** with granular token breakdowns
- The `aggregate()` function in [`benchmarks.py`](https://github.com/tirth8205/code-review-graph/blob/main/benchmarks.py) enables multi-repository analysis
- All measurement logic is validated in [`tests/test_eval.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py) according to the `tirth8205/code-review-graph` source code

## Frequently Asked Questions

### How is the efficiency ratio calculated in the token benchmark?

The ratio is computed by dividing naïve token consumption by graph-optimised token consumption: `ratio = naive_tokens / graph_tokens`. This value is produced by `compute_token_efficiency()` in [`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py). A ratio of 4.0 indicates the graph-optimised pipeline required one-quarter the tokens of the baseline.

### What does the token benchmark CSV output include?

Each CSV row contains the repository name, overall naïve token count, efficiency ratio, and three detailed columns: `naive_tokens`, `standard_tokens`, and `graph_tokens`. This structure enables both high-level summaries and granular analysis of where token savings occur.

### Where is the token benchmark implementation located?

The primary implementation resides in [`code_review_graph/eval/token_benchmark.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/token_benchmark.py), with supporting aggregation logic in [`code_review_graph/eval/benchmarks.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks.py) and scoring utilities in [`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py). Unit tests validating the measurement accuracy are found in [`tests/test_eval.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py).

### Can I run the token benchmark on multiple repositories?

Yes. Call `run()` for each repository to collect individual results, then pass the combined list to `aggregate()` from [`code_review_graph/eval/benchmarks.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks.py). This produces a consolidated summary suitable for cross-repository efficiency comparisons.