# Token Efficiency Benchmark in Code Review Graph: How It Measures LLM Context Savings

> Discover the token efficiency benchmark for code review graphs. Learn how it quantifies LLM context savings by using graph-based context over raw code or diffs.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: performance
- Published: 2026-08-18

---

**The token efficiency benchmark quantifies how many tokens are saved by using graph-based context (`get_review_context`) instead of sending raw source code or git diffs to an LLM for code review.**

The benchmark is implemented in the `tirth8205/code-review-graph` repository to validate the core value proposition of the graph-based code review system: reducing token consumption while preserving relevant context. By comparing three distinct token-counting strategies, it produces concrete metrics that demonstrate the efficiency gains of the graph approach.

## What the Token Efficiency Benchmark Measures

The benchmark evaluates three different ways of preparing code for LLM review, measuring token counts for each:

- **Naïve** – full file contents of all changed files
- **Standard** – git diff output between commits
- **Graph-based** – minimal context derived from the code graph

The primary metric is the **ratio of tokens saved** when using graph-based context compared to the baseline approaches.

## Three Token Counting Strategies Explained

Each strategy is implemented as a distinct function in [`code_review_graph/eval/benchmarks/token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py).

### Naïve File Content Strategy

The naïve approach reads the complete contents of every changed file and approximates tokens as `len(text) // 4`, assuming roughly 4 characters per token.

```python

# From token_efficiency.py, lines 45-53

def _count_file_tokens(file_path: Path) -> int:
    """Count tokens in a file using a naïve approximation."""
    try:
        content = file_path.read_text()
        return len(content) // 4  # ~4 chars per token

    except Exception:
        return 0

```

This represents the worst-case scenario: sending entire files to the LLM regardless of what actually changed.

### Standard Diff Strategy

The standard approach runs `git diff` between the commit and its parent, then applies the same token approximation to the diff output.

```python

# From token_efficiency.py, lines 58-73

def _count_diff_tokens(repo_path: Path, commit_sha: str) -> int:
    """Count tokens in git diff output."""
    import subprocess
    result = subprocess.run(
        ["git", "diff", f"{commit_sha}^..{commit_sha}"],
        cwd=repo_path,
        capture_output=True,
        text=True
    )
    return len(result.stdout) // 4

```

This reflects common real-world practice: sending only the changed lines rather than full files.

### Graph-Based Context Strategy

The graph-based approach calls `code_review_graph.tools.get_review_context` to obtain minimal, semantically relevant context, serializes it to JSON, and counts tokens in the result.

```python

# From token_efficiency.py, lines 101-115

def _count_tokens(review_context: dict) -> int:
    """Count tokens in the graph-based context."""
    try:
        json_str = json.dumps(review_context, indent=2)
        return len(json_str) // 4
    except Exception:
        return 0  # Errors are caught and marked in results

```

Errors are explicitly caught and marked with `status != "ok"` so failed context generations don't skew aggregate statistics.

## Benchmark Metrics and Aggregation

For each test commit, the benchmark records five key fields:

| Field | Description |
|-------|-------------|
| `naive_tokens` | Token count from full file contents |
| `standard_tokens` | Token count from git diff output |
| `graph_tokens` | Token count from graph-derived context |
| `naive_to_graph_ratio` | `naive_tokens / max(graph_tokens, 1)` |
| `standard_to_graph_ratio` | `standard_tokens / max(graph_tokens, 1)` |

The `aggregate()` function computes summary statistics across all successful runs, filtering out rows where `status != "ok"`:

```python

# Aggregate implementation, lines 25-44

def aggregate(results: list[dict]) -> dict:
    """Compute median ratios from benchmark results."""
    ok_results = [r for r in results if r.get("status") == "ok"]
    return {
        "median_naive_to_graph_ratio": median(
            r["naive_tokens"] / max(r["graph_tokens"], 1) for r in ok_results
        ),
        "median_standard_to_graph_ratio": median(
            r["standard_tokens"] / max(r["graph_tokens"], 1) for r in ok_results
        ),
        "total_commits": len(results),
        "successful_commits": len(ok_results),
    }

```

## Running the Token Efficiency Benchmark

Execute the benchmark programmatically using the repository's graph store:

```python
from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.incremental import full_build, get_db_path
from code_review_graph.eval.benchmarks import token_efficiency

# Build graph for target repository

repo_path = Path("/path/to/repo")
db_path = get_db_path(repo_path)
store = GraphStore(db_path)
full_build(repo_path, store)

# Configure benchmark with test commits

config = {
    "name": "myrepo",
    "test_commits": [
        {"sha": "abc1234", "description": "feature X"},
        {"sha": "def5678", "description": "bugfix Y"},
    ],
}

# Execute benchmark and aggregate

results = token_efficiency.run(repo_path, store, config)
summary = token_efficiency.aggregate(results)
print(f"Median naive→graph ratio: {summary['median_naive_to_graph_ratio']:.2f}x")
print(f"Median standard→graph ratio: {summary['median_standard_to_graph_ratio']:.2f}x")

```

The `run()` function returns a list of result dictionaries that can be serialized to CSV or analyzed directly.

## Computing Token Efficiency Metrics

For single measurements or unit testing, use `compute_token_efficiency` from [`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py):

```python
from code_review_graph.eval.scorer import compute_token_efficiency

metrics = compute_token_efficiency(raw_tokens=1200, graph_tokens=300)

# Result:

# {

#   "raw_tokens": 1200,

#   "graph_tokens": 300,

#   "ratio": 0.25,

#   "reduction_percent": 75.0,

# }

```

This helper returns both the ratio (`graph / raw`) and the percentage reduction, useful for reporting and assertions in [`tests/test_eval.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py).

## How Measurement Errors Are Handled

The benchmark implements defensive error handling to ensure reliable aggregates:

- **File read failures** in `_count_file_tokens` return 0 tokens
- **Git diff failures** in `_count_diff_tokens` return 0 tokens
- **Context generation failures** set `status = "error"` and exclude the row from aggregation

This design prevents transient failures (network issues, corrupted commits, graph build errors) from invalidating the entire benchmark run.

## Summary

- The **token efficiency benchmark** compares three LLM context strategies: naïve file contents, standard git diffs, and graph-derived minimal context.
- Token counts use a **4-character-per-token approximation** (`len(text) // 4`) across all strategies for consistency.
- **Core metrics** are `naive_to_graph_ratio` and `standard_to_graph_ratio`, showing how many times larger the baseline approaches are compared to the graph approach.
- **Aggregation filters failed runs** via `status == "ok"` checks, ensuring reliable median calculations.
- The implementation lives in [`code_review_graph/eval/benchmarks/token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py) with supporting utilities in [`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py).

## Frequently Asked Questions

### What tokenization method does the benchmark use?

The benchmark uses a **naïve character-based approximation**: `len(text) // 4`. This assumes approximately 4 characters per token, which correlates reasonably well with LLM tokenizers like GPT-4's without requiring heavy dependencies. All three strategies use the same approximation to ensure fair comparison.

### Can the benchmark handle repositories with binary files or encoding issues?

Yes. The `_count_file_tokens` function wraps file reads in try-except blocks and returns 0 on any exception. Similarly, the graph context strategy catches serialization errors and marks rows with `status = "error"`. These rows are excluded from aggregate calculations, so problematic files don't corrupt the overall metrics.

### How do I interpret the ratio values in benchmark results?

A `naive_to_graph_ratio` of 10 means the naïve approach uses 10× as many tokens as the graph-based approach—equivalently, the graph approach achieves a **90% token reduction**. The `standard_to_graph_ratio` similarly compares git diffs to graph context. Higher ratios indicate greater efficiency gains from using the graph.

### Where can I find example benchmark output data?

The benchmark writes CSV files with columns matching the result dictionary fields: `repo`, `commit_sha`, `naive_tokens`, `standard_tokens`, `graph_tokens`, `naive_to_graph_ratio`, `standard_to_graph_ratio`, and `status`. Inspect these files directly or load them with Python's `csv.DictReader` for custom analysis.