Token Efficiency Benchmark in Code Review Graph: How It Measures LLM Context Savings

The token efficiency benchmark quantifies how many tokens are saved by using graph-based context (get_review_context) instead of sending raw source code or git diffs to an LLM for code review.

The benchmark is implemented in the tirth8205/code-review-graph repository to validate the core value proposition of the graph-based code review system: reducing token consumption while preserving relevant context. By comparing three distinct token-counting strategies, it produces concrete metrics that demonstrate the efficiency gains of the graph approach.

What the Token Efficiency Benchmark Measures

The benchmark evaluates three different ways of preparing code for LLM review, measuring token counts for each:

  • Naïve – full file contents of all changed files
  • Standard – git diff output between commits
  • Graph-based – minimal context derived from the code graph

The primary metric is the ratio of tokens saved when using graph-based context compared to the baseline approaches.

Three Token Counting Strategies Explained

Each strategy is implemented as a distinct function in code_review_graph/eval/benchmarks/token_efficiency.py.

Naïve File Content Strategy

The naïve approach reads the complete contents of every changed file and approximates tokens as len(text) // 4, assuming roughly 4 characters per token.


# From token_efficiency.py, lines 45-53

def _count_file_tokens(file_path: Path) -> int:
    """Count tokens in a file using a naïve approximation."""
    try:
        content = file_path.read_text()
        return len(content) // 4  # ~4 chars per token

    except Exception:
        return 0

This represents the worst-case scenario: sending entire files to the LLM regardless of what actually changed.

Standard Diff Strategy

The standard approach runs git diff between the commit and its parent, then applies the same token approximation to the diff output.


# From token_efficiency.py, lines 58-73

def _count_diff_tokens(repo_path: Path, commit_sha: str) -> int:
    """Count tokens in git diff output."""
    import subprocess
    result = subprocess.run(
        ["git", "diff", f"{commit_sha}^..{commit_sha}"],
        cwd=repo_path,
        capture_output=True,
        text=True
    )
    return len(result.stdout) // 4

This reflects common real-world practice: sending only the changed lines rather than full files.

Graph-Based Context Strategy

The graph-based approach calls code_review_graph.tools.get_review_context to obtain minimal, semantically relevant context, serializes it to JSON, and counts tokens in the result.


# From token_efficiency.py, lines 101-115

def _count_tokens(review_context: dict) -> int:
    """Count tokens in the graph-based context."""
    try:
        json_str = json.dumps(review_context, indent=2)
        return len(json_str) // 4
    except Exception:
        return 0  # Errors are caught and marked in results

Errors are explicitly caught and marked with status != "ok" so failed context generations don't skew aggregate statistics.

Benchmark Metrics and Aggregation

For each test commit, the benchmark records five key fields:

Field Description
naive_tokens Token count from full file contents
standard_tokens Token count from git diff output
graph_tokens Token count from graph-derived context
naive_to_graph_ratio naive_tokens / max(graph_tokens, 1)
standard_to_graph_ratio standard_tokens / max(graph_tokens, 1)

The aggregate() function computes summary statistics across all successful runs, filtering out rows where status != "ok":


# Aggregate implementation, lines 25-44

def aggregate(results: list[dict]) -> dict:
    """Compute median ratios from benchmark results."""
    ok_results = [r for r in results if r.get("status") == "ok"]
    return {
        "median_naive_to_graph_ratio": median(
            r["naive_tokens"] / max(r["graph_tokens"], 1) for r in ok_results
        ),
        "median_standard_to_graph_ratio": median(
            r["standard_tokens"] / max(r["graph_tokens"], 1) for r in ok_results
        ),
        "total_commits": len(results),
        "successful_commits": len(ok_results),
    }

Running the Token Efficiency Benchmark

Execute the benchmark programmatically using the repository's graph store:

from pathlib import Path
from code_review_graph.graph import GraphStore
from code_review_graph.incremental import full_build, get_db_path
from code_review_graph.eval.benchmarks import token_efficiency

# Build graph for target repository

repo_path = Path("/path/to/repo")
db_path = get_db_path(repo_path)
store = GraphStore(db_path)
full_build(repo_path, store)

# Configure benchmark with test commits

config = {
    "name": "myrepo",
    "test_commits": [
        {"sha": "abc1234", "description": "feature X"},
        {"sha": "def5678", "description": "bugfix Y"},
    ],
}

# Execute benchmark and aggregate

results = token_efficiency.run(repo_path, store, config)
summary = token_efficiency.aggregate(results)
print(f"Median naive→graph ratio: {summary['median_naive_to_graph_ratio']:.2f}x")
print(f"Median standard→graph ratio: {summary['median_standard_to_graph_ratio']:.2f}x")

The run() function returns a list of result dictionaries that can be serialized to CSV or analyzed directly.

Computing Token Efficiency Metrics

For single measurements or unit testing, use compute_token_efficiency from code_review_graph/eval/scorer.py:

from code_review_graph.eval.scorer import compute_token_efficiency

metrics = compute_token_efficiency(raw_tokens=1200, graph_tokens=300)

# Result:

# {

#   "raw_tokens": 1200,

#   "graph_tokens": 300,

#   "ratio": 0.25,

#   "reduction_percent": 75.0,

# }

This helper returns both the ratio (graph / raw) and the percentage reduction, useful for reporting and assertions in tests/test_eval.py.

How Measurement Errors Are Handled

The benchmark implements defensive error handling to ensure reliable aggregates:

  • File read failures in _count_file_tokens return 0 tokens
  • Git diff failures in _count_diff_tokens return 0 tokens
  • Context generation failures set status = "error" and exclude the row from aggregation

This design prevents transient failures (network issues, corrupted commits, graph build errors) from invalidating the entire benchmark run.

Summary

  • The token efficiency benchmark compares three LLM context strategies: naïve file contents, standard git diffs, and graph-derived minimal context.
  • Token counts use a 4-character-per-token approximation (len(text) // 4) across all strategies for consistency.
  • Core metrics are naive_to_graph_ratio and standard_to_graph_ratio, showing how many times larger the baseline approaches are compared to the graph approach.
  • Aggregation filters failed runs via status == "ok" checks, ensuring reliable median calculations.
  • The implementation lives in code_review_graph/eval/benchmarks/token_efficiency.py with supporting utilities in code_review_graph/eval/scorer.py.

Frequently Asked Questions

What tokenization method does the benchmark use?

The benchmark uses a naïve character-based approximation: len(text) // 4. This assumes approximately 4 characters per token, which correlates reasonably well with LLM tokenizers like GPT-4's without requiring heavy dependencies. All three strategies use the same approximation to ensure fair comparison.

Can the benchmark handle repositories with binary files or encoding issues?

Yes. The _count_file_tokens function wraps file reads in try-except blocks and returns 0 on any exception. Similarly, the graph context strategy catches serialization errors and marks rows with status = "error". These rows are excluded from aggregate calculations, so problematic files don't corrupt the overall metrics.

How do I interpret the ratio values in benchmark results?

A naive_to_graph_ratio of 10 means the naïve approach uses 10× as many tokens as the graph-based approach—equivalently, the graph approach achieves a 90% token reduction. The standard_to_graph_ratio similarly compares git diffs to graph context. Higher ratios indicate greater efficiency gains from using the graph.

Where can I find example benchmark output data?

The benchmark writes CSV files with columns matching the result dictionary fields: repo, commit_sha, naive_tokens, standard_tokens, graph_tokens, naive_to_graph_ratio, standard_to_graph_ratio, and status. Inspect these files directly or load them with Python's csv.DictReader for custom analysis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →