How Token Savings Estimate Is Calculated and Verified in Code-Review-Graph

Token savings are calculated by comparing three token-counting methods—naïve (full file contents), standard (git diff), and graph-based (knowledge graph JSON)—then verified through unit tests on compute_token_efficiency and integration tests on the benchmark runner.

The token savings estimate is a core metric in tirth8205/code-review-graph that quantifies how much the knowledge-graph representation reduces token usage compared to traditional approaches. This article explains the calculation methodology implemented in code_review_graph/eval/benchmarks/token_efficiency.py, the aggregation logic that produces median savings ratios, and the test suite in tests/test_eval.py that ensures accuracy.

Three Methods for Token Counting

The benchmark evaluates token efficiency by measuring three distinct approaches for the same commit or change set:

  • Naïve token count — counts all characters from full file contents of changed files, approximating 1 token per 4 characters via _count_file_tokens
  • Standard token count — counts tokens in raw git diff output via _count_diff_tokens
  • Graph-based token count — counts tokens in the JSON payload returned by get_review_context, representing the knowledge-graph structure

Each method serves as a baseline to demonstrate the compression achieved by representing code changes as a structured graph rather than raw text or diffs.

Computing the Token Savings Ratio

The core calculation lives in code_review_graph/eval/scorer.py. The compute_token_efficiency function accepts raw and graph token counts, returning a structured result with the savings ratio:

from code_review_graph.eval.scorer import compute_token_efficiency

# Example: 10,000 tokens with naïve approach, 3,000 with graph approach

result = compute_token_efficiency(raw_tokens=10_000, graph_tokens=3_000)

# Returns: {

#   "raw_tokens": 10000,

#   "graph_tokens": 3000,

#   "ratio": 3.3

# }

The ratio represents approximate tokens saved per graph token consumed—higher values indicate greater efficiency.

Benchmark Implementation and Aggregation

The token_efficiency.py benchmark orchestrates the three-way comparison across commits:


# From code_review_graph/eval/benchmarks/token_efficiency.py

naive_to_graph_ratio = round(naive_tokens / max(graph_tokens, 1), 1)
standard_to_graph_ratio = round(standard_tokens / max(graph_tokens, 1), 1)

The max(graph_tokens, 1) guard prevents division-by-zero while preserving accuracy for typical cases.

Benchmark results are written as CSV rows with fields for each token count, both ratios, and a status field. The aggregation logic in the same file computes median and mean statistics while excluding rows where status="error"—ensuring failed graph generations don't skew savings estimates.

Run the full benchmark with:

python -m code_review_graph.eval.runner \
    --benchmark token_efficiency \
    --repo /path/to/target/repo \
    --config path/to/config.yaml

Verification Through Automated Testing

The token savings calculation is verified at two levels:

Unit Tests for Core Logic

tests/test_eval.py validates compute_token_efficiency directly, confirming:

  • Correct field population (raw_tokens, graph_tokens, ratio)
  • Accurate ratio computation
  • Proper handling of edge cases

Integration Tests for Benchmark Integrity

tests/test_token_efficiency.py exercises the complete pipeline:

  • Benchmark execution on synthetic commits
  • CSV row structure validation
  • Aggregation logic correctness, including error-row exclusion

This dual-layer testing ensures both the mathematical correctness of ratios and the robustness of real-world benchmark execution.

Practical Example: Analyzing Savings

Process benchmark output programmatically:

from code_review_graph.eval.benchmarks.token_efficiency import aggregate

# rows loaded from benchmark CSV output

summary = aggregate(rows)

print(f"Median naïve→graph ratio: {summary['median_naive_to_graph_ratio']}")
print(f"Mean naïve→graph ratio: {summary['mean_naive_to_graph_ratio']}")

Typical results show ratios of 2.5–5×, meaning the knowledge graph reduces token consumption by 60–80% compared to full file contents.

Key Source Files

File Purpose
[code_review_graph/eval/scorer.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py) compute_token_efficiency function and scoring utilities
[code_review_graph/eval/benchmarks/token_efficiency.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py) Three-way token comparison and aggregation
[tests/test_eval.py](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py) Unit tests for token efficiency computation
[tests/test_token_efficiency.py](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_token_efficiency.py) Integration tests for benchmark runner and aggregation

Summary

  • Token savings estimate compares naïve, standard (diff), and graph-based token counts
  • Ratios are computed with division-by-zero guards and rounded to one decimal place
  • Aggregation excludes failed runs to prevent inflated savings estimates
  • Verification spans unit tests on core math and integration tests on full pipeline
  • Typical savings range from 2.5× to 5× reduction in token usage

Frequently Asked Questions

How does the benchmark handle commits where graph generation fails?

Failed graph generations are marked with status="error" in the CSV output. The aggregation logic explicitly filters these rows before computing median and mean statistics, ensuring that incomplete data doesn't artificially inflate token savings estimates.

Why compare three methods instead of just naïve versus graph-based?

The standard (diff-based) method serves as a practical middle ground that developers commonly use today. Including it demonstrates whether the knowledge graph improves upon both the worst-case scenario (full files) and the typical current practice (diffs).

What tokenization approximation does the benchmark use?

The implementation uses a character-based heuristic of approximately 1 token per 4 characters via _count_tokens and related helper functions. This matches common tokenizer behavior (e.g., GPT-2/3 family) without requiring external dependencies or API calls during benchmarking.

Can I run the token efficiency benchmark on my own repository?

Yes. Install the package, create a configuration file specifying target commits, and execute python -m code_review_graph.eval.runner --benchmark token_efficiency with your repo path and config. Results are emitted as CSV for custom analysis or use the built-in aggregate function for summary statistics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →