How Token Savings Estimate Is Calculated and Verified in Code-Review-Graph
Token savings are calculated by comparing three token-counting methods—naïve (full file contents), standard (git diff), and graph-based (knowledge graph JSON)—then verified through unit tests on compute_token_efficiency and integration tests on the benchmark runner.
The token savings estimate is a core metric in tirth8205/code-review-graph that quantifies how much the knowledge-graph representation reduces token usage compared to traditional approaches. This article explains the calculation methodology implemented in code_review_graph/eval/benchmarks/token_efficiency.py, the aggregation logic that produces median savings ratios, and the test suite in tests/test_eval.py that ensures accuracy.
Three Methods for Token Counting
The benchmark evaluates token efficiency by measuring three distinct approaches for the same commit or change set:
- Naïve token count — counts all characters from full file contents of changed files, approximating 1 token per 4 characters via
_count_file_tokens - Standard token count — counts tokens in raw
git diffoutput via_count_diff_tokens - Graph-based token count — counts tokens in the JSON payload returned by
get_review_context, representing the knowledge-graph structure
Each method serves as a baseline to demonstrate the compression achieved by representing code changes as a structured graph rather than raw text or diffs.
Computing the Token Savings Ratio
The core calculation lives in code_review_graph/eval/scorer.py. The compute_token_efficiency function accepts raw and graph token counts, returning a structured result with the savings ratio:
from code_review_graph.eval.scorer import compute_token_efficiency
# Example: 10,000 tokens with naïve approach, 3,000 with graph approach
result = compute_token_efficiency(raw_tokens=10_000, graph_tokens=3_000)
# Returns: {
# "raw_tokens": 10000,
# "graph_tokens": 3000,
# "ratio": 3.3
# }
The ratio represents approximate tokens saved per graph token consumed—higher values indicate greater efficiency.
Benchmark Implementation and Aggregation
The token_efficiency.py benchmark orchestrates the three-way comparison across commits:
# From code_review_graph/eval/benchmarks/token_efficiency.py
naive_to_graph_ratio = round(naive_tokens / max(graph_tokens, 1), 1)
standard_to_graph_ratio = round(standard_tokens / max(graph_tokens, 1), 1)
The max(graph_tokens, 1) guard prevents division-by-zero while preserving accuracy for typical cases.
Benchmark results are written as CSV rows with fields for each token count, both ratios, and a status field. The aggregation logic in the same file computes median and mean statistics while excluding rows where status="error"—ensuring failed graph generations don't skew savings estimates.
Run the full benchmark with:
python -m code_review_graph.eval.runner \
--benchmark token_efficiency \
--repo /path/to/target/repo \
--config path/to/config.yaml
Verification Through Automated Testing
The token savings calculation is verified at two levels:
Unit Tests for Core Logic
tests/test_eval.py validates compute_token_efficiency directly, confirming:
- Correct field population (
raw_tokens,graph_tokens,ratio) - Accurate ratio computation
- Proper handling of edge cases
Integration Tests for Benchmark Integrity
tests/test_token_efficiency.py exercises the complete pipeline:
- Benchmark execution on synthetic commits
- CSV row structure validation
- Aggregation logic correctness, including error-row exclusion
This dual-layer testing ensures both the mathematical correctness of ratios and the robustness of real-world benchmark execution.
Practical Example: Analyzing Savings
Process benchmark output programmatically:
from code_review_graph.eval.benchmarks.token_efficiency import aggregate
# rows loaded from benchmark CSV output
summary = aggregate(rows)
print(f"Median naïve→graph ratio: {summary['median_naive_to_graph_ratio']}")
print(f"Mean naïve→graph ratio: {summary['mean_naive_to_graph_ratio']}")
Typical results show ratios of 2.5–5×, meaning the knowledge graph reduces token consumption by 60–80% compared to full file contents.
Key Source Files
| File | Purpose |
|---|---|
[code_review_graph/eval/scorer.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py) |
compute_token_efficiency function and scoring utilities |
[code_review_graph/eval/benchmarks/token_efficiency.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py) |
Three-way token comparison and aggregation |
[tests/test_eval.py](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py) |
Unit tests for token efficiency computation |
[tests/test_token_efficiency.py](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_token_efficiency.py) |
Integration tests for benchmark runner and aggregation |
Summary
- Token savings estimate compares naïve, standard (diff), and graph-based token counts
- Ratios are computed with division-by-zero guards and rounded to one decimal place
- Aggregation excludes failed runs to prevent inflated savings estimates
- Verification spans unit tests on core math and integration tests on full pipeline
- Typical savings range from 2.5× to 5× reduction in token usage
Frequently Asked Questions
How does the benchmark handle commits where graph generation fails?
Failed graph generations are marked with status="error" in the CSV output. The aggregation logic explicitly filters these rows before computing median and mean statistics, ensuring that incomplete data doesn't artificially inflate token savings estimates.
Why compare three methods instead of just naïve versus graph-based?
The standard (diff-based) method serves as a practical middle ground that developers commonly use today. Including it demonstrates whether the knowledge graph improves upon both the worst-case scenario (full files) and the typical current practice (diffs).
What tokenization approximation does the benchmark use?
The implementation uses a character-based heuristic of approximately 1 token per 4 characters via _count_tokens and related helper functions. This matches common tokenizer behavior (e.g., GPT-2/3 family) without requiring external dependencies or API calls during benchmarking.
Can I run the token efficiency benchmark on my own repository?
Yes. Install the package, create a configuration file specifying target commits, and execute python -m code_review_graph.eval.runner --benchmark token_efficiency with your repo path and config. Results are emitted as CSV for custom analysis or use the built-in aggregate function for summary statistics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →