# How Token Savings Estimate Is Calculated and Verified in Code-Review-Graph

> Learn how code-review-graph calculates and verifies token savings using naïve standard and graph-based methods. Discover the testing methods for accurate efficiency.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: how-to-guide
- Published: 2026-08-18

---

**Token savings are calculated by comparing three token-counting methods—naïve (full file contents), standard (git diff), and graph-based (knowledge graph JSON)—then verified through unit tests on `compute_token_efficiency` and integration tests on the benchmark runner.**

The **token savings estimate** is a core metric in tirth8205/code-review-graph that quantifies how much the knowledge-graph representation reduces token usage compared to traditional approaches. This article explains the calculation methodology implemented in [`code_review_graph/eval/benchmarks/token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py), the aggregation logic that produces median savings ratios, and the test suite in [`tests/test_eval.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py) that ensures accuracy.

## Three Methods for Token Counting

The benchmark evaluates token efficiency by measuring three distinct approaches for the same commit or change set:

- **Naïve token count** — counts all characters from full file contents of changed files, approximating 1 token per 4 characters via `_count_file_tokens`
- **Standard token count** — counts tokens in raw `git diff` output via `_count_diff_tokens`
- **Graph-based token count** — counts tokens in the JSON payload returned by `get_review_context`, representing the knowledge-graph structure

Each method serves as a baseline to demonstrate the compression achieved by representing code changes as a structured graph rather than raw text or diffs.

## Computing the Token Savings Ratio

The core calculation lives in [`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py). The `compute_token_efficiency` function accepts raw and graph token counts, returning a structured result with the savings ratio:

```python
from code_review_graph.eval.scorer import compute_token_efficiency

# Example: 10,000 tokens with naïve approach, 3,000 with graph approach

result = compute_token_efficiency(raw_tokens=10_000, graph_tokens=3_000)

# Returns: {

#   "raw_tokens": 10000,

#   "graph_tokens": 3000,

#   "ratio": 3.3

# }

```

The ratio represents approximate tokens saved per graph token consumed—higher values indicate greater efficiency.

## Benchmark Implementation and Aggregation

The [`token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/token_efficiency.py) benchmark orchestrates the three-way comparison across commits:

```python

# From code_review_graph/eval/benchmarks/token_efficiency.py

naive_to_graph_ratio = round(naive_tokens / max(graph_tokens, 1), 1)
standard_to_graph_ratio = round(standard_tokens / max(graph_tokens, 1), 1)

```

The `max(graph_tokens, 1)` guard prevents division-by-zero while preserving accuracy for typical cases.

Benchmark results are written as CSV rows with fields for each token count, both ratios, and a status field. The aggregation logic in the same file computes median and mean statistics while **excluding rows where `status="error"`**—ensuring failed graph generations don't skew savings estimates.

Run the full benchmark with:

```bash
python -m code_review_graph.eval.runner \
    --benchmark token_efficiency \
    --repo /path/to/target/repo \
    --config path/to/config.yaml

```

## Verification Through Automated Testing

The token savings calculation is verified at two levels:

### Unit Tests for Core Logic

[`tests/test_eval.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py) validates `compute_token_efficiency` directly, confirming:
- Correct field population (`raw_tokens`, `graph_tokens`, `ratio`)
- Accurate ratio computation
- Proper handling of edge cases

### Integration Tests for Benchmark Integrity

[`tests/test_token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_token_efficiency.py) exercises the complete pipeline:
- Benchmark execution on synthetic commits
- CSV row structure validation
- Aggregation logic correctness, including error-row exclusion

This dual-layer testing ensures both the mathematical correctness of ratios and the robustness of real-world benchmark execution.

## Practical Example: Analyzing Savings

Process benchmark output programmatically:

```python
from code_review_graph.eval.benchmarks.token_efficiency import aggregate

# rows loaded from benchmark CSV output

summary = aggregate(rows)

print(f"Median naïve→graph ratio: {summary['median_naive_to_graph_ratio']}")
print(f"Mean naïve→graph ratio: {summary['mean_naive_to_graph_ratio']}")

```

Typical results show ratios of **2.5–5×**, meaning the knowledge graph reduces token consumption by 60–80% compared to full file contents.

## Key Source Files

| File | Purpose |
|------|---------|
| [[`code_review_graph/eval/scorer.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/scorer.py) | `compute_token_efficiency` function and scoring utilities |
| [[`code_review_graph/eval/benchmarks/token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py) | Three-way token comparison and aggregation |
| [[`tests/test_eval.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py)](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_eval.py) | Unit tests for token efficiency computation |
| [[`tests/test_token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_token_efficiency.py)](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_token_efficiency.py) | Integration tests for benchmark runner and aggregation |

## Summary

- **Token savings estimate** compares naïve, standard (diff), and graph-based token counts
- **Ratios are computed** with division-by-zero guards and rounded to one decimal place
- **Aggregation excludes failed runs** to prevent inflated savings estimates
- **Verification spans** unit tests on core math and integration tests on full pipeline
- **Typical savings** range from 2.5× to 5× reduction in token usage

## Frequently Asked Questions

### How does the benchmark handle commits where graph generation fails?

Failed graph generations are marked with `status="error"` in the CSV output. The aggregation logic explicitly filters these rows before computing median and mean statistics, ensuring that incomplete data doesn't artificially inflate token savings estimates.

### Why compare three methods instead of just naïve versus graph-based?

The standard (diff-based) method serves as a practical middle ground that developers commonly use today. Including it demonstrates whether the knowledge graph improves upon both the worst-case scenario (full files) and the typical current practice (diffs).

### What tokenization approximation does the benchmark use?

The implementation uses a character-based heuristic of approximately 1 token per 4 characters via `_count_tokens` and related helper functions. This matches common tokenizer behavior (e.g., GPT-2/3 family) without requiring external dependencies or API calls during benchmarking.

### Can I run the token efficiency benchmark on my own repository?

Yes. Install the package, create a configuration file specifying target commits, and execute `python -m code_review_graph.eval.runner --benchmark token_efficiency` with your repo path and config. Results are emitted as CSV for custom analysis or use the built-in `aggregate` function for summary statistics.