# Token Savings Achieved by code-review-graph in Benchmarks: A 65× Reduction

> Discover how code-review-graph slashes token usage by up to 376x in benchmarks, offering a massive 65x median reduction per review. Optimize your code reviews today.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: performance
- Published: 2026-08-11

---

**code-review-graph achieves a median token reduction of approximately 65× per review question, with reductions ranging from 36× to 376× depending on repository size and query complexity.**

The `code-review-graph` project dramatically cuts LLM token consumption by replacing naive full-corpus context with targeted graph queries. This article breaks down the benchmark methodology, the specific savings metrics, and how to verify these reductions in your own workflow.

## How Token Savings Are Measured

The benchmark suite evaluates two realistic scenarios against six real-world repositories (13 commits total). Both scenarios produce identical median results, confirming the robustness of the approach.

### Benchmark Scenarios

- **Token-efficiency benchmark** — Compares naive whole-repo token count against tokens returned by a graph query
- **Agent-baseline benchmark** — Compares top-3 grep-matched files (realistic agent behavior) against graph query results

| Metric | Value |
|--------|-------|
| **Median reduction** | ~65× per question |
| **Range** | 36× – 376× |
| **Maximum observed** | 376× (fastapi repository) |

The headline figures appear in the README's **Benchmarks** section at lines 121–128, with the detailed per-repository breakdown at lines 226–233.

## Why the Graph Approach Saves Tokens

Three architectural decisions drive these savings:

- **Graph-driven impact radius** — Instead of dumping every file, the tool traverses a pre-computed knowledge graph to fetch only relevant nodes (functions, classes, imports, tests)
- **Incremental parsing** — Only files with changed hashes are reparsed, minimizing overhead
- **Conservative token estimation** — Uses a 4-characters-per-token approximation calibrated against OpenAI's `cl100k_base` tokenizer (within ~1% of actual GPT-4 counts)

The token savings calculation lives in [`code_review_graph/context_savings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/context_savings.py), specifically in the `format_context_savings_panel` function at lines 201–220.

## Viewing Token Savings in Practice

### CLI Output

Run any detection with `--brief` to see the boxed savings panel:

```bash
code-review-graph detect-changes --brief path/to/changed/file.py

```

The `--brief` flag triggers this output path in [`cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/cli.py) around line 1743. Sample panel:

```

┌──────────────── Token Savings ────────────────┐
│ Full context would be: 12,932 tokens          │
│ Graph context used:        773 tokens          │
│ Saved:                  12,159 tokens (~94%)   │
│ Breakdown: Functions 580 · Tests 120 · …      │
└───────────────────────────────────────────────┘

```

### Programmatic Access

Access savings metadata directly in Python:

```python
from code_review_graph import detect_changes

result = detect_changes(
    repo_path="my_repo",
    changed_files=["app/main.py"],
    brief=True,
)

print(result["context_savings"])

# → {'estimated': True, 'saved_tokens': 12159, 'saved_percent': 94, ...}

```

### Verify Against Real Tokenizer

Cross-check estimates with `tiktoken`:

```bash
code-review-graph detect-changes --brief --verify path/to/file.py

```

The `--verify` flag uses `tiktoken`'s `cl100k` tokenizer and prints a "Verified" line inside the panel.

## Key Source Files

| File | Purpose |
|------|---------|
| [`README.md`](https://github.com/tirth8205/code-review-graph/blob/main/README.md) (lines 121–128, 226–233) | Benchmark summary and per-repo breakdown |
| [`code_review_graph/context_savings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/context_savings.py) | Token-savings computation and panel formatting |
| [`code_review_graph/cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/cli.py) (line ~1743) | CLI integration for panel display |
| [`code_review_graph/eval/benchmarks/token_efficiency.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/token_efficiency.py) | Formal token-efficiency benchmark |
| [`code_review_graph/eval/benchmarks/agent_baseline.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/agent_baseline.py) | Realistic grep-baseline benchmark |
| [`docs/USAGE.md`](https://github.com/tirth8205/code-review-graph/blob/main/docs/USAGE.md) | Feature documentation |

## Summary

- **Primary result**: code-review-graph achieves ~65× median token reduction per review question
- **Range**: 36× to 376× depending on repository structure
- **Method**: Graph queries replace full-corpus context, fetching only relevant code nodes
- **Verification**: Built-in estimation (4 chars/token) with optional `tiktoken` verification
- **Integration**: Available via CLI `--brief` flag or programmatic `context_savings` field

## Frequently Asked Questions

### How accurate is the token savings estimate?

The default estimate uses a conservative 4-characters-per-token approximation calibrated against OpenAI's `cl100k_base` tokenizer. According to the README at lines 242–244, this stays within ~1% of actual GPT-4 token counts. Use `--verify` with `tiktoken` installed for exact counts.

### What repositories were used in the benchmarks?

Six real-world repositories totaling 13 commits. The detailed table at README lines 226–233 includes `fastapi`, which achieved the maximum 376× reduction due to its large codebase relative to typical change scopes.

### Can I see token savings without using the CLI?

Yes. When calling `detect_changes()` programmatically with `brief=True`, the returned dictionary includes a `context_savings` key with `saved_tokens`, `saved_percent`, and breakdown by node type (functions, tests, etc.).

### Why is the agent-baseline benchmark important?

It validates that the 65× figure isn't artificially inflated by an unrealistic "dump everything" comparison. By matching against top-3 grep results—what a typical code agent would fetch—the benchmark confirms graph queries outperform practical baselines by the same margin.