Token Savings Achieved by code-review-graph in Benchmarks: A 65× Reduction
code-review-graph achieves a median token reduction of approximately 65× per review question, with reductions ranging from 36× to 376× depending on repository size and query complexity.
The code-review-graph project dramatically cuts LLM token consumption by replacing naive full-corpus context with targeted graph queries. This article breaks down the benchmark methodology, the specific savings metrics, and how to verify these reductions in your own workflow.
How Token Savings Are Measured
The benchmark suite evaluates two realistic scenarios against six real-world repositories (13 commits total). Both scenarios produce identical median results, confirming the robustness of the approach.
Benchmark Scenarios
- Token-efficiency benchmark — Compares naive whole-repo token count against tokens returned by a graph query
- Agent-baseline benchmark — Compares top-3 grep-matched files (realistic agent behavior) against graph query results
| Metric | Value |
|---|---|
| Median reduction | ~65× per question |
| Range | 36× – 376× |
| Maximum observed | 376× (fastapi repository) |
The headline figures appear in the README's Benchmarks section at lines 121–128, with the detailed per-repository breakdown at lines 226–233.
Why the Graph Approach Saves Tokens
Three architectural decisions drive these savings:
- Graph-driven impact radius — Instead of dumping every file, the tool traverses a pre-computed knowledge graph to fetch only relevant nodes (functions, classes, imports, tests)
- Incremental parsing — Only files with changed hashes are reparsed, minimizing overhead
- Conservative token estimation — Uses a 4-characters-per-token approximation calibrated against OpenAI's
cl100k_basetokenizer (within ~1% of actual GPT-4 counts)
The token savings calculation lives in code_review_graph/context_savings.py, specifically in the format_context_savings_panel function at lines 201–220.
Viewing Token Savings in Practice
CLI Output
Run any detection with --brief to see the boxed savings panel:
code-review-graph detect-changes --brief path/to/changed/file.py
The --brief flag triggers this output path in cli.py around line 1743. Sample panel:
┌──────────────── Token Savings ────────────────┐
│ Full context would be: 12,932 tokens │
│ Graph context used: 773 tokens │
│ Saved: 12,159 tokens (~94%) │
│ Breakdown: Functions 580 · Tests 120 · … │
└───────────────────────────────────────────────┘
Programmatic Access
Access savings metadata directly in Python:
from code_review_graph import detect_changes
result = detect_changes(
repo_path="my_repo",
changed_files=["app/main.py"],
brief=True,
)
print(result["context_savings"])
# → {'estimated': True, 'saved_tokens': 12159, 'saved_percent': 94, ...}
Verify Against Real Tokenizer
Cross-check estimates with tiktoken:
code-review-graph detect-changes --brief --verify path/to/file.py
The --verify flag uses tiktoken's cl100k tokenizer and prints a "Verified" line inside the panel.
Key Source Files
| File | Purpose |
|---|---|
README.md (lines 121–128, 226–233) |
Benchmark summary and per-repo breakdown |
code_review_graph/context_savings.py |
Token-savings computation and panel formatting |
code_review_graph/cli.py (line ~1743) |
CLI integration for panel display |
code_review_graph/eval/benchmarks/token_efficiency.py |
Formal token-efficiency benchmark |
code_review_graph/eval/benchmarks/agent_baseline.py |
Realistic grep-baseline benchmark |
docs/USAGE.md |
Feature documentation |
Summary
- Primary result: code-review-graph achieves ~65× median token reduction per review question
- Range: 36× to 376× depending on repository structure
- Method: Graph queries replace full-corpus context, fetching only relevant code nodes
- Verification: Built-in estimation (4 chars/token) with optional
tiktokenverification - Integration: Available via CLI
--briefflag or programmaticcontext_savingsfield
Frequently Asked Questions
How accurate is the token savings estimate?
The default estimate uses a conservative 4-characters-per-token approximation calibrated against OpenAI's cl100k_base tokenizer. According to the README at lines 242–244, this stays within ~1% of actual GPT-4 token counts. Use --verify with tiktoken installed for exact counts.
What repositories were used in the benchmarks?
Six real-world repositories totaling 13 commits. The detailed table at README lines 226–233 includes fastapi, which achieved the maximum 376× reduction due to its large codebase relative to typical change scopes.
Can I see token savings without using the CLI?
Yes. When calling detect_changes() programmatically with brief=True, the returned dictionary includes a context_savings key with saved_tokens, saved_percent, and breakdown by node type (functions, tests, etc.).
Why is the agent-baseline benchmark important?
It validates that the 65× figure isn't artificially inflated by an unrealistic "dump everything" comparison. By matching against top-3 grep results—what a typical code agent would fetch—the benchmark confirms graph queries outperform practical baselines by the same margin.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →