Token Savings Achieved by code-review-graph in Benchmarks: A 65× Reduction

code-review-graph achieves a median token reduction of approximately 65× per review question, with reductions ranging from 36× to 376× depending on repository size and query complexity.

The code-review-graph project dramatically cuts LLM token consumption by replacing naive full-corpus context with targeted graph queries. This article breaks down the benchmark methodology, the specific savings metrics, and how to verify these reductions in your own workflow.

How Token Savings Are Measured

The benchmark suite evaluates two realistic scenarios against six real-world repositories (13 commits total). Both scenarios produce identical median results, confirming the robustness of the approach.

Benchmark Scenarios

  • Token-efficiency benchmark — Compares naive whole-repo token count against tokens returned by a graph query
  • Agent-baseline benchmark — Compares top-3 grep-matched files (realistic agent behavior) against graph query results
Metric Value
Median reduction ~65× per question
Range 36× – 376×
Maximum observed 376× (fastapi repository)

The headline figures appear in the README's Benchmarks section at lines 121–128, with the detailed per-repository breakdown at lines 226–233.

Why the Graph Approach Saves Tokens

Three architectural decisions drive these savings:

  • Graph-driven impact radius — Instead of dumping every file, the tool traverses a pre-computed knowledge graph to fetch only relevant nodes (functions, classes, imports, tests)
  • Incremental parsing — Only files with changed hashes are reparsed, minimizing overhead
  • Conservative token estimation — Uses a 4-characters-per-token approximation calibrated against OpenAI's cl100k_base tokenizer (within ~1% of actual GPT-4 counts)

The token savings calculation lives in code_review_graph/context_savings.py, specifically in the format_context_savings_panel function at lines 201–220.

Viewing Token Savings in Practice

CLI Output

Run any detection with --brief to see the boxed savings panel:

code-review-graph detect-changes --brief path/to/changed/file.py

The --brief flag triggers this output path in cli.py around line 1743. Sample panel:


┌──────────────── Token Savings ────────────────┐
│ Full context would be: 12,932 tokens          │
│ Graph context used:        773 tokens          │
│ Saved:                  12,159 tokens (~94%)   │
│ Breakdown: Functions 580 · Tests 120 · …      │
└───────────────────────────────────────────────┘

Programmatic Access

Access savings metadata directly in Python:

from code_review_graph import detect_changes

result = detect_changes(
    repo_path="my_repo",
    changed_files=["app/main.py"],
    brief=True,
)

print(result["context_savings"])

# → {'estimated': True, 'saved_tokens': 12159, 'saved_percent': 94, ...}

Verify Against Real Tokenizer

Cross-check estimates with tiktoken:

code-review-graph detect-changes --brief --verify path/to/file.py

The --verify flag uses tiktoken's cl100k tokenizer and prints a "Verified" line inside the panel.

Key Source Files

File Purpose
README.md (lines 121–128, 226–233) Benchmark summary and per-repo breakdown
code_review_graph/context_savings.py Token-savings computation and panel formatting
code_review_graph/cli.py (line ~1743) CLI integration for panel display
code_review_graph/eval/benchmarks/token_efficiency.py Formal token-efficiency benchmark
code_review_graph/eval/benchmarks/agent_baseline.py Realistic grep-baseline benchmark
docs/USAGE.md Feature documentation

Summary

  • Primary result: code-review-graph achieves ~65× median token reduction per review question
  • Range: 36× to 376× depending on repository structure
  • Method: Graph queries replace full-corpus context, fetching only relevant code nodes
  • Verification: Built-in estimation (4 chars/token) with optional tiktoken verification
  • Integration: Available via CLI --brief flag or programmatic context_savings field

Frequently Asked Questions

How accurate is the token savings estimate?

The default estimate uses a conservative 4-characters-per-token approximation calibrated against OpenAI's cl100k_base tokenizer. According to the README at lines 242–244, this stays within ~1% of actual GPT-4 token counts. Use --verify with tiktoken installed for exact counts.

What repositories were used in the benchmarks?

Six real-world repositories totaling 13 commits. The detailed table at README lines 226–233 includes fastapi, which achieved the maximum 376× reduction due to its large codebase relative to typical change scopes.

Can I see token savings without using the CLI?

Yes. When calling detect_changes() programmatically with brief=True, the returned dictionary includes a context_savings key with saved_tokens, saved_percent, and breakdown by node type (functions, tests, etc.).

Why is the agent-baseline benchmark important?

It validates that the 65× figure isn't artificially inflated by an unrealistic "dump everything" comparison. By matching against top-3 grep results—what a typical code agent would fetch—the benchmark confirms graph queries outperform practical baselines by the same margin.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →