How code-review-graph Reduces Token Usage for AI Code Reviews by 90%+
code-review-graph cuts token consumption by building a persistent knowledge graph and sending only the impacted sub-graph to LLMs instead of full source files.
The code-review-graph open-source tool tackles a critical cost bottleneck in AI-powered code reviews: token usage. Rather than transmitting entire repositories to language models, it constructs a queryable knowledge graph of files, classes, functions, imports, and tests. When changes are detected, the tool isolates precisely the context an LLM needs to evaluate the modification—dramatically shrinking payloads and lowering API costs.
How Knowledge Graphs Enable Token Reduction
Traditional AI code reviews often dump full files or even entire repositories into the context window. This wastes tokens on unchanged, irrelevant code. code-review-graph takes a fundamentally different approach through graph-based context selection.
Persistent Repository Graph Structure
The tool parses the codebase once and stores a structured graph. In code_review_graph/graph.py, the Graph class manages nodes for every meaningful code unit—functions, classes, files, test dependencies—with edges representing relationships like calls, imports, and inheritance. This graph enables precise, programmatic querying of code relationships.
When a review is requested, the graph becomes the single source of truth for what matters.
The Five-Step Token Reduction Workflow
1. Calculate Impact Radius from Git Changes
Instead of reading changed files wholesale, code-review-graph queries the graph for affected nodes and their dependencies.
The tools.get_impact_radius function (called from CLI entry points in cli.py lines 123-129) traverses the graph to find:
- Directly modified functions and classes
- Their callers and callees
- Data flow paths
- Related test coverage gaps
Similarly, tools.list_flows and the flows_cmd handler (lines 190-196) extract execution paths relevant to the change set.
This impact-driven selection typically reduces context to 1-5% of raw source tokens.
2. Serialize Compact JSON Payloads
Selected nodes are converted to minimal JSON representations. The serialization logic in code_review_graph/graph.py (graph.to_json and related methods) captures:
- Node identifiers and types
- Function signatures
- Brief source snippets (not full files)
- Relationship edges
The code_review_graph/changes.py module assembles this into the final response payload. No full file contents traverse the network.
3. Attach Verified Token Savings Metadata
The code_review_graph/context_savings.py module computes and exposes concrete savings metrics. The estimate_context_savings() function implements a conservative token model (4 characters ≈ 1 token):
# context_savings.py – estimate_context_savings()
saved = max(0, baseline - returned)
percent = round((saved / baseline) * 100)
return {"estimated": True,
"saved_tokens": int(saved),
"saved_percent": int(percent)}
Results populate a context_savings key in every response, enabling transparency and optimization.
4. Optionally Verify with Real Tokenizers
For accurate measurement, pass --verify to enable tiktoken integration. The verify_with_tiktoken function (lines 156-190 in context_savings.py) tokenizes both the naive full-context baseline and the actual JSON payload using OpenAI's reference tokenizer:
code-review-graph detect-changes --brief --verify
This produces verified counts like:
┌──────────────── Token Savings ────────────────┐
│ Full context would be: 12,932 tokens │
│ Graph context used: 773 tokens │
│ Saved: 12,159 tokens (~94%)│
│ Verified (tiktoken): 12,120 tokens (~93%) [12,932 → 812]
│ Breakdown: Functions 580 · Tests 120 · …
└───────────────────────────────────────────────┘
5. Risk-Scored Prioritization
The graph enables intelligent token budgeting. Nodes receive risk scores based on complexity, test coverage gaps, and change sensitivity. Only the highest-risk items consume precious context space, ensuring LLM attention focuses where it matters most.
Command-Line Usage Examples
Run a brief review with automatic savings display:
code-review-graph detect-changes --brief
Update the graph first, then review with verification:
code-review-graph update --brief --verify
Programmatic Integration
Embed token savings tracking in custom tools:
from code_review_graph.context_savings import attach_context_savings
response = {"graph": my_subgraph}
response = attach_context_savings(
response,
original_context=repo_root, # naive baseline (all files)
returned_context=response["graph"], # compact payload
)
Key Source Files
| File | Purpose |
|---|---|
code_review_graph/context_savings.py |
Token estimation, verification, and CLI panel formatting |
code_review_graph/graph.py |
Core graph structures and JSON serialization |
code_review_graph/changes.py |
Response assembly with context_savings metadata |
code_review_graph/cli.py |
CLI flags --brief (line 779) and --verify (line 992) |
docs/USAGE.md |
User documentation for Token Savings features |
Summary
- Graph-based context selection isolates only impacted code paths, eliminating irrelevant file content
- Compact JSON serialization transmits identifiers and signatures instead of full sources
- Built-in savings measurement provides transparency via
context_savings.pyestimation and optionaltiktokenverification - Risk scoring prioritizes high-value context within limited token budgets
- Typical reduction: 90%+ tokens (often 94-95% in verified scenarios) for standard pull request reviews
Frequently Asked Questions
How accurate is the token savings estimate?
The default estimate uses a conservative 4-character-per-token heuristic. When --verify is passed, verify_with_tiktoken loads the actual cl100k_base tokenizer and computes precise counts. Verified results typically align within 1-3% of estimates.
Can I use code-review-graph with any LLM provider?
Yes. The tool outputs standard JSON payloads with compact context. Any provider accepting JSON input—OpenAI, Anthropic, local models via Ollama—receives the optimized context. Token savings calculations reference OpenAI's tokenizer but the payload format is universal.
What if my change affects many files across the codebase?
The impact radius query in tools.get_impact_radius scales with graph connectivity, not file count. Even large refactorings produce focused subgraphs because the tool follows actual code relationships (calls, imports, flows) rather than file boundaries. The CLI will display expanded context size and reduced—but still substantial—percentage savings.
Does the graph require frequent rebuilding?
The update command incrementally refreshes the graph. For active repositories, run this in CI or pre-commit hooks. Stale graphs may miss new relationships, causing slightly larger contexts until updated—no correctness impact, only token efficiency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →