How Context Savings Estimates Are Calculated for Token Reduction in Code Review Graph
The estimate_context_savings function in context_savings.py uses a 4-characters-per-token heuristic to approximate token reduction between original and graph-generated context, returning a dictionary with saved_tokens and saved_percent when a valid baseline is available.
The code-review-graph repository quantifies token efficiency gains when replacing full source code context with condensed graph-derived context. This metric helps developers understand how much context window they reclaim when using the tool's minimal context generation. The calculation intentionally avoids model-specific tokenizers in favor of a universal approximation that works across any language model.
The Core Calculation in estimate_context_savings
All token reduction logic resides in code_review_graph/context_savings.py. The estimate_context_savings function (lines 54-82) orchestrates the estimate through a deterministic four-step process that never invokes actual tokenizer APIs.
Step 1: Character-to-Token Conversion with estimate_tokens
The foundation is the CHARS_PER_TOKEN = 4 constant. The estimate_tokens helper function (lines 16-33) applies this heuristic:
- Converts any input to string representation (JSON-serializing non-strings)
- Computes
ceil(len(text) / 4)to guarantee at least one token for non-empty input
This conservative estimate deliberately overcounts slightly rather than undercounting, ensuring reported savings are achievable in practice.
Step 2: Establish the Baseline Token Count
The function accepts two paths for the original context baseline:
original_tokens— use caller-supplied count directlyoriginal_context— derive viaestimate_tokens()when explicit count unavailable
Step 3: Measure the Returned Context
Similarly for the graph-generated output:
returned_tokens— use provided valuereturned_context— estimate viaestimate_tokens()when needed
Step 4: Compute Savings and Percentage
The final arithmetic (lines 70-78) ensures non-negative results:
saved = max(0, baseline - returned)
percent = round((saved / baseline) * 100) if baseline else 0
Return Values and Edge Cases
When baseline > 0, estimate_context_savings returns a standardized dictionary:
{
"estimated": true,
"saved_tokens": <int>,
"saved_percent": <int>
}
If no baseline can be determined—when original_context is None and no original_tokens supplied—the function returns None. This signals to callers that no meaningful comparison is possible.
Practical Usage Examples
Example 1: String Context Comparison
from code_review_graph.context_savings import estimate_context_savings
original = "def foo():\n return 42\n" * 10 # ~300 characters
returned = "def foo():\n return 42\n" # ~30 characters
savings = estimate_context_savings(
original_context=original,
returned_context=returned,
)
# savings → {'estimated': True, 'saved_tokens': 68, 'saved_percent': 93}
print(savings)
The 300-character original yields ~75 tokens; the 30-character returned yields ~8 tokens. The 68-token reduction represents 93% savings.
Example 2: Explicit Token Counts
savings = estimate_context_savings(
original_tokens=1500,
returned_tokens=400,
)
# savings → {'estimated': True, 'saved_tokens': 1100, 'saved_percent': 73}
print(savings)
Supplying pre-computed token counts (e.g., from tiktoken or model-specific tokenizers) bypasses the heuristic entirely for greater precision.
Supporting Utilities
The context_savings.py module includes helpers that consume these estimates without altering them:
attach_context_savings— embeds the savings dictionary into tool result metadataformat_context_savings— renders the estimate for CLI display
These wrappers ensure consistent presentation across the codebase while keeping the core calculation isolated and testable.
Integration with Context Generation
The estimation triggers automatically during minimal context generation. In code_review_graph/tools/context.py, the get_minimal_context function invokes estimate_context_savings when producing graph-derived context, attaching the result to the response payload for downstream observability.
Validation and Testing
Unit tests in tests/test_context_savings.py verify:
- Correct token estimation for strings and objects
- Handling of
Noneinputs and empty contexts - Percentage calculation edge cases (zero baseline, equal inputs)
- Round-trip serialization behavior
These tests ensure the 4-characters-per-token heuristic behaves predictably across data types and boundary conditions.
Summary
estimate_context_savingsincode_review_graph/context_savings.pyis the sole source of token reduction calculations- 4 characters per token heuristic provides model-agnostic estimates via
estimate_tokens - Two input modes: explicit token counts or context strings requiring estimation
- Non-negative savings guarantee via
max(0, baseline - returned) Nonereturn when baseline cannot be determined- Conservative by design: slight overestimation ensures reported savings are achievable
Frequently Asked Questions
Why does the tool use 4 characters per token instead of a real tokenizer?
estimate_tokens uses CHARS_PER_TOKEN = 4 as a deliberately conservative, model-agnostic approximation. Real tokenizers vary significantly between models—GPT-4, Claude, and Llama families all use different encoding schemes. The 4-character heuristic provides a universal lower bound that works across any deployment without adding dependencies or model-specific configuration.
Can I supply my own token counts instead of using the heuristic?
Yes, both original_tokens and returned_tokens parameters accept explicit integers that bypass estimation entirely. Pass values from your preferred tokenizer (e.g., OpenAI's tiktoken, Anthropic's tokenizer API) to replace the heuristic with precise counts.
What happens when the returned context is larger than the original?
The calculation floors at zero: saved = max(0, baseline - returned) ensures saved_tokens never reports negative values. When returned exceeds baseline, you receive {'saved_tokens': 0, 'saved_percent': 0} rather than negative numbers. This reflects practical reality—no tokens were "saved" in this scenario.
Where is the context savings estimate actually displayed to users?
The attach_context_savings helper embeds estimates into tool results, while format_context_savings renders them for CLI output. The primary consumer is get_minimal_context in code_review_graph/tools/context.py, which attaches savings metadata to every graph-derived context response.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →