# How Context Savings Estimates Are Calculated for Token Reduction in Code Review Graph

> Discover how context savings estimates calculate token reduction in code review graphs. Learn about the 4-characters-per-token heuristic used for optimization.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: how-to-guide
- Published: 2026-08-16

---

**The `estimate_context_savings` function in [`context_savings.py`](https://github.com/tirth8205/code-review-graph/blob/main/context_savings.py) uses a 4-characters-per-token heuristic to approximate token reduction between original and graph-generated context, returning a dictionary with `saved_tokens` and `saved_percent` when a valid baseline is available.**

The `code-review-graph` repository quantifies token efficiency gains when replacing full source code context with condensed graph-derived context. This metric helps developers understand how much context window they reclaim when using the tool's minimal context generation. The calculation intentionally avoids model-specific tokenizers in favor of a universal approximation that works across any language model.

## The Core Calculation in `estimate_context_savings`

All token reduction logic resides in [`code_review_graph/context_savings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/context_savings.py). The `estimate_context_savings` function (lines 54-82) orchestrates the estimate through a deterministic four-step process that never invokes actual tokenizer APIs.

### Step 1: Character-to-Token Conversion with `estimate_tokens`

The foundation is the `CHARS_PER_TOKEN = 4` constant. The `estimate_tokens` helper function (lines 16-33) applies this heuristic:

- Converts any input to string representation (JSON-serializing non-strings)
- Computes `ceil(len(text) / 4)` to guarantee at least one token for non-empty input

This conservative estimate deliberately overcounts slightly rather than undercounting, ensuring reported savings are achievable in practice.

### Step 2: Establish the Baseline Token Count

The function accepts two paths for the original context baseline:

- **`original_tokens`** — use caller-supplied count directly
- **`original_context`** — derive via `estimate_tokens()` when explicit count unavailable

### Step 3: Measure the Returned Context

Similarly for the graph-generated output:

- **`returned_tokens`** — use provided value
- **`returned_context`** — estimate via `estimate_tokens()` when needed

### Step 4: Compute Savings and Percentage

The final arithmetic (lines 70-78) ensures non-negative results:

```python
saved   = max(0, baseline - returned)
percent = round((saved / baseline) * 100) if baseline else 0

```

## Return Values and Edge Cases

When `baseline > 0`, `estimate_context_savings` returns a standardized dictionary:

```json
{
  "estimated": true,
  "saved_tokens": <int>,
  "saved_percent": <int>
}

```

If no baseline can be determined—when `original_context` is `None` and no `original_tokens` supplied—the function returns `None`. This signals to callers that no meaningful comparison is possible.

## Practical Usage Examples

### Example 1: String Context Comparison

```python
from code_review_graph.context_savings import estimate_context_savings

original = "def foo():\n    return 42\n" * 10   # ~300 characters

returned = "def foo():\n    return 42\n"          # ~30 characters

savings = estimate_context_savings(
    original_context=original,
    returned_context=returned,
)

# savings → {'estimated': True, 'saved_tokens': 68, 'saved_percent': 93}

print(savings)

```

The 300-character original yields ~75 tokens; the 30-character returned yields ~8 tokens. The 68-token reduction represents 93% savings.

### Example 2: Explicit Token Counts

```python
savings = estimate_context_savings(
    original_tokens=1500,
    returned_tokens=400,
)

# savings → {'estimated': True, 'saved_tokens': 1100, 'saved_percent': 73}

print(savings)

```

Supplying pre-computed token counts (e.g., from `tiktoken` or model-specific tokenizers) bypasses the heuristic entirely for greater precision.

## Supporting Utilities

The [`context_savings.py`](https://github.com/tirth8205/code-review-graph/blob/main/context_savings.py) module includes helpers that consume these estimates without altering them:

- **`attach_context_savings`** — embeds the savings dictionary into tool result metadata
- **`format_context_savings`** — renders the estimate for CLI display

These wrappers ensure consistent presentation across the codebase while keeping the core calculation isolated and testable.

## Integration with Context Generation

The estimation triggers automatically during minimal context generation. In [`code_review_graph/tools/context.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/context.py), the `get_minimal_context` function invokes `estimate_context_savings` when producing graph-derived context, attaching the result to the response payload for downstream observability.

## Validation and Testing

Unit tests in [`tests/test_context_savings.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_context_savings.py) verify:

- Correct token estimation for strings and objects
- Handling of `None` inputs and empty contexts
- Percentage calculation edge cases (zero baseline, equal inputs)
- Round-trip serialization behavior

These tests ensure the 4-characters-per-token heuristic behaves predictably across data types and boundary conditions.

## Summary

- **`estimate_context_savings`** in [`code_review_graph/context_savings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/context_savings.py) is the sole source of token reduction calculations
- **4 characters per token** heuristic provides model-agnostic estimates via `estimate_tokens`
- **Two input modes**: explicit token counts or context strings requiring estimation
- **Non-negative savings guarantee** via `max(0, baseline - returned)`
- **`None` return** when baseline cannot be determined
- **Conservative by design**: slight overestimation ensures reported savings are achievable

## Frequently Asked Questions

### Why does the tool use 4 characters per token instead of a real tokenizer?

**`estimate_tokens` uses `CHARS_PER_TOKEN = 4` as a deliberately conservative, model-agnostic approximation.** Real tokenizers vary significantly between models—GPT-4, Claude, and Llama families all use different encoding schemes. The 4-character heuristic provides a universal lower bound that works across any deployment without adding dependencies or model-specific configuration.

### Can I supply my own token counts instead of using the heuristic?

**Yes, both `original_tokens` and `returned_tokens` parameters accept explicit integers that bypass estimation entirely.** Pass values from your preferred tokenizer (e.g., OpenAI's `tiktoken`, Anthropic's tokenizer API) to replace the heuristic with precise counts.

### What happens when the returned context is larger than the original?

**The calculation floors at zero: `saved = max(0, baseline - returned)` ensures `saved_tokens` never reports negative values.** When returned exceeds baseline, you receive `{'saved_tokens': 0, 'saved_percent': 0}` rather than negative numbers. This reflects practical reality—no tokens were "saved" in this scenario.

### Where is the context savings estimate actually displayed to users?

**The `attach_context_savings` helper embeds estimates into tool results, while `format_context_savings` renders them for CLI output.** The primary consumer is `get_minimal_context` in [`code_review_graph/tools/context.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/context.py), which attaches savings metadata to every graph-derived context response.