# What Is the Average Token Saving with Caveman? The 65% Benchmark Explained

> Discover how Caveman achieves an average 65% token saving. Learn about its benchmark performance and understand the impressive compression rates for LLM outputs.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: performance
- Published: 2026-07-11

---

**Caveman reduces output tokens by approximately 65% on average** across a benchmark suite of ten representative prompts, with individual results ranging from 22% to 87% compression.

Caveman is an open-source AI skill designed to compress verbose language model responses into concise, token-efficient output. According to the `JuliusBrussee/caveman` repository, this **average token saving** of 65% is achieved through output-token compression, measured systematically via the benchmark script in [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py).

## How Caveman Achieves 65% Average Token Savings

The 65% figure represents **output-token compression**, not net session savings. The benchmark suite evaluates ten diverse prompts against Claude-3-Sonnet, measuring the difference between standard verbose responses and Caveman-compressed equivalents.

The calculation logic resides in [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py), specifically lines 42-49, where the script computes `avg_savings` from per-prompt token counts. The README (lines 54-70) displays these results in a markdown table highlighting the 65% average, while [`docs/HONEST-NUMBERS.md`](https://github.com/JuliusBrussee/caveman/blob/main/docs/HONEST-NUMBERS.md) (lines 11-14) reiterates this figure as the definitive average output-reduction metric.

### Understanding the Compression Range

Individual prompt results vary significantly:

- **Minimum**: 22% token reduction (highly concise base prompts)
- **Maximum**: 87% token reduction (verbose, repetitive content)
- **Standard**: 65% mean across the benchmark suite

This variance explains why Caveman excels with long-form content but provides diminishing returns on already terse queries.

## Calculating Token Savings in Practice

### Running the Benchmark Locally

Verify the 65% claim using the official benchmark suite against live API data:

```bash

# Install dependencies

pip install -r benchmarks/requirements.txt

# Execute benchmark (model = claude-3-sonnet-20240229, 3 trials per prompt)

python benchmarks/run.py claude-3-sonnet-20240229 3

```

The script outputs a markdown table concluding with:

```text
| **Average** | **1214** | **294** | **65%** |

```

The `avg_savings` value is calculated from actual API token counts reported by the Anthropic API.

### Querying Saved Percentages Programmatically

Access historical benchmark data programmatically:

```python
import json
from pathlib import Path

# Load the most recent benchmark result

results_path = next(Path("benchmarks/results").glob("benchmark_*.json"))
data = json.loads(results_path.read_text())

# Extract the average token saving

print(f"Average output token saving: {data['summary']['avg_savings']}%")

```

This reads JSON files generated by [`run.py`](https://github.com/JuliusBrussee/caveman/blob/main/run.py) containing the `summary` object with the `avg_savings` key.

### Using the Built-in Stats Command

During active Claude Code, Cursor, or Gemini sessions, invoke:

```bash
/caveman-stats

```

This command parses the session log and estimates savings using the established 65% benchmark ratio, outputting:

```

🦖 Saved ~65% output tokens (≈ 295 tokens) on this session.

```

## Input Token Overhead vs. Net Savings

While Caveman delivers **65% output token reduction**, it introduces approximately **1,000–1,500 input tokens per turn** (the skill's system prompt). Consequently, net session-wide savings depend on response length:

- **Short replies** (<500 tokens): Net savings may be negative due to overhead
- **Long replies** (>1000 tokens): Significant cost reduction achieved
- **Multi-turn sessions**: Compounding benefits as input overhead amortizes across many compressed outputs

The [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md) file defines the system prompt triggering this compression, while [`evals/measure.py`](https://github.com/JuliusBrussee/caveman/blob/main/evals/measure.py) provides offline evaluation capabilities for token-statistics analysis without API calls.

## Summary

- **Average token saving**: 65% (range 22%–87%) according to [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py)
- **Measurement method**: Benchmark suite of 10 prompts against Claude-3-Sonnet
- **Primary cost**: Adds ~1–1.5k input tokens per invocation
- **Key files**: [`README.md`](https://github.com/JuliusBrussee/caveman/blob/main/README.md) (public claims), [`docs/HONEST-NUMBERS.md`](https://github.com/JuliusBrussee/caveman/blob/main/docs/HONEST-NUMBERS.md) (methodology), [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py) (calculation logic)
- **Real-time monitoring**: Use `/caveman-stats` during active sessions

## Frequently Asked Questions

### What is the exact average token saving percentage for Caveman?

Caveman achieves exactly **65% average output token reduction** according to the benchmark calculations in [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py). This figure represents the mean across 10 diverse test prompts, with individual results varying from 22% to 87% depending on content verbosity.

### How does Caveman's token compression affect input costs?

Caveman adds roughly 1,000–1,500 input tokens per turn to load the compression skill. While this increases front-end costs, the 65% output reduction typically produces net savings for long-form responses exceeding 1,000 tokens, as documented in [`docs/HONEST-NUMBERS.md`](https://github.com/JuliusBrussee/caveman/blob/main/docs/HONEST-NUMBERS.md).

### Which files in the Caveman repository contain the benchmark logic?

The primary calculation logic resides in [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py) (lines 42-49), which computes `avg_savings`. Supporting evaluation exists in [`evals/measure.py`](https://github.com/JuliusBrussee/caveman/blob/main/evals/measure.py) for offline analysis. Documentation appears in [`README.md`](https://github.com/JuliusBrussee/caveman/blob/main/README.md) (lines 54-70) and [`docs/HONEST-NUMBERS.md`](https://github.com/JuliusBrussee/caveman/blob/main/docs/HONEST-NUMBERS.md) (lines 11-14).

### Can I verify the 65% token saving claim with my own prompts?

Yes. Execute `python benchmarks/run.py <model-name> <trials>` with your Anthropic API key to test custom prompts against baseline outputs. The script generates JSON results in `benchmarks/results/` containing per-prompt savings and the calculated average.