What Is the Average Token Saving with Caveman? The 65% Benchmark Explained

Caveman reduces output tokens by approximately 65% on average across a benchmark suite of ten representative prompts, with individual results ranging from 22% to 87% compression.

Caveman is an open-source AI skill designed to compress verbose language model responses into concise, token-efficient output. According to the JuliusBrussee/caveman repository, this average token saving of 65% is achieved through output-token compression, measured systematically via the benchmark script in benchmarks/run.py.

How Caveman Achieves 65% Average Token Savings

The 65% figure represents output-token compression, not net session savings. The benchmark suite evaluates ten diverse prompts against Claude-3-Sonnet, measuring the difference between standard verbose responses and Caveman-compressed equivalents.

The calculation logic resides in benchmarks/run.py, specifically lines 42-49, where the script computes avg_savings from per-prompt token counts. The README (lines 54-70) displays these results in a markdown table highlighting the 65% average, while docs/HONEST-NUMBERS.md (lines 11-14) reiterates this figure as the definitive average output-reduction metric.

Understanding the Compression Range

Individual prompt results vary significantly:

  • Minimum: 22% token reduction (highly concise base prompts)
  • Maximum: 87% token reduction (verbose, repetitive content)
  • Standard: 65% mean across the benchmark suite

This variance explains why Caveman excels with long-form content but provides diminishing returns on already terse queries.

Calculating Token Savings in Practice

Running the Benchmark Locally

Verify the 65% claim using the official benchmark suite against live API data:


# Install dependencies

pip install -r benchmarks/requirements.txt

# Execute benchmark (model = claude-3-sonnet-20240229, 3 trials per prompt)

python benchmarks/run.py claude-3-sonnet-20240229 3

The script outputs a markdown table concluding with:

| **Average** | **1214** | **294** | **65%** |

The avg_savings value is calculated from actual API token counts reported by the Anthropic API.

Querying Saved Percentages Programmatically

Access historical benchmark data programmatically:

import json
from pathlib import Path

# Load the most recent benchmark result

results_path = next(Path("benchmarks/results").glob("benchmark_*.json"))
data = json.loads(results_path.read_text())

# Extract the average token saving

print(f"Average output token saving: {data['summary']['avg_savings']}%")

This reads JSON files generated by run.py containing the summary object with the avg_savings key.

Using the Built-in Stats Command

During active Claude Code, Cursor, or Gemini sessions, invoke:

/caveman-stats

This command parses the session log and estimates savings using the established 65% benchmark ratio, outputting:


🦖 Saved ~65% output tokens (≈ 295 tokens) on this session.

Input Token Overhead vs. Net Savings

While Caveman delivers 65% output token reduction, it introduces approximately 1,000–1,500 input tokens per turn (the skill's system prompt). Consequently, net session-wide savings depend on response length:

  • Short replies (<500 tokens): Net savings may be negative due to overhead
  • Long replies (>1000 tokens): Significant cost reduction achieved
  • Multi-turn sessions: Compounding benefits as input overhead amortizes across many compressed outputs

The skills/caveman/SKILL.md file defines the system prompt triggering this compression, while evals/measure.py provides offline evaluation capabilities for token-statistics analysis without API calls.

Summary

  • Average token saving: 65% (range 22%–87%) according to benchmarks/run.py
  • Measurement method: Benchmark suite of 10 prompts against Claude-3-Sonnet
  • Primary cost: Adds ~1–1.5k input tokens per invocation
  • Key files: README.md (public claims), docs/HONEST-NUMBERS.md (methodology), benchmarks/run.py (calculation logic)
  • Real-time monitoring: Use /caveman-stats during active sessions

Frequently Asked Questions

What is the exact average token saving percentage for Caveman?

Caveman achieves exactly 65% average output token reduction according to the benchmark calculations in benchmarks/run.py. This figure represents the mean across 10 diverse test prompts, with individual results varying from 22% to 87% depending on content verbosity.

How does Caveman's token compression affect input costs?

Caveman adds roughly 1,000–1,500 input tokens per turn to load the compression skill. While this increases front-end costs, the 65% output reduction typically produces net savings for long-form responses exceeding 1,000 tokens, as documented in docs/HONEST-NUMBERS.md.

Which files in the Caveman repository contain the benchmark logic?

The primary calculation logic resides in benchmarks/run.py (lines 42-49), which computes avg_savings. Supporting evaluation exists in evals/measure.py for offline analysis. Documentation appears in README.md (lines 54-70) and docs/HONEST-NUMBERS.md (lines 11-14).

Can I verify the 65% token saving claim with my own prompts?

Yes. Execute python benchmarks/run.py <model-name> <trials> with your Anthropic API key to test custom prompts against baseline outputs. The script generates JSON results in benchmarks/results/ containing per-prompt savings and the calculated average.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →