# How to Measure Real Session Token Savings vs the Advertised 65% Reduction in Caveman

> Discover how to accurately measure Caveman session token savings. Compare raw API token counts to advertised reductions, accounting for input overhead, to find honest net savings. Learn more.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: how-to-guide
- Published: 2026-07-08

---

**Use the `/caveman-stats` command to compare raw API token counts against the estimated 65% output reduction, then subtract the ~1,000–1,500 input-token overhead the skill injects per turn to determine your real net savings.**

Caveman (JuliusBrussee/caveman) advertises an approximate **65% reduction** in output tokens compared to vanilla model responses. However, to measure real session token savings accurately, you must account for the skill's own input costs—approximately 1,000–1,500 tokens per turn—that offset these gains. The repository provides built-in tooling and benchmark scripts to expose these "honest numbers" and calculate the true impact on your specific workload.

## Where the Advertised 65% Originates

The headline **≈ 65 %** figure (range 22‑87 %) comes from a controlled benchmark that measures **output tokens only** on a fixed set of ten representative prompts. According to the implementation in [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py), the script calls the Claude API twice—once with Caveman enabled and once without—then compares the resulting token counts.

| Metric | Value | Source |
|--------|-------|--------|
| Output reduction vs. vanilla | **≈ 65 %** | [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py) |
| Input reduction from skill | **0 %** | [`docs/HONEST-NUMBERS.md`](https://github.com/JuliusBrussee/caveman/blob/main/docs/HONEST-NUMBERS.md) |
| Extra input per turn | **≈ 1–1.5 k tokens** | [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md) |
| Persistent file compression | **≈ 46 %** per file | [`README.md`](https://github.com/JuliusBrussee/caveman/blob/main/README.md) benchmarks |

These figures are "honest" because they isolate **output** savings from **input** costs. The skill never compresses your prompt or context; it only changes the style of the reply by injecting a system prompt (~5 KB) plus skill-list entries.

## The Hidden Input Cost Per Turn

Every turn with Caveman enabled appends the skill's system prompt to the model request. As documented in [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md), this injection costs approximately **1,000 to 1,500 input tokens** per turn depending on the agent configuration.

When your outputs are short (≈ 150 tokens), this overhead can make net savings **negative**. In heavy-input workloads where context dominates the budget, the constant overhead may dwarf the 65% output cut, resulting in total session savings of only **14‑21 %**.

## How to Measure Real Session Token Savings

To determine your actual savings, you must extract raw token counts from the provider's API responses and subtract the skill's overhead.

### Using /caveman-stats to View Raw Token Counts

When Caveman is installed, it writes a session flag (`caveman.flag`) and appends the system prompt to each request. The `/caveman-stats` command reads the session log ([`.caveman/log.json`](https://github.com/JuliusBrussee/caveman/blob/main/.caveman/log.json)) and extracts the provider-reported input and output counters.

```bash

# In any supported agent (Claude Code, Cursor, etc.)

/caveman-stats

```

Typical output:

```

[ CAVEMAN ] Session token usage:
  Input  : 12,842 tokens
  Output :  8,123 tokens
  Saved  : 5,210 tokens (≈ 65% est.)   ← based on benchmark ratio
  Net    : –1,000 tokens (overhead)    ← real impact

```

The `Saved` line estimates what the reply would have been without Caveman using the benchmark ratio. The `Net` line shows the true balance after subtracting the input overhead. See the implementation in [`src/tools/caveman-stats.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/tools/caveman-stats.js) for the calculation logic.

### Running Reproducible Benchmarks

To verify the headline percentage on your own infrastructure:

```bash

# Install dependencies (requires Anthropic API key)

pip install -r benchmarks/requirements.txt

# Run the benchmark suite

python benchmarks/run.py

```

This script prints raw token counts for each prompt and the resulting percentage reduction, matching the table in the README.

### Conducting A/B Verification Tests

The only fully honest test is to run the same task twice—once with Caveman enabled and once without—then compare the provider's usage page. The repository provides a reproducible harness for this:

```bash

# Export your session log to JSON

/caveman-export-log > mysession.json

# Compute net savings with the helper script

python evals/measure.py --log mysession.json

```

[`evals/measure.py`](https://github.com/JuliusBrussee/caveman/blob/main/evals/measure.py) parses the log, subtracts the skill's overhead, and reports the final net token delta.

## When the Headline Number Misleads

Real-world conditions often diverge from the benchmark average:

- **Short replies** (≈ 150 tokens): The ~1k input overhead exceeds output savings, resulting in net **negative** savings.
- **Per-request billing** (e.g., Copilot): Token count is irrelevant; shorter replies do not reduce the number of requests.
- **Heavy-input workloads**: When prompts and context dominate, the 65% output cut may be dwarfed by the constant input overhead, yielding only **14‑21 %** total session savings.

These limitations are documented in [`docs/HONEST-NUMBERS.md`](https://github.com/JuliusBrussee/caveman/blob/main/docs/HONEST-NUMBERS.md) and flagged in the repository's issue tracker.

## Summary

- **The 65% claim** applies only to output tokens measured on fixed benchmarks in [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py).
- **Input costs** are constant at ~1,000–1,500 tokens per turn, as defined in [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md).
- **Real measurement** requires `/caveman-stats` to view raw API counters, or [`evals/measure.py`](https://github.com/JuliusBrussee/caveman/blob/main/evals/measure.py) for offline analysis.
- **Net savings** depend on reply length; short outputs may result in higher total costs despite the output reduction.

## Frequently Asked Questions

### Why do advertised savings differ from real session measurements?

The advertised **≈ 65 %** measures only output tokens on a curated prompt set. Real sessions include both output *and* input tokens, and Caveman injects ~1,000–1,500 input tokens per turn that are not present in the benchmark calculation. This overhead is documented in [`docs/HONEST-NUMBERS.md`](https://github.com/JuliusBrussee/caveman/blob/main/docs/HONEST-NUMBERS.md) and [`skills/caveman/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman/SKILL.md).

### How does the /caveman-stats command calculate net savings?

The command reads [`.caveman/log.json`](https://github.com/JuliusBrussee/caveman/blob/main/.caveman/log.json) to extract provider-reported counts, then computes `saved_output = output_without_caveman * (1 - reduction_ratio)` using the benchmark ratio. It subtracts the known input overhead to derive `net_savings`. See the logic in [`src/tools/caveman-stats.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/tools/caveman-stats.js).

### Can Caveman reduce costs with per-request billing models?

No. Per-request billing models (such as GitHub Copilot) charge per API call regardless of token count. Since Caveman reduces output length but not the number of requests, it provides no cost savings on these platforms. The token-reduction benefits only apply to usage-based billing models.

### What persistent savings does /caveman-compress provide?

The `/caveman-compress` command reduces memory file sizes by approximately **46 %** per file per session. This creates a persistent input reduction for subsequent turns, partially offsetting the per-turn overhead. This metric is separate from the headline 65% output reduction and is measured in the [`README.md`](https://github.com/JuliusBrussee/caveman/blob/main/README.md) compression table.