How to Measure Real Session Token Savings vs the Advertised 65% Reduction in Caveman
Use the /caveman-stats command to compare raw API token counts against the estimated 65% output reduction, then subtract the ~1,000–1,500 input-token overhead the skill injects per turn to determine your real net savings.
Caveman (JuliusBrussee/caveman) advertises an approximate 65% reduction in output tokens compared to vanilla model responses. However, to measure real session token savings accurately, you must account for the skill's own input costs—approximately 1,000–1,500 tokens per turn—that offset these gains. The repository provides built-in tooling and benchmark scripts to expose these "honest numbers" and calculate the true impact on your specific workload.
Where the Advertised 65% Originates
The headline ≈ 65 % figure (range 22‑87 %) comes from a controlled benchmark that measures output tokens only on a fixed set of ten representative prompts. According to the implementation in benchmarks/run.py, the script calls the Claude API twice—once with Caveman enabled and once without—then compares the resulting token counts.
| Metric | Value | Source |
|---|---|---|
| Output reduction vs. vanilla | ≈ 65 % | benchmarks/run.py |
| Input reduction from skill | 0 % | docs/HONEST-NUMBERS.md |
| Extra input per turn | ≈ 1–1.5 k tokens | skills/caveman/SKILL.md |
| Persistent file compression | ≈ 46 % per file | README.md benchmarks |
These figures are "honest" because they isolate output savings from input costs. The skill never compresses your prompt or context; it only changes the style of the reply by injecting a system prompt (~5 KB) plus skill-list entries.
The Hidden Input Cost Per Turn
Every turn with Caveman enabled appends the skill's system prompt to the model request. As documented in skills/caveman/SKILL.md, this injection costs approximately 1,000 to 1,500 input tokens per turn depending on the agent configuration.
When your outputs are short (≈ 150 tokens), this overhead can make net savings negative. In heavy-input workloads where context dominates the budget, the constant overhead may dwarf the 65% output cut, resulting in total session savings of only 14‑21 %.
How to Measure Real Session Token Savings
To determine your actual savings, you must extract raw token counts from the provider's API responses and subtract the skill's overhead.
Using /caveman-stats to View Raw Token Counts
When Caveman is installed, it writes a session flag (caveman.flag) and appends the system prompt to each request. The /caveman-stats command reads the session log (.caveman/log.json) and extracts the provider-reported input and output counters.
# In any supported agent (Claude Code, Cursor, etc.)
/caveman-stats
Typical output:
[ CAVEMAN ] Session token usage:
Input : 12,842 tokens
Output : 8,123 tokens
Saved : 5,210 tokens (≈ 65% est.) ← based on benchmark ratio
Net : –1,000 tokens (overhead) ← real impact
The Saved line estimates what the reply would have been without Caveman using the benchmark ratio. The Net line shows the true balance after subtracting the input overhead. See the implementation in src/tools/caveman-stats.js for the calculation logic.
Running Reproducible Benchmarks
To verify the headline percentage on your own infrastructure:
# Install dependencies (requires Anthropic API key)
pip install -r benchmarks/requirements.txt
# Run the benchmark suite
python benchmarks/run.py
This script prints raw token counts for each prompt and the resulting percentage reduction, matching the table in the README.
Conducting A/B Verification Tests
The only fully honest test is to run the same task twice—once with Caveman enabled and once without—then compare the provider's usage page. The repository provides a reproducible harness for this:
# Export your session log to JSON
/caveman-export-log > mysession.json
# Compute net savings with the helper script
python evals/measure.py --log mysession.json
evals/measure.py parses the log, subtracts the skill's overhead, and reports the final net token delta.
When the Headline Number Misleads
Real-world conditions often diverge from the benchmark average:
- Short replies (≈ 150 tokens): The ~1k input overhead exceeds output savings, resulting in net negative savings.
- Per-request billing (e.g., Copilot): Token count is irrelevant; shorter replies do not reduce the number of requests.
- Heavy-input workloads: When prompts and context dominate, the 65% output cut may be dwarfed by the constant input overhead, yielding only 14‑21 % total session savings.
These limitations are documented in docs/HONEST-NUMBERS.md and flagged in the repository's issue tracker.
Summary
- The 65% claim applies only to output tokens measured on fixed benchmarks in
benchmarks/run.py. - Input costs are constant at ~1,000–1,500 tokens per turn, as defined in
skills/caveman/SKILL.md. - Real measurement requires
/caveman-statsto view raw API counters, orevals/measure.pyfor offline analysis. - Net savings depend on reply length; short outputs may result in higher total costs despite the output reduction.
Frequently Asked Questions
Why do advertised savings differ from real session measurements?
The advertised ≈ 65 % measures only output tokens on a curated prompt set. Real sessions include both output and input tokens, and Caveman injects ~1,000–1,500 input tokens per turn that are not present in the benchmark calculation. This overhead is documented in docs/HONEST-NUMBERS.md and skills/caveman/SKILL.md.
How does the /caveman-stats command calculate net savings?
The command reads .caveman/log.json to extract provider-reported counts, then computes saved_output = output_without_caveman * (1 - reduction_ratio) using the benchmark ratio. It subtracts the known input overhead to derive net_savings. See the logic in src/tools/caveman-stats.js.
Can Caveman reduce costs with per-request billing models?
No. Per-request billing models (such as GitHub Copilot) charge per API call regardless of token count. Since Caveman reduces output length but not the number of requests, it provides no cost savings on these platforms. The token-reduction benefits only apply to usage-based billing models.
What persistent savings does /caveman-compress provide?
The /caveman-compress command reduces memory file sizes by approximately 46 % per file per session. This creates a persistent input reduction for subsequent turns, partially offsetting the per-turn overhead. This metric is separate from the headline 65% output reduction and is measured in the README.md compression table.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →