# Caveman Input vs Output Token Savings: How /caveman-compress and Regular Mode Reduce Costs

> Discover input token savings with /caveman-compress and output savings with regular mode. Learn how JuliusBrussee caveman reduces Claude API costs by compressing input and enforcing concise output.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: performance
- Published: 2026-07-08

---

**Input token savings from `/caveman-compress` reduce the text sent *to* Claude by compressing Markdown files before the API call, while output token savings from regular mode reduce the text returned *from* Claude by enforcing concise responses via system prompts.**

The Caveman repository by JuliusBrussee provides two distinct mechanisms for reducing API costs when working with Claude. Understanding the difference between these approaches helps you optimize your workflow depending on whether your bottlenecks are large input files or verbose LLM responses.

## Input Token Savings: Pre-Processing with /caveman-compress

The `/caveman-compress` skill operates **before** the LLM receives your content, reducing the token count of your input files.

### How Compression Works in compress.py

In [`skills/caveman-compress/scripts/compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/compress.py), the script reads your Markdown files and strips away natural-language prose while preserving essential structure such as headings, code blocks, and URLs. The `build_compress_prompt` function constructs a prompt that instructs Claude to rewrite the content in a condensed "caveman" format.

The process flow works as follows:

1. The script reads the original Markdown file
2. It removes YAML front-matter and sends only the body to Claude
3. Claude returns a compressed version with minimal verbosity
4. The original file is backed up to `~/.local/share/caveman-compress/backups`
5. The compressed content replaces the original file

Because the LLM never receives the uncompressed text, you save input tokens that would otherwise be consumed by verbose prose. This is particularly effective for large documentation files or knowledge bases where you want to maximize context window efficiency.

```bash

# Compress a Markdown file to reduce input tokens

python -m caveman-compress.scripts.compress path/to/file.md

```

## Output Token Savings: Regular Mode Benchmarking

Output token savings occur **after** the LLM generates content, measured by comparing response lengths between standard and caveman system prompts.

### How the Benchmark Measures Savings

The [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py) script quantifies output token savings by running each prompt in [`benchmarks/prompts.json`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/prompts.json) twice: once with `NORMAL_SYSTEM = "You are a helpful assistant."` and once using `load_caveman_system()` which loads the concise caveman system prompt.

The benchmark logic calculates savings using the formula:

```python
savings = 1 - (caveman_medians / normal_medians)

```

After recording `output_tokens` from each API response across multiple trials, the `compute_stats` function derives the median output token count for each mode. The resulting percentage shows how much shorter caveman responses are compared to standard responses.

```bash

# Run the benchmark to measure output token savings

python benchmarks/run.py --model claude-sonnet-4-20250514 --trials 3 --update-readme

```

The results are automatically injected into [`README.md`](https://github.com/JuliusBrussee/caveman/blob/main/README.md) between the `BENCHMARK-TABLE-START` and `BENCHMARK-TABLE-END` markers, providing a live dashboard of efficiency gains.

## Key Differences: Input vs Output Token Savings

| Aspect | Input Token Savings (`/caveman-compress`) | Output Token Savings (Regular Mode) |
|--------|-------------------------------------------|-------------------------------------|
| **When it operates** | Before the API call | After the API call |
| **What it reduces** | Size of the prompt sent to Claude | Size of the response from Claude |
| **Implementation** | [`compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/compress.py) rewrites files locally | [`run.py`](https://github.com/JuliusBrussee/caveman/blob/main/run.py) compares system prompt effects |
| **Measurement** | Implicit (smaller files = fewer input tokens) | Explicit (calculated percentage via `savings = 1 - (caveman_medians / normal_medians)`) |
| **Use case** | Large Markdown files, documentation compression | General conversation efficiency, API cost reduction |

Input savings require you to compress files beforehand using the caveman-compress skill, while output savings work transparently whenever you use the caveman system prompt, regardless of input size.

## Summary

- **Input token savings** from `/caveman-compress` work by preprocessing Markdown files in [`skills/caveman-compress/scripts/compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/compress.py) to remove verbose content before sending to Claude
- **Output token savings** from regular mode are measured in [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py) by comparing response lengths between standard and caveman system prompts using the formula `1 - (caveman_medians / normal_medians)`
- Input reduction happens **before** the API call (file compression), while output reduction happens **during** generation (concise instructions)
- Both mechanisms contribute to overall cost efficiency but target different sides of the token usage equation

## Frequently Asked Questions

### How do I know if I need input or output token savings?

If you are processing large Markdown files or documentation where the input context exceeds your needs, use `/caveman-compress` to reduce file sizes before sending them to Claude. If your inputs are small but Claude's responses are unnecessarily verbose, rely on the regular caveman system prompt to generate output token savings. Many workflows benefit from combining both approaches.

### Can I use /caveman-compress and regular mode together?

Yes. You can compress your input files using `python -m caveman-compress.scripts.compress` to reduce input tokens, then query Claude using the caveman system prompt (loaded via `load_caveman_system()` in the benchmark) to ensure concise outputs. This stacks both savings mechanisms for maximum cost efficiency.

### Where does the benchmark store its output token measurements?

The [`benchmarks/run.py`](https://github.com/JuliusBrussee/caveman/blob/main/benchmarks/run.py) script records median output token counts for both normal and caveman modes, then writes the calculated savings percentage back into [`README.md`](https://github.com/JuliusBrussee/caveman/blob/main/README.md) between specific markers (`BENCHMARK-TABLE-START` and `BENCHMARK-TABLE-END`). You can view historical efficiency data directly in the repository's main documentation.

### Does file compression affect the quality of Claude's understanding?

According to the implementation in [`compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/compress.py), the compression preserves essential structural elements like headings, code blocks, and URLs while removing natural-language prose. This maintains the semantic content necessary for Claude to understand the material while eliminating filler text that consumes tokens without adding value.