/caveman‑compress vs Regular Caveman Mode: Token Savings Explained
Regular Caveman mode reduces output tokens by roughly 65% by rewriting the model's replies in terse "caveman-speak," while /caveman-compress reduces input tokens by ~46% by permanently shrinking memory files like CLAUDE.md before the model reads them.
The JuliusBrussee/caveman repository provides two complementary mechanisms to slash token costs in AI-assisted workflows. Understanding the difference between /caveman-compress and regular Caveman mode helps you target the right token bottleneck—whether you are paying for verbose responses or bloated context windows.
How Regular Caveman Mode Cuts Output Tokens
Regular Caveman mode targets the text the model writes in every reply. When enabled, the system intercepts the assistant's response and rewrites it in a terse, minimal style before sending it back to you.
What It Trims
This mode drops filler words, elaborate explanations, and decorative formatting while preserving code blocks, URLs, and shell commands intact. The model still sees your full prompt and thinks through the complete answer, but the final output is stripped to essentials.
Implementation Details
In src/plugins/opencode/commands/caveman.md, the hook wraps each assistant reply with a style-enforcing prompt. The original verbose reply is never stored or transmitted; only the compressed version reaches your screen. This happens dynamically on every turn, regardless of which files are loaded in the context.
Typical Savings
According to the benchmark tables in the main README.md, regular Caveman mode achieves approximately 65% output token reduction per reply when using the default "full" setting.
# Enable regular Caveman mode (default is "full")
/caveman
# or explicitly
/caveman full
How /caveman‑compress Cuts Input Tokens
While regular mode trims what the model writes, /caveman-compress attacks what the model reads. It permanently rewrites persistent memory files—such as CLAUDE.md—into the same terse caveman style, reducing the token count of your system context.
The Compression Workflow
The script skills/caveman-compress/scripts/compress.py handles the transformation. When you run /caveman-compress <file>, the script:
- Reads the target file and strips YAML front‑matter
- Sends the body to Claude via the
call_claude()function - Validates the result using the
validate()function - Writes a backup with the
*.original.mdsuffix - Overwrites the original with the compressed content
This process runs a retry loop on validation failures and includes safety checks to protect sensitive file paths (documented in SECURITY.md).
Persistent Savings Across Sessions
Once compressed, the file remains smaller permanently. Future sessions load the compressed version automatically, meaning the model reads ~46% fewer tokens every time that file is referenced. This compounds significantly for long‑lived memory files that persist across dozens of conversations.
Typical Savings
As documented in docs/HONEST-NUMBERS.md and the "caveman‑compress receipts" table in the README, /caveman-compress achieves approximately 46% average input reduction per compressed file.
# Compress a memory file once for permanent savings
/caveman-compress CLAUDE.md
# Creates CLAUDE.original.md as backup
Key Differences at a Glance
| Aspect | Regular Caveman Mode | /caveman‑compress |
|---|---|---|
| Token target | Output (what the model writes) | Input (what the model reads) |
| Scope | Every reply in the session | Specific files you designate |
| Persistence | Temporary (per reply) | Permanent (until you restore from backup) |
| Implementation | Hook in src/plugins/opencode/commands/caveman.md |
Script in skills/caveman-compress/scripts/compress.py |
| Typical savings | ~65% fewer output tokens | ~46% fewer input tokens per file |
Using Both Modes Together for Maximum Savings
For the biggest token‑budget win, apply both strategies simultaneously. First, compress your long‑lived memory files with /caveman-compress to shrink your permanent context. Then enable regular Caveman mode to ensure every subsequent reply remains terse.
This dual approach minimizes both the input cost of loading your project context and the output cost of each turn, creating a compounding effect on your overall token expenditure.
Summary
- Regular Caveman mode intercepts and rewrites assistant replies at runtime, cutting output tokens by ~65% per turn.
/caveman-compresspermanently shrinks memory files before the model reads them, saving ~46% on input tokens for every future session that loads those files.- The前者 is implemented in
src/plugins/opencode/commands/caveman.md, while the latter runs viaskills/caveman-compress/scripts/compress.pyusingcall_claude()andvalidate()functions. - Combine both methods to optimize both sides of the token equation: what the model reads and what it writes.
Frequently Asked Questions
Can I use both regular Caveman mode and /caveman-compress together?
Yes. These tools target different token streams and work independently. Compress your memory files once with /caveman-compress to reduce input costs, then keep regular Caveman mode enabled to trim every reply. This layering provides the most aggressive token savings.
Is the compression reversible if I need the original file back?
Yes. The compress.py script automatically creates a backup with the *.original.md suffix before overwriting the original. If you need to restore the verbose version, simply rename the backup file or run the compression command again to regenerate the compressed version from the restored original.
Does regular Caveman mode affect my project files or memory files?
No. Regular Caveman mode only transforms the assistant's output before it reaches your screen. It never modifies your source code, documentation, or memory files like CLAUDE.md. Only /caveman-compress writes changes to disk, and it only does so for the specific files you explicitly target.
Which files should I prioritize for /caveman-compress?
Prioritize long, static memory files that load at the start of every session, such as CLAUDE.md, CONTEXT.md, or similar project documentation. These files are read repeatedly but change infrequently, making the ~46% input token savings compound across many conversations. Avoid compressing files that change frequently or contain sensitive credentials, as the script includes specific protections for such paths.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →