How Caveman Reduces AI Agent Output Tokens: Architecture and Implementation
Caveman reduces AI agent output tokens by injecting a prompt template that instructs the model to eliminate natural-language filler while preserving code, URLs, and command snippets exactly, achieving an average 65% reduction without post-processing overhead.
Caveman is an open-source skill (or plugin) designed for AI-coding agents that minimizes token usage by rewriting responses before they reach the user. According to the JuliusBrussee/caveman repository, the tool achieves this by manipulating the model's generation parameters rather than applying external filters. This guide explains how Caveman reduces AI agent output tokens through its architectural components and prompt-level compression strategies.
How Caveman Compresses Output at the Prompt Level
Unlike traditional token-reduction methods that post-process generated text, Caveman operates at the generation layer. The skill injects a specific prompt template into the host agent that constrains how the model formulates its response.
The Prompt-Level Approach
In dist/caveman.skill, the prompt template explicitly instructs the model to "drop filler, keep substance, use fragments." This means the compression happens during the LLM's generation step, not after. The model receives constraints that prohibit verbose connective phrasing ("the reason ... is that ...", "I would recommend ...") while maintaining exact byte-for-byte preservation of code blocks, error messages, and URLs.
Because the model itself produces the trimmed output, no additional post-processing overhead is required. The resulting token count is genuinely lower, not merely filtered after the fact.
Preserving Substance vs. Eliminating Filler
Caveman distinguishes between two types of content:
- Substance: Code snippets, command-line instructions, error messages, URLs, and file paths. These remain unchanged.
- Filler: Natural-language verbosity, explanatory transitions, and conversational padding. These are compressed or eliminated.
This selective approach ensures that functional content retains its utility while conversational overhead is minimized.
Architectural Components
The Caveman ecosystem consists of several interconnected components that handle activation, persistence, and measurement of token savings.
Skill File and Activation
The caveman.skill file declares the prompt template that tells the host agent to apply compression rules. When installed, the skill automatically activates through src/hooks/caveman-activate.js, which writes a tiny flag file (.caveman) to signal the agent that the skill is active from the first message.
This eliminates the need for explicit /caveman commands on every session start in supported agents like Claude Code.
Mode Tracking and Compression Levels
The src/hooks/caveman-mode-tracker.js component persists the chosen compression level across a session. Caveman offers four distinct levels that determine how aggressively filler is trimmed:
- lite: Minimal trimming for readability
- full: Standard compression (default)
- ultra: Aggressive fragment-based responses
- wenyan: Classical Chinese-style extreme brevity
Each level modifies the prompt constraints to produce progressively terser output while maintaining the same substance preservation rules.
Statistics and Status Line
After each reply, src/hooks/caveman-stats.js parses the host agent's log file, counts tokens before and after compression, and updates a statistics file. The src/hooks/caveman-statusline.sh script (and its PowerShell equivalent) reads this data to emit a concise banner displaying lifetime token savings, typically formatted as [CAVEMAN] ⛏ 12.4k.
Practical Usage Examples
Install Caveman using the one-line installer:
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash
Verify activation by checking the current compression level:
/caveman
Switch between compression modes based on your needs:
/caveman lite # Minimal trimming
/caveman ultra # Aggressive, ultra-short fragments
/caveman wenyan # Classical Chinese-style brevity
View cumulative token savings:
/caveman-stats
Sample output:
[CAVEMAN] ⛏ 12.4k ← lifetime tokens saved so far
Compress static memory files permanently to reduce input tokens on every agent launch:
/caveman-compress CLAUDE.md
This command rewrites the specified file in "caveman-speak," cutting approximately 46% of its input tokens according to repository benchmarks.
Generate concise conventional commit messages:
/caveman-commit
This returns a commit message limited to ≤ 50 characters for the subject line.
Key Source Files
Understanding the implementation requires examining these specific files:
dist/caveman.skill: Contains the serialized skill with the compression prompt template that injects constraints into the host agent.src/hooks/caveman-activate.js: Handles automatic activation by writing the.cavemanflag file.src/hooks/caveman-mode-tracker.js: Persists selected compression levels across sessions.src/hooks/caveman-stats.js: Computes token savings by parsing agent logs and writing statistics.src/hooks/caveman-statusline.sh: Shell wrapper that generates the status banner for UI display.src/plugins/opencode/commands/caveman.md: Defines the/cavemanslash command interface.src/plugins/opencode/commands/caveman-compress.md: Implements the/caveman-compresscommand for static file optimization.
Summary
- Caveman reduces AI agent output tokens by prompting the model to generate concise responses during the generation phase, eliminating the need for post-processing filters.
- The architecture separates concerns into activation hooks (
caveman-activate.js), mode tracking (caveman-mode-tracker.js), and statistics collection (caveman-stats.js). - Four compression levels (
lite,full,ultra,wenyan) provide granular control over verbosity. - Code, URLs, and error messages remain byte-for-byte identical while natural-language filler is removed.
- Real-time token savings display through status-line scripts provides immediate feedback on efficiency gains.
Frequently Asked Questions
How does Caveman differ from post-processing compression tools?
Caveman operates at the prompt level rather than filtering generated text. Traditional tools receive the full verbose output and then remove words, which still incurs the generation cost. Caveman instructs the model via the caveman.skill prompt template to never generate filler in the first place, resulting in genuinely fewer output tokens and no post-processing overhead.
What are the four compression levels in Caveman?
Caveman offers lite, full, ultra, and wenyan modes. The lite mode applies minimal trimming for readability, while full provides standard compression. ultra forces aggressive fragment-based responses, and wenyan produces extreme brevity styled after classical Chinese literary conventions. Users toggle these via the /caveman command.
Does Caveman modify code snippets or URLs?
No, Caveman preserves substance byte-for-byte. According to the prompt implementation in dist/caveman.skill, the model must keep code blocks, command-line snippets, error messages, and URLs exactly unchanged. Only natural-language filler connecting these elements is subject to compression.
How can I measure token savings with Caveman?
Use the /caveman-stats command or check your status line. The caveman-stats.js hook parses agent logs after each reply, calculates the difference between original and compressed token counts, and updates a JSON file. The caveman-statusline.sh script reads this data to display lifetime savings (e.g., [CAVEMAN] ⛏ 12.4k) in your agent's interface.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →