How Caveman Reduces AI Agent Output Tokens: Architecture and Implementation

Caveman reduces AI agent output tokens by injecting a prompt template that instructs the model to eliminate natural-language filler while preserving code, URLs, and command snippets exactly, achieving an average 65% reduction without post-processing overhead.

Caveman is an open-source skill (or plugin) designed for AI-coding agents that minimizes token usage by rewriting responses before they reach the user. According to the JuliusBrussee/caveman repository, the tool achieves this by manipulating the model's generation parameters rather than applying external filters. This guide explains how Caveman reduces AI agent output tokens through its architectural components and prompt-level compression strategies.

How Caveman Compresses Output at the Prompt Level

Unlike traditional token-reduction methods that post-process generated text, Caveman operates at the generation layer. The skill injects a specific prompt template into the host agent that constrains how the model formulates its response.

The Prompt-Level Approach

In dist/caveman.skill, the prompt template explicitly instructs the model to "drop filler, keep substance, use fragments." This means the compression happens during the LLM's generation step, not after. The model receives constraints that prohibit verbose connective phrasing ("the reason ... is that ...", "I would recommend ...") while maintaining exact byte-for-byte preservation of code blocks, error messages, and URLs.

Because the model itself produces the trimmed output, no additional post-processing overhead is required. The resulting token count is genuinely lower, not merely filtered after the fact.

Preserving Substance vs. Eliminating Filler

Caveman distinguishes between two types of content:

  • Substance: Code snippets, command-line instructions, error messages, URLs, and file paths. These remain unchanged.
  • Filler: Natural-language verbosity, explanatory transitions, and conversational padding. These are compressed or eliminated.

This selective approach ensures that functional content retains its utility while conversational overhead is minimized.

Architectural Components

The Caveman ecosystem consists of several interconnected components that handle activation, persistence, and measurement of token savings.

Skill File and Activation

The caveman.skill file declares the prompt template that tells the host agent to apply compression rules. When installed, the skill automatically activates through src/hooks/caveman-activate.js, which writes a tiny flag file (.caveman) to signal the agent that the skill is active from the first message.

This eliminates the need for explicit /caveman commands on every session start in supported agents like Claude Code.

Mode Tracking and Compression Levels

The src/hooks/caveman-mode-tracker.js component persists the chosen compression level across a session. Caveman offers four distinct levels that determine how aggressively filler is trimmed:

  • lite: Minimal trimming for readability
  • full: Standard compression (default)
  • ultra: Aggressive fragment-based responses
  • wenyan: Classical Chinese-style extreme brevity

Each level modifies the prompt constraints to produce progressively terser output while maintaining the same substance preservation rules.

Statistics and Status Line

After each reply, src/hooks/caveman-stats.js parses the host agent's log file, counts tokens before and after compression, and updates a statistics file. The src/hooks/caveman-statusline.sh script (and its PowerShell equivalent) reads this data to emit a concise banner displaying lifetime token savings, typically formatted as [CAVEMAN] ⛏ 12.4k.

Practical Usage Examples

Install Caveman using the one-line installer:

curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash

Verify activation by checking the current compression level:

/caveman

Switch between compression modes based on your needs:

/caveman lite      # Minimal trimming

/caveman ultra     # Aggressive, ultra-short fragments

/caveman wenyan    # Classical Chinese-style brevity

View cumulative token savings:

/caveman-stats

Sample output:


[CAVEMAN] ⛏ 12.4k  ← lifetime tokens saved so far

Compress static memory files permanently to reduce input tokens on every agent launch:

/caveman-compress CLAUDE.md

This command rewrites the specified file in "caveman-speak," cutting approximately 46% of its input tokens according to repository benchmarks.

Generate concise conventional commit messages:

/caveman-commit

This returns a commit message limited to ≤ 50 characters for the subject line.

Key Source Files

Understanding the implementation requires examining these specific files:

Summary

  • Caveman reduces AI agent output tokens by prompting the model to generate concise responses during the generation phase, eliminating the need for post-processing filters.
  • The architecture separates concerns into activation hooks (caveman-activate.js), mode tracking (caveman-mode-tracker.js), and statistics collection (caveman-stats.js).
  • Four compression levels (lite, full, ultra, wenyan) provide granular control over verbosity.
  • Code, URLs, and error messages remain byte-for-byte identical while natural-language filler is removed.
  • Real-time token savings display through status-line scripts provides immediate feedback on efficiency gains.

Frequently Asked Questions

How does Caveman differ from post-processing compression tools?

Caveman operates at the prompt level rather than filtering generated text. Traditional tools receive the full verbose output and then remove words, which still incurs the generation cost. Caveman instructs the model via the caveman.skill prompt template to never generate filler in the first place, resulting in genuinely fewer output tokens and no post-processing overhead.

What are the four compression levels in Caveman?

Caveman offers lite, full, ultra, and wenyan modes. The lite mode applies minimal trimming for readability, while full provides standard compression. ultra forces aggressive fragment-based responses, and wenyan produces extreme brevity styled after classical Chinese literary conventions. Users toggle these via the /caveman command.

Does Caveman modify code snippets or URLs?

No, Caveman preserves substance byte-for-byte. According to the prompt implementation in dist/caveman.skill, the model must keep code blocks, command-line snippets, error messages, and URLs exactly unchanged. Only natural-language filler connecting these elements is subject to compression.

How can I measure token savings with Caveman?

Use the /caveman-stats command or check your status line. The caveman-stats.js hook parses agent logs after each reply, calculates the difference between original and compressed token counts, and updates a JSON file. The caveman-statusline.sh script reads this data to display lifetime savings (e.g., [CAVEMAN] ⛏ 12.4k) in your agent's interface.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →