Does Caveman Compress Input Tokens or Only Output Tokens?

Caveman compresses input tokens, shrinking the project files that Claude reads into its context window, while leaving Claude's output token generation unchanged.

The caveman repository by JuliusBrussee provides a preprocessing skill designed to reduce token costs in Claude workflows. Understanding whether Caveman compresses input tokens or only output tokens is essential for calculating cost savings and optimizing context window usage.

How Caveman Targets Input Tokens

Caveman-compress specifically reduces input token consumption by preprocessing files before Claude reads them. According to the source code in skills/caveman-compress/scripts/compress.py, the skill consumes tokens only during the compression phase, not when Claude generates responses.

The implementation follows a strict pipeline where token usage is isolated to specific LLM calls:

  • Detection (detect.py): Determines file compressibility using local Python—zero tokens
  • Compression (compress.py): Calls Claude once via call_claude() to shrink file bodies—token consuming
  • Validation (validate.py): Checks structural integrity (headings, code blocks, URLs) locally—zero tokens
  • Targeted fixes: If validation fails, invokes Claude only for specific broken fragments—limited tokens

Token Consumption in the Compression Pipeline

The Single LLM Compression Call

In skills/caveman-compress/scripts/compress.py (lines 222-240), the compress_file() function triggers the only mandatory token-consuming operation:


# From skills/caveman-compress/scripts/compress.py

compressed_body = call_claude(compression_prompt, content)

This call_claude() invocation sends the file body to Claude for compression, receives the shortened version, and stores it for validation.

Local Validation and Backup Handling

After compression, the validate.py module validates the output structure without additional LLM calls. If validation passes, the system writes both the compressed file and the .original.md backup using standard file I/O operations—no tokens consumed.

Targeted Fix Strategy

If validation fails, the system does not recompress the entire file. Instead, it calls Claude only for the specific problematic sections:


# From skills/caveman-compress/scripts/compress.py logic

if not result.is_valid:
    fixed_body = call_claude(fix_prompt)  # Tokens consumed only for broken parts

Input Token Savings vs Output Token Usage

Caveman specifically optimizes input tokens. The README at skills/caveman-compress/README.md (lines 27-28) confirms: "Only two things use tokens: initial compression + targeted fix if validation fails."

This means:

  • Input tokens: Reduced every time Claude reads the compressed file during a session
  • Output tokens: Unaffected—Claude generates responses at its normal length
  • One-time cost: The initial compression consumes tokens, but subsequent reads of the compressed file save tokens repeatedly

Practical Code Examples

Compress a File from the Command Line

caveman-compress CLAUDE.md

This command detects the file type locally (no tokens), sends the markdown body to Claude for compression (token-consuming), and writes both CLAUDE.md (compressed) and CLAUDE.original.md (backup).

Using the Python API Directly

from pathlib import Path
from scripts.compress import compress_file

path = Path("docs/preferences.md")
success = compress_file(path)   # Returns True on successful compression

print("Compressed!" if success else "No change")

Only the compress_file() call triggers Claude (consuming tokens); everything else executes as local Python.

Handling Validation Failures

from scripts.validate import validate
from scripts.compress import call_claude

# Validation check runs locally (no tokens)

result = validate(original_path, compressed_path)

if not result.is_valid:
    # Only problematic sections are sent again

    fixed_body = call_claude(fix_prompt)  # Limited token consumption

The fix step consumes tokens only for the problematic fragments, not for a full recompression.

Key Implementation Files

Summary

  • Caveman-compress reduces input tokens by preprocessing files before Claude reads them, not by limiting Claude's output.
  • Only two operations consume tokens: the initial call_claude() compression and targeted fixes if validation fails.
  • File detection, validation, and file I/O run as local Python with zero token usage.
  • The token savings compound over multiple sessions, as Claude repeatedly reads the compressed files.

Frequently Asked Questions

Does Caveman reduce Claude's output tokens?

No. Caveman only compresses the input context—the files that Claude reads. Claude's response generation and output token usage remain unchanged. The savings come from shrinking the context window that Claude must process before generating each response.

What happens if compression validation fails?

If the validate.py module detects structural issues (broken headings, malformed code blocks, or invalid URLs), the system invokes Claude only for targeted fixes on the specific problematic fragments. It does not consume tokens to recompress the entire file.

Is file type detection performed by Claude?

No. The detect.py module determines file compressibility using local Python heuristics. Detection runs entirely offline with zero token consumption before any LLM calls occur.

Where is the compression logic implemented?

The primary orchestration resides in skills/caveman-compress/scripts/compress.py, specifically within the compress_file() function (lines 222-240) and the call_claude() helper. This file manages the complete workflow from detection through validation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →