# How the Caveman Skill Compresses Agent Output: A Deep Dive into the Compression Pipeline

> Discover how the Caveman skill compresses agent output. Learn about its pipeline: prose detection, front-matter stripping, prompt engineering, and validation for efficient markdown.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-08-22

---

**The caveman-compress skill rewrites LLM-generated markdown into compact caveman format by detecting compressible prose, stripping front-matter, prompting Claude with strict safety rules, and validating the output through a retry loop before atomically writing the result.**

The `JuliusBrussee/caveman` repository implements a specialized compression skill that transforms verbose agent outputs into concise, LLM-friendly caveman notation without breaking code blocks, URLs, or structural elements. This article examines the exact implementation details found in [`skills/caveman-compress/scripts/compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/compress.py) and its supporting modules to explain how the skill safely compresses agent output while maintaining data integrity.

## Overview of the Compression Architecture

The compression workflow follows a defensive, eight-step pipeline designed to prevent data loss and accidental secret leakage. Each step is implemented as a discrete function within [`compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/compress.py), orchestrated by the main `compress_file` entry point. The system treats every file operation as potentially volatile, employing atomic writes and backup verification to ensure that failures leave the original document untouched.

## Step 1: Detecting Compressible Content

Before processing begins, the skill must determine whether a file qualifies for compression. The `detect.should_compress` function in [`skills/caveman-compress/scripts/detect.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/detect.py) applies heuristics to filter out non-prose content, including:

- **Binary files** and **executables**
- **Source code files** that would lose meaning if compressed
- **Files exceeding 500KB** (hard size limit to prevent token overflow)
- **Empty or whitespace-only files**
- **Sensitive paths** (`.env`, `*.pem`, certificates) rejected by `is_sensitive_path`

Only natural-language markdown files pass this gate, ensuring Claude receives appropriate content for compression.

## Step 2: Safe File Handling and Backup Creation

Once a file passes detection, the `read_source` function reads the content **exactly** as UTF-8 bytes. The skill immediately creates a backup copy with the extension `*.original.md` in a platform-aware out-of-tree directory (`~/.local/share/caveman-compress/backups/`). 

This backup strategy ensures the skill’s auto-loader never re-ingests the backup file, preventing infinite compression loops. If any subsequent step fails, the original file restores from this backup, guaranteeing no data loss occurs during the transformation.

## Step 3: Preserving Metadata with Front-Matter Extraction

The `split_frontmatter` function (lines 37-45 in [`compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/compress.py)) isolates YAML front-matter from the document body. Because LLMs tend to rewrite or corrupt metadata blocks during compression, the skill strips front-matter before sending content to Claude and reattaches it verbatim to the final output. This preserves critical metadata like titles, dates, and tags without trusting the language model to maintain their accuracy.

## Step 4: Constructing the Compression Prompt

The `build_compress_prompt` function constructs a strict instructional prompt that defines exactly what may and may not change during compression:

- **Allowed changes**: Natural language prose, fluff removal, verbosity reduction
- **Protected elements**: Code fences, inline code, URLs, heading structure, list markers, and tables

This prompt engineering ensures the caveman format remains functionally equivalent to the original while significantly reducing token count.

## Step 5: Claude API Integration and Response Processing

The `call_claude` function (lines 68-104) interfaces with Anthropic’s models through two pathways:

1. **Anthropic SDK**: Used when `ANTHROPIC_API_KEY` environment variable is set
2. **Claude CLI fallback**: Direct subprocess calls to the local `claude` binary

After receiving the response, `strip_llm_wrapper` removes any outer markdown fences the model might have added, ensuring clean text extraction before validation begins.

## Step 6: Validation and Automated Fixing

The validation layer checks structural integrity through the `validate` function. If validation fails—indicating the model altered protected elements like code fences or URLs—the skill enters a fix loop:

- `build_fix_prompt` generates a correction prompt based on specific validation errors
- The corrected content resubmits to Claude
- This retry cycle executes up to `MAX_RETRIES` (default: 2) before aborting

If validation fails after exhausting retries, the operation aborts and restores the original from backup.

## Step 7: Atomic Write and Failure Recovery

Final output writes occur through `write_text_atomic` and `write_bytes_atomic` (internal `_write_target` helper). These functions ensure that readers never encounter partial writes by:

1. Writing to a temporary file on the same filesystem
2. Performing an atomic rename operation to the target path
3. Verifying backup integrity before removing the backup file

If any write operation fails or the verification check detects corruption, the skill automatically restores the original file from the `*.original.md` backup.

## Safety Guardrails and Secret Protection

The compression skill implements multiple defensive layers beyond basic file handling:

- **Secret path detection**: Patterns matching `.env`, `*.pem`, `id_rsa`, and certificate files trigger immediate rejection via `is_sensitive_path` (lines 102-113)
- **Backup verification**: Post-write hash comparison ensures the backup matches the original before cleanup
- **UTF-8 enforcement**: Invalid encoding detection aborts processing before sending to external APIs
- **Size limits**: Hard 500KB cap prevents accidental token cost explosions and timeout scenarios

## How to Use the Compress Skill

### Command-Line Usage

Execute compression directly against markdown files:

```bash

# Compress a single markdown file

python -m caveman.skills.caveman-compress.scripts.compress /path/to/file.md

```

The script creates backups at `~/.local/share/caveman-compress/backups/<parent-dir>/file.original.md` and replaces the target with the compressed version, printing progress to stdout.

### Python API Integration

Import the compression logic directly into agent workflows:

```python
from pathlib import Path
from skills.caveman_compress.scripts.compress import compress_file

# Path to the file you want to compress

target = Path("docs/long-story.md")

# Returns True on success, False if the file was not compressible

if compress_file(target):
    print("✅ Compression succeeded")
else:
    print("⚠️ No changes made")

```

### Advanced: Custom Prompt Construction

For debugging or custom implementations, reuse the internal prompt builders:

```python
from skills.caveman_compress.scripts.compress import build_compress_prompt, call_claude

original_md = Path("docs/notes.md").read_text()
prompt = build_compress_prompt(original_md)
compressed_body = call_claude(prompt)
print(compressed_body)   # Raw compressed markdown without outer fence

```

## Summary

- The **caveman-compress** skill selectively processes natural-language markdown while rejecting code, binaries, and secrets via `detect.should_compress` and `is_sensitive_path`.
- **Atomic write operations** and **out-of-tree backups** ensure zero data loss during the compression pipeline.
- **Front-matter extraction** preserves YAML metadata by processing it separately from the compressible body content.
- **Strict prompt engineering** in `build_compress_prompt` protects code blocks, URLs, and structural elements from modification.
- **Automated validation** with a retry loop guarantees output integrity before final writes occur.
- **Dual API support** allows operation through either the Anthropic SDK or local Claude CLI tooling.

## Frequently Asked Questions

### What file types does the Caveman compress skill support?

The skill targets markdown files containing natural language prose. It explicitly excludes source code, binary formats, files over 500KB, and sensitive paths (like `.env` or `*.pem` files) according to the heuristics in [`skills/caveman-compress/scripts/detect.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/detect.py). The `should_compress` function validates UTF-8 encoding and content type before processing begins.

### How does the skill prevent data loss during compression?

The implementation uses defensive programming with three recovery mechanisms: **atomic file writes** via `write_text_atomic` that prevent partial file corruption, **automatic backup restoration** if validation fails or writes error out, and **backup verification** checks that ensure the `*.original.md` copy matches the source file before any destructive operations occur.

### Can I adjust the number of validation retries?

The skill defines `MAX_RETRIES` with a default value of 2 attempts for fixing validation errors. While the current implementation uses this constant, advanced users can modify the value in [`skills/caveman-compress/scripts/compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/compress.py) before the validation loop at lines 61-86, though this requires editing the source code directly as no CLI flag currently exposes this parameter.

### Does the compression skill work without an Anthropic API key?

Yes. The `call_claude` function checks for the `ANTHROPIC_API_KEY` environment variable and falls back to the local `claude` CLI binary if the SDK credentials are unavailable. This dual-path design ensures the skill functions in environments with either cloud API access or a locally installed Claude command-line tool.