# How to Use the caveman-compress Skill to Compress a File

> Learn to compress files with caveman-compress. Shrink markdown or text files into terse caveman style for LLM prompts, preserving all technical content. Run caveman-compress <filepath>.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: how-to-guide
- Published: 2026-08-22

---

**The `caveman-compress` skill shrinks natural-language markdown or text files into terse "caveman" style by running `caveman-compress <filepath>`, which preserves all technical content while reducing token count for LLM prompts.**

The **caveman-compress** skill is a specialized utility in the [JuliusBrussee/caveman](https://github.com/JuliusBrussee/caveman) repository designed to optimize documentation files for AI context windows. It processes eligible files through Claude-powered compression, removes filler words and redundant phrasing, and maintains semantic integrity.

## Command Syntax and Basic Usage

Invoke the skill directly from your terminal using the command defined in [`src/plugins/opencode/commands/caveman-compress.md`](https://github.com/JuliusBrussee/caveman/blob/main/src/plugins/opencode/commands/caveman-compress.md).

```bash
caveman-compress docs/README.md

```

The command accepts a single argument: the **absolute or relative filepath** to the target file. Upon execution, the system confirms success and returns a short confirmation message to the user.

## Step-by-Step Execution Flow

Understanding the internals helps troubleshoot edge cases and verify output quality.

### Triggering the CLI Command

When you execute `caveman-compress`, the command definition in [`src/plugins/opencode/commands/caveman-compress.md`](https://github.com/JuliusBrussee/caveman/blob/main/src/plugins/opencode/commands/caveman-compress.md) parses your input and delegates to the Python entry point. The system checks file eligibility immediately—only natural-language files (`.md`, `.txt`, `.typ`, `.tex`, or extension-less) are processed, while source code files (`.py`, `.js`, `.json`, etc.) and existing `*.original.md` backups are explicitly rejected.

### Script Execution and Claude Integration

The command runs the Python module located at [`plugins/caveman/skills/caveman-compress/scripts/__main__.py`](https://github.com/JuliusBrussee/caveman/blob/main/plugins/caveman/skills/caveman-compress/scripts/__main__.py) using the pattern:

```bash
python3 -m scripts <absolute_filepath>

```

This script performs three critical operations:
1. **File type detection** – Validates the extension against allowlists before processing.
2. **Claude API interaction** – Sends file contents to Claude for intelligent compression.
3. **Output validation** – Ensures the compressed text preserves **code blocks**, **inline code**, **URLs**, **file paths**, **commands**, and **markdown structure** exactly as documented in [`plugins/caveman/skills/caveman-compress/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/plugins/caveman/skills/caveman-compress/SKILL.md).

### Validation and Retry Logic

If the initial compression attempt fails validation—meaning technical elements were altered or the structure was corrupted—the system automatically retries up to **two additional times**. This robust handling ensures high reliability when processing complex documentation files found in `tests/caveman-compress/`.

### Backup Creation

Before overwriting the original file, the skill creates a human-readable backup at:

```

$XDG_DATA_HOME/caveman-compress/backups/<relative_path>/<filename>.original.md

```

This out-of-tree storage strategy prevents the backup from being re-ingested as a live file during subsequent compression runs. The original file is then replaced with the compressed version.

## File Type Requirements and Restrictions

The skill applies strict eligibility criteria to prevent corruption of executable code or configuration files.

**Eligible formats:**
- Markdown (`.md`)
- Plain text (`.txt`)
- Typst (`.typ`)
- LaTeX (`.tex`)
- Extension-less natural-language files

**Rejected formats:**
- Python (`.py`), JavaScript (`.js`), JSON (`.json`), and other source code or configuration files
- Existing backup files matching `*.original.md`

## Programmatic Usage

Integrate the skill into automation scripts using Python's `subprocess` module.

### Running Compression from Python

```python
import subprocess
import os

def compress_file(filepath: str) -> str:
    """Compress a file using the caveman-compress skill."""
    result = subprocess.run(
        ["caveman-compress", filepath],
        capture_output=True,
        text=True,
        check=True
    )
    return result.stdout

# Example usage

output = compress_file("notes/project-notes.md")
print(output)

```

### Verifying Backup Integrity

Confirm that the safety backup was created correctly by checking the XDG data directory:

```python
import pathlib
import os

def verify_backup(original_path: str) -> bool:
    """Check if backup exists in the XDG data home."""
    xdg_home = os.getenv("XDG_DATA_HOME", os.path.expanduser("~/.local/share"))
    backup_path = pathlib.Path(xdg_home) / "caveman-compress" / "backups"
    
    # Construct backup filepath based on original

    original = pathlib.Path(original_path)
    backup_file = backup_path / original.parent / f"{original.stem}.original.md"
    
    return backup_file.is_file()

# Verify specific backup

exists = verify_backup("notes/project-notes.md")
assert exists, "Backup file not found in XDG data home"

```

## Summary

- **Trigger** the skill with `caveman-compress <filepath>` as defined in [`src/plugins/opencode/commands/caveman-compress.md`](https://github.com/JuliusBrussee/caveman/blob/main/src/plugins/opencode/commands/caveman-compress.md).
- **Processing** occurs in [`plugins/caveman/skills/caveman-compress/scripts/__main__.py`](https://github.com/JuliusBrussee/caveman/blob/main/plugins/caveman/skills/caveman-compress/scripts/__main__.py), which sends content to Claude and validates output while preserving all technical formatting.
- **Retries** happen automatically (up to two times) if validation fails.
- **Backups** are stored in `$XDG_DATA_HOME/caveman-compress/backups/` as `<file>.original.md` before overwriting originals.
- **Eligibility** is limited to natural-language files (`.md`, `.txt`, `.typ`, `.tex`, extension-less); source code and existing backups are rejected.

## Frequently Asked Questions

### How does caveman-compress handle validation failures?

According to the JuliusBrussee/caveman source code, if Claude's output fails validation—meaning it altered code blocks, removed URLs, or damaged markdown structure—the system automatically retries the compression up to two additional times before failing. This ensures technical content remains intact.

### Where are the backup files stored?

Backup files are stored in an out-of-tree data directory at `$XDG_DATA_HOME/caveman-compress/backups/` (falling back to `~/.local/share` if the environment variable is unset). This prevents backup files from being accidentally processed as input files during subsequent compression runs.

### Can I compress Python or JSON files with this skill?

No. The skill explicitly rejects source code and configuration files including `.py`, `.js`, `.json`, `.yaml`, and similar formats. Only natural-language documentation files (`.md`, `.txt`, `.typ`, `.tex`, or extension-less text files) are eligible for compression per the rules in [`plugins/caveman/skills/caveman-compress/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/plugins/caveman/skills/caveman-compress/SKILL.md).

### What happens if I run the command on a file that is already compressed?

If you attempt to compress a file that has already been processed, the skill will proceed normally provided a backup doesn't already exist. However, if an `*.original.md` backup already exists for that file, the skill rejects the operation to prevent double-compression and potential data loss.