# What Happens to Original Files When Using /caveman-compress: Complete Backup Workflow Explained

> Discover what happens to original files when using caveman-compress. Learn how caveman protects your data with verified backups before compressing, ensuring a safe workflow.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: how-to-guide
- Published: 2026-07-11

---

**When you invoke `/caveman-compress`, the original file is never deleted or overwritten without protection; instead, the system creates a verified [`.original.md`](https://github.com/JuliusBrussee/caveman/blob/main/.original.md) backup outside the repository, validates the compressed output, and only then atomically replaces the source file.**

The `/caveman-compress` command in the JuliusBrussee/caveman repository offers AI-powered markdown compression while implementing a defensive, multi-layered backup strategy. Understanding exactly what happens to your source files during this automated process helps developers trust the tool with critical documentation. This guide breaks down the complete workflow from file validation through safe replacement.

## Pre-Compression Safety Checks

Before any data leaves your machine, [`skills/caveman-compress/scripts/compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/compress.py) performs several defensive checks to protect your original content.

### File Validation and Size Limits

The script begins by canonicalizing the target path and validating file constraints. At lines 22–30, it resolves the absolute path and enforces a 500 KB maximum file size:

```python
filepath = filepath.resolve()
if filepath.stat().st_size > MAX_FILE_SIZE:
    # Abort if file too large

```

### Sensitive File Protection

Between lines 31–42, the script checks for potentially dangerous paths using `is_sensitive_path(filepath)`. Files matching patterns like `.env`, `*.pem`, or `credentials` are rejected immediately with a `ValueError`, ensuring secrets never leave your machine.

### Content Type Detection

At lines 45–48, the script invokes `should_compress(filepath)` from [`detect.py`](https://github.com/JuliusBrussee/caveman/blob/main/detect.py) to verify the file contains natural language content worth compressing. If the heuristics suggest the file isn't compressible, the operation exits early and the original remains untouched.

## The Backup Creation Process

Once pre-checks pass, the system implements a rigorous backup strategy that stores originals outside the source tree.

### Out-of-Tree Backup Location

At lines 53–57, the script calculates a backup directory in `~/.local/share/caveman-compress/backups/<parent-dir>/`. This out-of-tree location prevents the skill loader from re-ingesting backups as live source files. The backup receives the suffix [`.original.md`](https://github.com/JuliusBrussee/caveman/blob/main/.original.md):

```python
backup_dir = backup_dir_for(filepath)
backup_path = backup_dir / (filepath.stem + ".original.md")

```

### Duplicate Protection

Lines 61–66 implement a critical safety guard: if `backup_path.exists()`, the script aborts immediately. This prevents accidental overwrites of previous backups if you attempt to compress the same file twice.

### Frontmatter Preservation

For markdown files containing YAML frontmatter, lines 68–73 split the content using `split_frontmatter(original_text)`. The frontmatter is stored verbatim and later reattached to the compressed output, ensuring metadata remains intact.

## Compression and Validation Workflow

Only after backup verification does the system proceed with AI compression and atomic replacement.

### Identity Check Before Processing

At lines 79–82, the script sends the body content to Claude via `call_claude(build_compress_prompt(body))`. Upon receiving the compressed result, lines 88–94 perform an identity check:

```python
if compressed_body.strip() == body.strip():
    # Abort if Claude returned identical text

    return False

```

If the AI returns unchanged content, the operation aborts and the original file remains untouched.

### Verified Backup Write

Lines 99–106 implement a write-verify pattern: the original content is written to the backup path, immediately read back, and compared to the source. If `backup_readback` doesn't match `original_text`, the backup is deleted and the process aborts.

### Atomic Replacement with Rollback

After successful verification, lines 111–112 write the compressed content to the original file path. The script then runs validation via [`validate.py`](https://github.com/JuliusBrussee/caveman/blob/main/validate.py) (lines 115–136). If validation fails, Claude attempts up to two retries (`MAX_RETRIES`). If all attempts fail, the script restores the original content from the backup and removes the backup file, returning the system to its original state.

## Code Examples

### Command Line Usage

```bash

# Basic compression from command line

$ python skills/caveman-compress/scripts/compress.py docs/guide.md
Processing: /home/user/project/docs/guide.md
Detected YAML frontmatter (123 chars) — preserving verbatim
Compressing with Claude...
✅ Validation passed

```

### Slash Command Usage

```bash

# Using the chat-agent command

/user: /caveman-compress docs/guide.md
→ Caveman compresses the file, backs up the original to
   ~/.local/share/caveman-compress/backups/docs/guide.original.md
   and overwrites docs/guide.md with the compressed markdown.

```

### Programmatic Integration

```python
from pathlib import Path
from skills.caveman-compress.scripts.compress import compress_file

if compress_file(Path("docs/guide.md")):
    print("File compressed and original safely backed up.")
else:
    print("Compression skipped or failed; original untouched.")

```

## Key Implementation Files

| File | Role |
|------|------|
| [`skills/caveman-compress/scripts/compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/compress.py) | Core orchestrator that creates backups, calls Claude, validates, and writes the compressed file (lines 22–136). |
| [`skills/caveman-compress/scripts/detect.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/detect.py) | Contains `should_compress` – heuristics that decide whether a file is natural language and therefore compressible. |
| [`skills/caveman-compress/scripts/validate.py`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/scripts/validate.py) | Implements post-compression validation logic (checks for missing URLs, code-block mismatches, heading integrity). |
| [`src/plugins/opencode/commands/caveman-compress.md`](https://github.com/JuliusBrussee/caveman/blob/main/src/plugins/opencode/commands/caveman-compress.md) | Defines the slash-command `/caveman-compress` that forwards requests to the Python script. |
| [`src/mcp-servers/caveman-shrink/compress.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/mcp-servers/caveman-shrink/compress.js) | JavaScript counterpart used by the MCP server for on-the-fly compression (mirrors the Python flow). |

## Summary

- **Original files are never deleted**: They are preserved with the [`.original.md`](https://github.com/JuliusBrussee/caveman/blob/main/.original.md) suffix in `~/.local/share/caveman-compress/backups/`
- **Backups are verified before replacement**: The script reads back the backup to confirm integrity before overwriting the source
- **Sensitive files are protected**: Files matching `.env`, `*.pem`, or credential patterns are rejected before processing
- **Automatic rollback**: Validation failures trigger restoration of the original content and removal of the backup
- **Duplicate protection**: The script blocks re-compression attempts if a backup already exists to prevent accidental overwrites

## Frequently Asked Questions

### Does /caveman-compress delete my original file?

No. The original file is never deleted. Instead, it is preserved as a backup with the [`.original.md`](https://github.com/JuliusBrussee/caveman/blob/main/.original.md) extension in an out-of-tree directory (`~/.local/share/caveman-compress/backups/`). The source file is only overwritten after the backup is verified and the compressed content passes validation checks.

### Where does caveman-compress store backup files?

Backups are stored in `~/.local/share/caveman-compress/backups/` following the directory structure of the original file's parent directory. For example, compressing [`docs/guide.md`](https://github.com/JuliusBrussee/caveman/blob/main/docs/guide.md) creates a backup at `~/.local/share/caveman-compress/backups/docs/guide.original.md`. This prevents the skill loader from treating backups as source files.

### What happens if the compression fails or produces invalid output?

If validation fails after two retry attempts (as implemented in [`compress.py`](https://github.com/JuliusBrussee/caveman/blob/main/compress.py) lines 115–136), the script automatically restores the original content from the backup and removes the backup file. The original file remains unchanged throughout the entire process if any step fails, including network errors, empty responses, or validation mismatches.

### Can I compress the same file twice?

No. The script explicitly checks for existing backup files at lines 61–66 (`if backup_path.exists()`) and aborts if found. This prevents accidental overwrites of your original backup. To re-compress a file, you must manually move or remove the existing [`.original.md`](https://github.com/JuliusBrussee/caveman/blob/main/.original.md) backup file from the `~/.local/share/caveman-compress/backups/` directory.