What Happens to Original Files When Using /caveman-compress: Complete Backup Workflow Explained

When you invoke /caveman-compress, the original file is never deleted or overwritten without protection; instead, the system creates a verified .original.md backup outside the repository, validates the compressed output, and only then atomically replaces the source file.

The /caveman-compress command in the JuliusBrussee/caveman repository offers AI-powered markdown compression while implementing a defensive, multi-layered backup strategy. Understanding exactly what happens to your source files during this automated process helps developers trust the tool with critical documentation. This guide breaks down the complete workflow from file validation through safe replacement.

Pre-Compression Safety Checks

Before any data leaves your machine, skills/caveman-compress/scripts/compress.py performs several defensive checks to protect your original content.

File Validation and Size Limits

The script begins by canonicalizing the target path and validating file constraints. At lines 22–30, it resolves the absolute path and enforces a 500 KB maximum file size:

filepath = filepath.resolve()
if filepath.stat().st_size > MAX_FILE_SIZE:
    # Abort if file too large

Sensitive File Protection

Between lines 31–42, the script checks for potentially dangerous paths using is_sensitive_path(filepath). Files matching patterns like .env, *.pem, or credentials are rejected immediately with a ValueError, ensuring secrets never leave your machine.

Content Type Detection

At lines 45–48, the script invokes should_compress(filepath) from detect.py to verify the file contains natural language content worth compressing. If the heuristics suggest the file isn't compressible, the operation exits early and the original remains untouched.

The Backup Creation Process

Once pre-checks pass, the system implements a rigorous backup strategy that stores originals outside the source tree.

Out-of-Tree Backup Location

At lines 53–57, the script calculates a backup directory in ~/.local/share/caveman-compress/backups/<parent-dir>/. This out-of-tree location prevents the skill loader from re-ingesting backups as live source files. The backup receives the suffix .original.md:

backup_dir = backup_dir_for(filepath)
backup_path = backup_dir / (filepath.stem + ".original.md")

Duplicate Protection

Lines 61–66 implement a critical safety guard: if backup_path.exists(), the script aborts immediately. This prevents accidental overwrites of previous backups if you attempt to compress the same file twice.

Frontmatter Preservation

For markdown files containing YAML frontmatter, lines 68–73 split the content using split_frontmatter(original_text). The frontmatter is stored verbatim and later reattached to the compressed output, ensuring metadata remains intact.

Compression and Validation Workflow

Only after backup verification does the system proceed with AI compression and atomic replacement.

Identity Check Before Processing

At lines 79–82, the script sends the body content to Claude via call_claude(build_compress_prompt(body)). Upon receiving the compressed result, lines 88–94 perform an identity check:

if compressed_body.strip() == body.strip():
    # Abort if Claude returned identical text

    return False

If the AI returns unchanged content, the operation aborts and the original file remains untouched.

Verified Backup Write

Lines 99–106 implement a write-verify pattern: the original content is written to the backup path, immediately read back, and compared to the source. If backup_readback doesn't match original_text, the backup is deleted and the process aborts.

Atomic Replacement with Rollback

After successful verification, lines 111–112 write the compressed content to the original file path. The script then runs validation via validate.py (lines 115–136). If validation fails, Claude attempts up to two retries (MAX_RETRIES). If all attempts fail, the script restores the original content from the backup and removes the backup file, returning the system to its original state.

Code Examples

Command Line Usage


# Basic compression from command line

$ python skills/caveman-compress/scripts/compress.py docs/guide.md
Processing: /home/user/project/docs/guide.md
Detected YAML frontmatter (123 chars) — preserving verbatim
Compressing with Claude...
✅ Validation passed

Slash Command Usage


# Using the chat-agent command

/user: /caveman-compress docs/guide.md
→ Caveman compresses the file, backs up the original to
   ~/.local/share/caveman-compress/backups/docs/guide.original.md
   and overwrites docs/guide.md with the compressed markdown.

Programmatic Integration

from pathlib import Path
from skills.caveman-compress.scripts.compress import compress_file

if compress_file(Path("docs/guide.md")):
    print("File compressed and original safely backed up.")
else:
    print("Compression skipped or failed; original untouched.")

Key Implementation Files

File Role
skills/caveman-compress/scripts/compress.py Core orchestrator that creates backups, calls Claude, validates, and writes the compressed file (lines 22–136).
skills/caveman-compress/scripts/detect.py Contains should_compress – heuristics that decide whether a file is natural language and therefore compressible.
skills/caveman-compress/scripts/validate.py Implements post-compression validation logic (checks for missing URLs, code-block mismatches, heading integrity).
src/plugins/opencode/commands/caveman-compress.md Defines the slash-command /caveman-compress that forwards requests to the Python script.
src/mcp-servers/caveman-shrink/compress.js JavaScript counterpart used by the MCP server for on-the-fly compression (mirrors the Python flow).

Summary

  • Original files are never deleted: They are preserved with the .original.md suffix in ~/.local/share/caveman-compress/backups/
  • Backups are verified before replacement: The script reads back the backup to confirm integrity before overwriting the source
  • Sensitive files are protected: Files matching .env, *.pem, or credential patterns are rejected before processing
  • Automatic rollback: Validation failures trigger restoration of the original content and removal of the backup
  • Duplicate protection: The script blocks re-compression attempts if a backup already exists to prevent accidental overwrites

Frequently Asked Questions

Does /caveman-compress delete my original file?

No. The original file is never deleted. Instead, it is preserved as a backup with the .original.md extension in an out-of-tree directory (~/.local/share/caveman-compress/backups/). The source file is only overwritten after the backup is verified and the compressed content passes validation checks.

Where does caveman-compress store backup files?

Backups are stored in ~/.local/share/caveman-compress/backups/ following the directory structure of the original file's parent directory. For example, compressing docs/guide.md creates a backup at ~/.local/share/caveman-compress/backups/docs/guide.original.md. This prevents the skill loader from treating backups as source files.

What happens if the compression fails or produces invalid output?

If validation fails after two retry attempts (as implemented in compress.py lines 115–136), the script automatically restores the original content from the backup and removes the backup file. The original file remains unchanged throughout the entire process if any step fails, including network errors, empty responses, or validation mismatches.

Can I compress the same file twice?

No. The script explicitly checks for existing backup files at lines 61–66 (if backup_path.exists()) and aborts if found. This prevents accidental overwrites of your original backup. To re-compress a file, you must manually move or remove the existing .original.md backup file from the ~/.local/share/caveman-compress/backups/ directory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →