# What Content Does Caveman Preserve During Compression? A Technical Breakdown

> Discover what content Caveman preserves during compression. Learn how it keeps code blocks, URLs, and technical terms while compressing prose.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-07-11

---

**Caveman preserves code blocks, inline code, URLs, file paths, command-line snippets, technical terms, headings, table structures, and numeric values while compressing natural-language prose to reduce token count.**

The `caveman-compress` tool in the JuliusBrussee/caveman repository optimizes project files for LLM context windows by stripping unnecessary verbiage. When you invoke the compression command, the tool specifically targets free-form natural language while ensuring critical technical elements remain intact. Understanding exactly what Caveman preserves during compression helps developers predict how their documentation and code will appear to AI models after processing.

## Structural Elements Preserved by Caveman

Caveman’s compression algorithm distinguishes between technical structural content and natural language. According to the source documentation in [`skills/caveman-compress/README.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/README.md), the tool maintains specific categories of content that are essential for code comprehension and navigation.

### Code Blocks and Inline Syntax

Caveman retains **all code blocks** regardless of formatting style. This includes:
- Fenced code blocks using triple backticks (```)
- Indented code blocks using spaces or tabs
- Inline code delimited by single backticks

Preserving these elements ensures that API references, function signatures, and syntax examples remain readable and accurate for LLM processing.

### Navigation and Path References

The compression process keeps all **URLs** and **markdown links** intact, ensuring documentation references remain accessible. File-system paths such as `/src/components/...` are preserved, along with command-line snippets like `npm install` or `git commit`. This maintains the technical accuracy of installation instructions and deployment workflows.

### Metadata and Structural Markers

Caveman keeps **headings** exactly as written, preserving document hierarchy and section organization. **Table structures** are retained with their layouts intact, though the text within individual cells may be shortened. Additionally, **dates**, **version numbers**, and other numeric values remain untouched during compression.

## Content That Gets Compressed

All **natural-language content** that is not technical in nature is shortened or abbreviated. This includes:
- Plain sentences and paragraphs
- Bullet-point explanations
- Explanatory prose within documentation

The tool reduces these elements to shorter forms that retain semantic meaning while significantly decreasing token count. This selective compression ensures LLMs receive the essential technical context without processing redundant verbiage.

## Caveman Compression Example

To run the compression skill on a project memory file:

```bash
/caveman-compress CLAUDE.md

```

**Before compression** (excerpt):

```markdown
I strongly prefer TypeScript with strict mode enabled for all new code. 
Please don't use `any` type unless there's genuinely no way around it, 
and if you do, leave a comment explaining the reasoning. 
Proper types catch bugs early.

```ts
// Example code block
export const foo = (arg: any) => { … }

```

Visit https://example.com for more details.

```

**After compression** (excerpt):

```markdown
Prefer TypeScript strict mode always. No `any` unless unavoidable — comment why if used. Proper types catch bugs early.

```ts
// Example code block
export const foo = (arg: any) => { … }

```

Visit https://example.com for more details.

```

Notice that the prose sentences have been shortened, but the code block, backticks, URL, and technical terms remain untouched.

## Implementation and Configuration Files

The preservation logic is implemented across several key files in the repository:

- **[`skills/caveman-compress/README.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/README.md)** — Documents what content is preserved and which file types are processed
- **[`skills/caveman-compress/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/SKILL.md)** — Contains the actual skill definition that drives the compression logic
- **[`bin/lib/settings.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/settings.js)** — Handles configuration for the Caveman toolkit, including compression settings
- **[`commands/caveman-compress.md`](https://github.com/JuliusBrussee/caveman/blob/main/commands/caveman-compress.md)** — Provides CLI documentation for invoking the compress command

## Summary

- **Caveman preserves during compression**: code blocks (fenced and indented), inline code, URLs, file paths, command-line snippets, technical terms, headings, table structures, dates, and numeric values.
- **Caveman compresses**: natural-language prose including sentences, paragraphs, and explanatory bullet points.
- The preservation behavior is documented in [`skills/caveman-compress/README.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/README.md) and implemented via the skill definition in [`skills/caveman-compress/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/SKILL.md).
- Compression reduces token count for LLM processing while maintaining technical accuracy and structural integrity.

## Frequently Asked Questions

### Does Caveman preserve code comments during compression?

Yes, Caveman preserves the content within code blocks and inline code delimiters, which includes code comments. However, free-form explanatory text outside of code blocks is compressed even if it discusses technical concepts.

### Will table formatting be lost when using Caveman compress?

No, table structures are preserved with their layouts intact. According to the documentation in [`skills/caveman-compress/README.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/README.md), only the text within table cells is shortened, while the table syntax and alignment remain unchanged.

### Does Caveman compression affect file paths and URLs?

File paths and URLs are explicitly preserved during compression. The tool recognizes patterns like `/src/components/...` and `https://example.com` as technical elements that must remain intact for proper navigation and reference.

### Where is the preservation logic defined in the Caveman repository?

The preservation rules are defined in [`skills/caveman-compress/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/SKILL.md), which contains the skill definition driving the compression logic. Documentation describing what is preserved can be found in [`skills/caveman-compress/README.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/caveman-compress/README.md), while configuration settings are managed in [`bin/lib/settings.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/settings.js).