# How Caveman Handles Inputs That Cannot Be Compressed or Become Larger After Transformation

> Caveman ensures lossless compression by returning original input unchanged if transformation results in equal or larger size. Prevent token-count regression with this guarantee.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Caveman guarantees lossless compression by returning the original input unchanged whenever a transformation would result in equal or larger byte size, preventing any token-count regression.**

The [JuliusBrussee/caveman](https://github.com/JuliusBrussee/caveman) repository implements a defensive compression pipeline designed specifically for MCP (Model Context Protocol) payloads. When processing text for token reduction, the system prioritizes data integrity over forced transformation, ensuring that edge cases—such as already-optimized strings or content rich in stop-words—never inflate the final output.

## Size Regression Protection in the Core Compressor

The primary defense against size bloat resides in [`src/mcp-servers/caveman-shrink/compress.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/mcp-servers/caveman-shrink/compress.js). The exported `compress` function executes a series of deterministic rewrites—including stop-word removal, whitespace collapsing, and URL shortening—then applies a strict size comparison before returning results.

### The Byte-Length Guard Clause

After applying transformations, the function compares the byte length of the processed text against the original. This logic ensures that only beneficial modifications survive:

- **If smaller**: The compressed text is returned, and a [`.original.md`](https://github.com/JuliusBrussee/caveman/blob/main/.original.md) backup is created for session restoration
- **If equal or larger**: The function discards the transformation and returns the pristine original, skipping backup creation entirely

This guard clause prevents the compressor from accidentally increasing token counts on inputs that are already minimal or contain high densities of incompressible tokens.

### Code Example: Basic Compression with Fallback

```javascript
const { compress } = require('./src/mcp-servers/caveman-shrink/compress');

function maybeCompress(text) {
  const { compressed, before, after } = compress(text);
  // `compressed` is the original text if compression didn’t help
  return compressed;
}

// Example – a sentence that compresses nicely
console.log(maybeCompress('The user is the owner of an account'));
// → "User is owner of account"

// Example – a sentence that would get larger after naive rewrites
console.log(maybeCompress('Sure, this just basically returns the value'));
// → "Sure, this just basically returns the value" (unchanged)

```

## Safeguarding Nested Schema Objects

For complex API responses containing nested tool descriptions, Caveman provides `compressDescriptionsInPlace`. This utility walks schema objects recursively, applying compression only where it provides measurable benefit.

### Selective Field Processing

The function inspects each string field within nested objects and invokes `compress` only when the value is a non-null string. Critically, it **skips** any field where compression would not shrink the text, ensuring that nested payloads never regress in size. This is essential for MCP server responses where unpredictable user-generated content might resist simplification.

### Code Example: Nested Object Processing

```javascript
const { compressDescriptionsInPlace } = require('./src/mcp-servers/caveman-shrink/compress');

const payload = {
  result: {
    description: 'The weather forecast for the city name is 75°F',
    subtool: {
      description: 'Connect to the database at https://example.com/api'
    }
  }
};

compressDescriptionsInPlace(payload.result, ['description']);
console.log(JSON.stringify(payload, null, 2));
/* The `description` fields are compressed only when the
   resulting string is shorter; otherwise they stay untouched. */

```

## Statistical Tracking of Effective Compression

The system maintains accurate telemetry by filtering out non-beneficial compression attempts. In [`src/hooks/caveman-stats.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/hooks/caveman-stats.js), the statistics collector specifically ignores file pairs where `compressedSize` is not strictly less than `originalSize`.

### Code Example: Statistics Filtering

```javascript
// Inside src/hooks/caveman-stats.js
function summarizeCompressed(pairs) {
  // Only count pairs where compressedSize < originalSize
  const good = pairs.filter(p => p.compressedSize < p.originalSize);
  return {
    count: good.length,
    tokensSaved: good.reduce((t, p) => t + (p.originalSize - p.compressedSize) * 0.75, 0)
  };
}

```

This filtering ensures that reported memory savings reflect only genuine compression wins, excluding no-op transformations from metrics.

## Test Coverage for Size Safety

The safety guarantees are enforced by multiple test suites:

- **[`tests/test_caveman_stats.js`](https://github.com/JuliusBrussee/caveman/blob/main/tests/test_caveman_stats.js)**: Explicitly verifies that "pairs where compressed is not actually smaller" are excluded from memory-compression statistics
- **[`tests/test_compress_safety.js`](https://github.com/JuliusBrussee/caveman/blob/main/tests/test_compress_safety.js)**: Contains safety tests ensuring the compressor never returns a larger string than provided
- **[`tests/test_mcp_shrink.js`](https://github.com/JuliusBrussee/caveman/blob/main/tests/test_mcp_shrink.js)**: Unit tests for the `compress` function demonstrating both successful compression cases and intentional no-op returns

These tests validate that the size regression protection works across edge cases, including strings with minimal stop-words, already-optimized technical content, and single-character inputs.

## Summary

- **Size-first validation**: The `compress` function in [`compress.js`](https://github.com/JuliusBrussee/caveman/blob/main/compress.js) always validates byte length post-transformation, returning the original when no benefit exists
- **No backup bloat**: [`.original.md`](https://github.com/JuliusBrussee/caveman/blob/main/.original.md) backup files are created only when compression succeeds, preventing filesystem overhead for unchanged content
- **Nested object safety**: `compressDescriptionsInPlace` skips fields where compression fails to reduce size, protecting complex schema payloads
- **Accurate metrics**: The statistics collector in [`caveman-stats.js`](https://github.com/JuliusBrussee/caveman/blob/main/caveman-stats.js) filters for strictly smaller outputs, ensuring reported savings are genuine
- **Comprehensive testing**: Multiple test suites verify that inputs resistant to compression never trigger size regression

## Frequently Asked Questions

### What happens if Caveman tries to compress content that is already minimal?

If the transformed text is equal to or larger than the original, Caveman discards the transformation and returns the original unchanged. No backup file is created, and the statistics system treats the file as uncompressed, ensuring zero overhead for suboptimal candidates.

### Does Caveman create backup files for every input processed?

No. Backup files with the [`.original.md`](https://github.com/JuliusBrussee/caveman/blob/main/.original.md) extension are created only when compression actually reduces the byte size. If the input cannot be compressed or grows larger after transformation, the function exits early without writing backup files, keeping the workspace clean.

### How does Caveman handle compression of nested JSON objects or API responses?

The `compressDescriptionsInPlace` function traverses nested objects and applies compression only to string fields that exceed their original length after processing. It automatically skips null values, non-strings, and any field where the transformation fails to shrink the content.

### Will compression attempts that increase size affect the reported token savings?

No. The statistics collector in [`src/hooks/caveman-stats.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/hooks/caveman-stats.js) explicitly filters the results array to include only pairs where `compressedSize < originalSize`. Attempts that result in equal or larger outputs are excluded from the `tokensSaved` calculation and compression count.