How Caveman Handles Inputs That Cannot Be Compressed or Become Larger After Transformation

Caveman guarantees lossless compression by returning the original input unchanged whenever a transformation would result in equal or larger byte size, preventing any token-count regression.

The JuliusBrussee/caveman repository implements a defensive compression pipeline designed specifically for MCP (Model Context Protocol) payloads. When processing text for token reduction, the system prioritizes data integrity over forced transformation, ensuring that edge cases—such as already-optimized strings or content rich in stop-words—never inflate the final output.

Size Regression Protection in the Core Compressor

The primary defense against size bloat resides in src/mcp-servers/caveman-shrink/compress.js. The exported compress function executes a series of deterministic rewrites—including stop-word removal, whitespace collapsing, and URL shortening—then applies a strict size comparison before returning results.

The Byte-Length Guard Clause

After applying transformations, the function compares the byte length of the processed text against the original. This logic ensures that only beneficial modifications survive:

  • If smaller: The compressed text is returned, and a .original.md backup is created for session restoration
  • If equal or larger: The function discards the transformation and returns the pristine original, skipping backup creation entirely

This guard clause prevents the compressor from accidentally increasing token counts on inputs that are already minimal or contain high densities of incompressible tokens.

Code Example: Basic Compression with Fallback

const { compress } = require('./src/mcp-servers/caveman-shrink/compress');

function maybeCompress(text) {
  const { compressed, before, after } = compress(text);
  // `compressed` is the original text if compression didn’t help
  return compressed;
}

// Example – a sentence that compresses nicely
console.log(maybeCompress('The user is the owner of an account'));
// → "User is owner of account"

// Example – a sentence that would get larger after naive rewrites
console.log(maybeCompress('Sure, this just basically returns the value'));
// → "Sure, this just basically returns the value" (unchanged)

Safeguarding Nested Schema Objects

For complex API responses containing nested tool descriptions, Caveman provides compressDescriptionsInPlace. This utility walks schema objects recursively, applying compression only where it provides measurable benefit.

Selective Field Processing

The function inspects each string field within nested objects and invokes compress only when the value is a non-null string. Critically, it skips any field where compression would not shrink the text, ensuring that nested payloads never regress in size. This is essential for MCP server responses where unpredictable user-generated content might resist simplification.

Code Example: Nested Object Processing

const { compressDescriptionsInPlace } = require('./src/mcp-servers/caveman-shrink/compress');

const payload = {
  result: {
    description: 'The weather forecast for the city name is 75°F',
    subtool: {
      description: 'Connect to the database at https://example.com/api'
    }
  }
};

compressDescriptionsInPlace(payload.result, ['description']);
console.log(JSON.stringify(payload, null, 2));
/* The `description` fields are compressed only when the
   resulting string is shorter; otherwise they stay untouched. */

Statistical Tracking of Effective Compression

The system maintains accurate telemetry by filtering out non-beneficial compression attempts. In src/hooks/caveman-stats.js, the statistics collector specifically ignores file pairs where compressedSize is not strictly less than originalSize.

Code Example: Statistics Filtering

// Inside src/hooks/caveman-stats.js
function summarizeCompressed(pairs) {
  // Only count pairs where compressedSize < originalSize
  const good = pairs.filter(p => p.compressedSize < p.originalSize);
  return {
    count: good.length,
    tokensSaved: good.reduce((t, p) => t + (p.originalSize - p.compressedSize) * 0.75, 0)
  };
}

This filtering ensures that reported memory savings reflect only genuine compression wins, excluding no-op transformations from metrics.

Test Coverage for Size Safety

The safety guarantees are enforced by multiple test suites:

  • tests/test_caveman_stats.js: Explicitly verifies that "pairs where compressed is not actually smaller" are excluded from memory-compression statistics
  • tests/test_compress_safety.js: Contains safety tests ensuring the compressor never returns a larger string than provided
  • tests/test_mcp_shrink.js: Unit tests for the compress function demonstrating both successful compression cases and intentional no-op returns

These tests validate that the size regression protection works across edge cases, including strings with minimal stop-words, already-optimized technical content, and single-character inputs.

Summary

  • Size-first validation: The compress function in compress.js always validates byte length post-transformation, returning the original when no benefit exists
  • No backup bloat: .original.md backup files are created only when compression succeeds, preventing filesystem overhead for unchanged content
  • Nested object safety: compressDescriptionsInPlace skips fields where compression fails to reduce size, protecting complex schema payloads
  • Accurate metrics: The statistics collector in caveman-stats.js filters for strictly smaller outputs, ensuring reported savings are genuine
  • Comprehensive testing: Multiple test suites verify that inputs resistant to compression never trigger size regression

Frequently Asked Questions

What happens if Caveman tries to compress content that is already minimal?

If the transformed text is equal to or larger than the original, Caveman discards the transformation and returns the original unchanged. No backup file is created, and the statistics system treats the file as uncompressed, ensuring zero overhead for suboptimal candidates.

Does Caveman create backup files for every input processed?

No. Backup files with the .original.md extension are created only when compression actually reduces the byte size. If the input cannot be compressed or grows larger after transformation, the function exits early without writing backup files, keeping the workspace clean.

How does Caveman handle compression of nested JSON objects or API responses?

The compressDescriptionsInPlace function traverses nested objects and applies compression only to string fields that exceed their original length after processing. It automatically skips null values, non-strings, and any field where the transformation fails to shrink the content.

Will compression attempts that increase size affect the reported token savings?

No. The statistics collector in src/hooks/caveman-stats.js explicitly filters the results array to include only pairs where compressedSize < originalSize. Attempts that result in equal or larger outputs are excluded from the tokensSaved calculation and compression count.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →