# How Does the caveman-shrink MCP Middleware Compress Tool Descriptions?

> Discover how caveman-shrink MCP middleware compresses tool descriptions. Learn its token-saving strategy by removing filler words while preserving essential code and paths.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-07-12

---

**The caveman-shrink MCP middleware compresses tool descriptions by removing articles, filler words, and hedging phrases while protecting code blocks, URLs, and file paths through a sentinel-based placeholder system, typically achieving 15–30% token savings.**

The **caveman-shrink** middleware is a specialized proxy for the Model-Context-Protocol (MCP) ecosystem. According to the JuliusBrussee/caveman source code, it sits between AI clients and upstream MCP servers, intercepting JSON-RPC responses to reduce the token footprint of tool descriptions before they reach the language model. This article explains the exact compression algorithm implemented in the repository and how it preserves critical syntax while minimizing prose.

## What Is caveman-shrink?

**caveman-shrink** functions as a transparent middleware layer that spawns an upstream MCP server and proxies communication between the client and server. As implemented in [`src/mcp-servers/caveman-shrink/index.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/mcp-servers/caveman-shrink/index.js), the middleware:

1. Spawns the upstream server using configuration from [`src/mcp-servers/caveman-shrink/spawn-options.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/mcp-servers/caveman-shrink/spawn-options.js)
2. Passes through stdin/stdout for request forwarding
3. Intercepts JSON-RPC responses containing arrays of `tools`, `prompts`, `resources`, or `resourceTemplates`
4. Applies the `transformResponse` function to compress configured fields before returning data to the client

The middleware deliberately avoids modifying request payloads sent to the upstream server, ensuring compatibility with existing MCP server implementations.

## How the Compression Algorithm Works

The core compression logic resides in [`src/mcp-servers/caveman-shrink/compress.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/mcp-servers/caveman-shrink/compress.js). The algorithm operates through a two-phase process: **sentinel protection** followed by **text normalization**.

### Protected Sentinels and Preservation

Before any text reduction occurs, the algorithm identifies and protects code-like content that must remain verbatim. The `compress()` function replaces the following patterns with unique placeholder tokens:

- **Code fences** (triple backtick blocks)
- **Inline code** (single backticks)
- **URLs** and **file paths**
- **Constant-style identifiers** (e.g., `UPPER_SNAKE_CASE`)
- **Enum strings** and other machine-readable values

These protected sections are replaced by sentinels, allowing the text normalization phase to operate safely without corrupting functional syntax. After processing, the original values are restored exactly as they appeared in the source.

### Text Normalization and Token Reduction

Once sentinels mask the protected content, the algorithm applies aggressive prose compression:

- **Article removal**: Eliminates "a", "an", and "the"
- **Filler elimination**: Strips words like "sure", "just", "basically", and "simply"
- **Hedging reduction**: Removes phrases like "I will", "you can", and other conversational prefixes
- **Whitespace normalization**: Collapses multiple spaces and normalizes punctuation spacing

For example, the description "The user is the owner of an account and can manage it freely" becomes "User is owner of account can manage freely" after compression.

## Implementation Details in the Source Code

The compression pipeline involves several key functions across the codebase:

**`compress()` in [`compress.js`](https://github.com/JuliusBrussee/caveman/blob/main/compress.js)**
Implements the sentinel protection and text normalization algorithm described above. This function handles the heavy lifting of pattern matching and word removal.

**`transformResponse()` in [`index.js`](https://github.com/JuliusBrussee/caveman/blob/main/index.js)**
Intercepts JSON-RPC responses and identifies arrays containing `tools`, `prompts`, `resources`, or `resourceTemplates`. For each object in these arrays, it extracts the fields designated for compression and passes them to the compression helper.

**`compressDescriptionsInPlace()`**
A recursive helper that traverses the entire response payload when top-level array matching fails. This ensures deep-nested descriptions—such as those within tool parameter objects—are also compressed according to the configuration.

## Configuration and Usage

The middleware behavior is controlled through environment variables:

- **`CAVEMAN_SHRINK_FIELDS`**: Comma-separated list of field names to compress (default: `description`)
- **`CAVEMAN_SHRINK_DEBUG`**: Set to `1` to log byte-size deltas for each compressed field to stderr

To run the middleware, wrap your existing MCP server command:

```bash
npx caveman-shrink npx @modelcontextprotocol/server-filesystem /path/to/files

```

With debug output enabled:

```bash
CAVEMAN_SHRINK_DEBUG=1 CAVEMAN_SHRINK_FIELDS=description,summary \
  npx caveman-shrink npx @modelcontextprotocol/server-filesystem /tmp

```

Sample debug output from [`index.js`](https://github.com/JuliusBrussee/caveman/blob/main/index.js) shows the compression efficiency:

```

[caveman-shrink] tools.get_weather.description: 78→52 bytes

```

## Summary

- **caveman-shrink** acts as a proxy middleware between MCP clients and upstream servers, implemented in [`src/mcp-servers/caveman-shrink/index.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/mcp-servers/caveman-shrink/index.js)
- The compression algorithm in [`src/mcp-servers/caveman-shrink/compress.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/mcp-servers/caveman-shrink/compress.js) uses **sentinel placeholders** to protect code, URLs, and file paths before applying text normalization
- The system removes **articles**, **filler words**, and **hedging phrases** while preserving semantic meaning and functional syntax
- **Configuration** via `CAVEMAN_SHRINK_FIELDS` allows targeting of arbitrary fields beyond the default `description`
- Typical token savings range from **15–30%**, with debug logging available to verify compression ratios

## Frequently Asked Questions

### Does caveman-shrink modify requests sent to the upstream server?

No. The middleware only intercepts and transforms JSON-RPC responses flowing from the upstream server to the client. Request payloads sent to the server remain untouched to prevent parsing errors in the upstream MCP implementation.

### What fields does caveman-shrink compress by default?

By default, only the `description` field is compressed. You can extend this to other fields by setting the `CAVEMAN_SHRINK_FIELDS` environment variable to a comma-separated list (e.g., `description,summary,help`).

### How does the algorithm protect code snippets during compression?

The `compress()` function replaces code fences, inline code, URLs, and file paths with unique sentinel tokens before text processing begins. These placeholders prevent the normalization rules from modifying protected content. After compression, the original values are restored verbatim.

### What is the typical token savings when using caveman-shrink?

The caveman-shrink MCP middleware typically reduces token usage by **15–30%** for tool descriptions. The exact savings depend on the verbosity of the original descriptions and the density of protected content (code blocks, URLs) that must remain unchanged.