# How caveman-shrink MCP Server Performs Compression Tasks

> Discover how the caveman-shrink MCP server optimizes text by normalizing whitespace removing stop words shortening URLs masking secrets and preserving code blocks Get efficient data compression.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: internals
- Published: 2026-07-13

---

**The caveman-shrink MCP server compresses textual descriptions by normalizing whitespace, removing stop-words, shortening URLs and file paths, masking secrets, and preserving code blocks, returning both the optimized string and before/after character statistics.**

The `caveman-shrink` middleware in the JuliusBrussee/caveman repository functions as a Model-Context-Protocol (MCP) proxy that intercepts responses from upstream servers to reduce payload size through rule-based text transformation. This pure Node.js compression system processes tool descriptions and other textual fields in real-time before forwarding minimized payloads to clients.

## Core Compression Pipeline in [`compress.js`](https://github.com/JuliusBrussee/caveman/blob/main/compress.js)

The compression engine resides in [`src/mcp-servers/caveman-shrink/compress.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/mcp-servers/caveman-shrink/compress.js), implementing a seven-stage pipeline within the `compress()` function. According to the JuliusBrussee/caveman source code, this module handles text optimization while protecting sensitive data and code integrity.

### Input Normalization and Whitespace Handling

The function begins by collapsing multiple whitespace characters into single spaces and trimming leading and trailing boundaries. It specifically preserves code-fence delimiters (```) and inline back-ticks to ensure markdown formatting remains intact during processing.

### Stop-Word Removal and Filler Elimination

Using regex-based replacements, the compressor strips common articles (`the`, `a`, `an`) and low-value filler words (`sure`, `just`, `basically`, `perhaps`). These terms consume token space without contributing significant semantic meaning to LLM contexts.

### URL and File Path Shortening

For URLs (`https://example.com/...`) and file paths (`/tmp/...`), the system analyzes individual path segments and removes those matching the stop-word list. This produces shorter but still valid references that maintain navigability while reducing character count.

### Secret Masking for Safe Logging

The compressor detects typical secret patterns (e.g., `API_KEY_...`) and replaces them with generic placeholders like `API_KEY_VALUE`. This masking occurs during a dedicated processing stage in `compress()`, ensuring credentials never leak into logs while indicating that authentication tokens were present in the original text.

### Code Snippet Preservation

Content inside fenced code blocks or inline back-ticks bypasses stop-word removal entirely. This protection prevents syntax breakage that would occur if structural words were removed from code examples, making the compressor safe for technical documentation.

### Sentinel Cleanup and Statistics Generation

Temporary sentinel strings such as ` N ` used during internal processing are removed in a final sweep. The function returns an object containing the `compressed` string, the original `before` character count, and the final `after` character count, enabling precise compression ratio calculations.

## Recursive Description Processing

The `compressDescriptionsInPlace()` function, also defined in `src/mcp-servers/caveman-shrink/compress.js`, recursively traverses nested response objects typically found in MCP tool-list payloads. It identifies fields specified in a description map (commonly `description` keys) and replaces their values in-place with compressed versions using the `compress()` function.

This recursive approach ensures consistent optimization across complex nested structures without modifying non-description fields or breaking object references in the upstream response.

## MCP Proxy Integration and Request Flow

The compression pipeline integrates into the middleware architecture through `src/mcp-servers/caveman-shrink/index.js`. This entry point creates the MCP proxy server that spawns the upstream MCP server process using configuration logic from `src/mcp-servers/caveman-shrink/spawn-options.js`.

The execution flow follows this sequence:

- The client connects to `caveman-shrink` acting as the MCP proxy
- The proxy spawns the upstream MCP server via `spawn-options.js`
- Upon receiving responses, the proxy invokes `compressDescriptionsInPlace()` on the payload
- The compressed result returns to the client with reduced token count

This architecture allows transparent middleware operation requiring no modifications to upstream server implementations while delivering significant bandwidth and context-window savings.

## Practical Implementation Examples

You can utilize the compression module directly for standalone text processing or custom MCP workflows.

```javascript
// Example: manually compress a single string
const { compress } = require(
  path.join(__dirname, 'src', 'mcp-servers', 'caveman-shrink', 'compress.js')
);

const { compressed, before, after } = compress(
  'The user is the owner of an account'
);
console.log(compressed);   // => "User is owner of account"
console.log(`Reduced ${before} → ${after} chars`);

```

```javascript
// Example: compress all description fields inside a nested MCP result
const { compressDescriptionsInPlace } = require(
  path.join(__dirname, 'src', 'mcp-servers', 'caveman-shrink', 'compress.js')
);

const payload = {
  result: {
    tools: [
      { description: 'Sure, this just basically returns the value' },
      { description: 'I will perhaps connect to the database' }
    ]
  }
};

compressDescriptionsInPlace(payload.result, ['description']);
console.log(payload.result.tools[0].description); // "returns the value"
console.log(payload.result.tools[1].description); // "connect to the database"

```

## Summary

- The compression engine lives in `src/mcp-servers/caveman-shrink/compress.js` and exports `compress()` and `compressDescriptionsInPlace()`.
- **Seven processing stages** handle normalization, stop-word removal, URL/path shortening, secret masking, code preservation, sentinel cleanup, and statistics generation.
- The `compress()` function returns an object with `compressed`, `before`, and `after` properties to track compression ratios.
- `compressDescriptionsInPlace()` recursively processes nested MCP response objects, targeting specific field names like `description`.
- The middleware entry point at `src/mcp-servers/caveman-shrink/index.js` automatically applies compression to all upstream responses via the MCP proxy pattern.
- Code blocks and inline back-ticks remain untouched to prevent syntax corruption during text optimization.

## Frequently Asked Questions

### How does caveman-shrink handle sensitive information during compression?

The compressor masks secrets using regex patterns that detect tokens like `API_KEY_...` and replace them with generic placeholders such as `API_KEY_VALUE`. This occurs in `src/mcp-servers/caveman-shrink/compress.js` during the dedicated secret-masking stage, ensuring credentials never appear in logs or compressed outputs while maintaining the structural indication that a secret was present.

### What types of text does the compression algorithm modify?

The algorithm primarily targets natural language descriptions, removing articles and filler words like `the`, `a`, `an`, `sure`, `just`, and `basically`. It shortens URLs and file paths by dropping stop-word segments. However, it explicitly preserves content within code fences (```) and inline back-ticks to prevent breaking syntax, making it safe for technical documentation and API descriptions.

### How can I measure the effectiveness of the compression?

The `compress()` function returns statistics alongside the processed text. The returned object includes `before` (original character count), `after` (compressed character count), and the `compressed` string itself. You can calculate the ratio by comparing `before` and `after` values, as demonstrated in the test suite at [`tests/test_mcp_shrink.js`](https://github.com/JuliusBrussee/caveman/blob/main/tests/test_mcp_shrink.js).

### Where does the compression occur in the MCP request lifecycle?

Compression happens at the response stage within the middleware layer. The [`src/mcp-servers/caveman-shrink/index.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/mcp-servers/caveman-shrink/index.js) file creates an MCP proxy that forwards requests to an upstream server. When the upstream returns a response, `compressDescriptionsInPlace()` processes the payload before the proxy sends the final result to the client, ensuring minimal latency impact while maximizing token efficiency.