How Does the caveman-shrink MCP Middleware Compress Tool Descriptions?
The caveman-shrink MCP middleware compresses tool descriptions by removing articles, filler words, and hedging phrases while protecting code blocks, URLs, and file paths through a sentinel-based placeholder system, typically achieving 15–30% token savings.
The caveman-shrink middleware is a specialized proxy for the Model-Context-Protocol (MCP) ecosystem. According to the JuliusBrussee/caveman source code, it sits between AI clients and upstream MCP servers, intercepting JSON-RPC responses to reduce the token footprint of tool descriptions before they reach the language model. This article explains the exact compression algorithm implemented in the repository and how it preserves critical syntax while minimizing prose.
What Is caveman-shrink?
caveman-shrink functions as a transparent middleware layer that spawns an upstream MCP server and proxies communication between the client and server. As implemented in src/mcp-servers/caveman-shrink/index.js, the middleware:
- Spawns the upstream server using configuration from
src/mcp-servers/caveman-shrink/spawn-options.js - Passes through stdin/stdout for request forwarding
- Intercepts JSON-RPC responses containing arrays of
tools,prompts,resources, orresourceTemplates - Applies the
transformResponsefunction to compress configured fields before returning data to the client
The middleware deliberately avoids modifying request payloads sent to the upstream server, ensuring compatibility with existing MCP server implementations.
How the Compression Algorithm Works
The core compression logic resides in src/mcp-servers/caveman-shrink/compress.js. The algorithm operates through a two-phase process: sentinel protection followed by text normalization.
Protected Sentinels and Preservation
Before any text reduction occurs, the algorithm identifies and protects code-like content that must remain verbatim. The compress() function replaces the following patterns with unique placeholder tokens:
- Code fences (triple backtick blocks)
- Inline code (single backticks)
- URLs and file paths
- Constant-style identifiers (e.g.,
UPPER_SNAKE_CASE) - Enum strings and other machine-readable values
These protected sections are replaced by sentinels, allowing the text normalization phase to operate safely without corrupting functional syntax. After processing, the original values are restored exactly as they appeared in the source.
Text Normalization and Token Reduction
Once sentinels mask the protected content, the algorithm applies aggressive prose compression:
- Article removal: Eliminates "a", "an", and "the"
- Filler elimination: Strips words like "sure", "just", "basically", and "simply"
- Hedging reduction: Removes phrases like "I will", "you can", and other conversational prefixes
- Whitespace normalization: Collapses multiple spaces and normalizes punctuation spacing
For example, the description "The user is the owner of an account and can manage it freely" becomes "User is owner of account can manage freely" after compression.
Implementation Details in the Source Code
The compression pipeline involves several key functions across the codebase:
compress() in compress.js
Implements the sentinel protection and text normalization algorithm described above. This function handles the heavy lifting of pattern matching and word removal.
transformResponse() in index.js
Intercepts JSON-RPC responses and identifies arrays containing tools, prompts, resources, or resourceTemplates. For each object in these arrays, it extracts the fields designated for compression and passes them to the compression helper.
compressDescriptionsInPlace()
A recursive helper that traverses the entire response payload when top-level array matching fails. This ensures deep-nested descriptions—such as those within tool parameter objects—are also compressed according to the configuration.
Configuration and Usage
The middleware behavior is controlled through environment variables:
CAVEMAN_SHRINK_FIELDS: Comma-separated list of field names to compress (default:description)CAVEMAN_SHRINK_DEBUG: Set to1to log byte-size deltas for each compressed field to stderr
To run the middleware, wrap your existing MCP server command:
npx caveman-shrink npx @modelcontextprotocol/server-filesystem /path/to/files
With debug output enabled:
CAVEMAN_SHRINK_DEBUG=1 CAVEMAN_SHRINK_FIELDS=description,summary \
npx caveman-shrink npx @modelcontextprotocol/server-filesystem /tmp
Sample debug output from index.js shows the compression efficiency:
[caveman-shrink] tools.get_weather.description: 78→52 bytes
Summary
- caveman-shrink acts as a proxy middleware between MCP clients and upstream servers, implemented in
src/mcp-servers/caveman-shrink/index.js - The compression algorithm in
src/mcp-servers/caveman-shrink/compress.jsuses sentinel placeholders to protect code, URLs, and file paths before applying text normalization - The system removes articles, filler words, and hedging phrases while preserving semantic meaning and functional syntax
- Configuration via
CAVEMAN_SHRINK_FIELDSallows targeting of arbitrary fields beyond the defaultdescription - Typical token savings range from 15–30%, with debug logging available to verify compression ratios
Frequently Asked Questions
Does caveman-shrink modify requests sent to the upstream server?
No. The middleware only intercepts and transforms JSON-RPC responses flowing from the upstream server to the client. Request payloads sent to the server remain untouched to prevent parsing errors in the upstream MCP implementation.
What fields does caveman-shrink compress by default?
By default, only the description field is compressed. You can extend this to other fields by setting the CAVEMAN_SHRINK_FIELDS environment variable to a comma-separated list (e.g., description,summary,help).
How does the algorithm protect code snippets during compression?
The compress() function replaces code fences, inline code, URLs, and file paths with unique sentinel tokens before text processing begins. These placeholders prevent the normalization rules from modifying protected content. After compression, the original values are restored verbatim.
What is the typical token savings when using caveman-shrink?
The caveman-shrink MCP middleware typically reduces token usage by 15–30% for tool descriptions. The exact savings depend on the verbosity of the original descriptions and the density of protected content (code blocks, URLs) that must remain unchanged.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →