How Caveman Handles Code Blocks During Compression: Protected Patterns in caveman-shrink
Caveman preserves markdown fenced code blocks and inline code spans unchanged during compression by identifying protected patterns and excluding them from the token-shortening pipeline.
The caveman-shrink MCP server in the JuliusBrussee/caveman repository implements a specialized compression algorithm designed to reduce token counts while maintaining critical syntax. Understanding how Caveman handles code blocks during compression is essential for developers who need to minimize context window usage without breaking markdown formatting or code examples.
Protected Content Patterns in Caveman Compression
The compression logic in src/mcp-servers/caveman-shrink/compress.js explicitly shields specific syntactic constructs from modification. This protection ensures that functional code and references remain intact for downstream language models.
Markdown Fenced Code Blocks
Fenced code blocks delimited by triple backticks are treated as protected regions. The algorithm identifies the opening and closing fences and stores the entire block—including the language identifier and inner content—before any prose compression occurs. According to the source implementation, these blocks are spliced back into their original positions after the surrounding text is shortened.
Inline Code Spans and Identifiers
Single backtick inline code spans, URLs, file paths, and identifiers are also excluded from compression. The tokenizer recognizes these patterns and marks them as immutable, ensuring that `npm install` or `compress()` remain exactly as written while adjacent prose is minimized.
The Compression Algorithm Pipeline
The compress function operates through a deliberate three-stage pipeline that separates protected content from compressible prose.
Tokenization and Boundary Detection
First, the input string is split into tokens while preserving the exact character offsets of protected patterns. The algorithm scans for regex matches that indicate code fences, inline backticks, and path-like strings, storing these segments in a separate registry.
Selective Shortening of Prose
For tokens that are not part of a protected region, Caveman applies aggressive reduction techniques: stop-word removal, simple lemmatization, and deletion of unimportant adjectives. This stage targets only natural language content, leaving the protected registry untouched.
Re-assembly of Protected Content
Finally, the compressed prose tokens are reassembled with the original protected substrings inserted at their recorded positions. The result is a token-efficient string where markdown code blocks remain syntactically identical to the input.
Code Examples
Compressing Markdown with Fenced Code Blocks
When processing a document containing a JavaScript example:
const { compress } = require('./src/mcp-servers/caveman-shrink/compress');
const doc = `
# Setup
Install the package first.
\`\`\`bash
npm install caveman
\`\`\`
Then import the module.
`;
const { compressed } = compress(doc);
console.log(compressed);
The output retains the exact fenced block while compressing the surrounding explanation:
# Setup
Install package first.
```bash
npm install caveman
Then import module.
### Preserving Inline Code During Compression
Inline code spans are similarly protected:
```javascript
const text = "Use the `compress()` method to shrink descriptions.";
const { compressed } = compress(text);
// Result: "Use `compress()` method shrink descriptions."
The backtick-delimited function name remains unchanged while the surrounding words are minimized.
Source Files and Implementation
The compression behavior is implemented across three key files in the repository:
src/mcp-servers/caveman-shrink/compress.js– Contains the corecompressfunction and protected pattern detection logic.src/mcp-servers/caveman-shrink/index.js– Wraps the compression engine and applies it to MCP response fields.tests/test_mcp_shrink.js– Validates that fenced code blocks and inline code survive compression unchanged, including assertions that verify exact character preservation within backtick-delimited regions.
Summary
- Caveman identifies markdown fenced code blocks and inline code spans as protected patterns before compression begins.
- The algorithm in
compress.jsseparates these protected regions from prose during tokenization. - Only natural language tokens undergo shortening; code blocks are re-inserted unchanged during final assembly.
- This approach ensures that syntax-critical content remains valid for language models while minimizing overall token count.
Frequently Asked Questions
Does Caveman compress code inside markdown fenced blocks?
No. Caveman treats the entire content between triple backticks as a protected region. The opening fence, language identifier, code body, and closing fence are all preserved exactly as they appear in the input, ensuring that syntax highlighting and code integrity remain intact for downstream processing.
What other content does Caveman protect during compression?
Beyond code blocks, Caveman protects inline code spans (single backticks), URLs, file system paths, and identifiers. These patterns are detected via regex matching during the tokenization phase and are excluded from the selective shortening algorithm to prevent breaking references or executable commands.
How can I verify that my code blocks remain unchanged after compression?
You can run the compression locally using the compress function from src/mcp-servers/caveman-shrink/compress.js and compare outputs, or examine the test suite in tests/test_mcp_shrink.js which contains assertions verifying that fenced blocks survive the compression pipeline with character-perfect fidelity.
Is the compression lossy for text outside code blocks?
Yes. The compression algorithm aggressively removes stop words, articles, and unnecessary adjectives from prose sections to reduce token count. This lossy compression applies only to natural language; protected patterns including code blocks, URLs, and inline code remain lossless.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →