How Streaming Request Compression Differs from Standard Non-Streaming Transforms in Caveman
Streaming request compression in Caveman requires the optional MCP (Model-Cache-Proxy) tool to apply marker-only metadata recording, whereas standard compression processes the entire payload through a direct synchronous API call to the compression engine.
Caveman is an LLM gateway optimization toolkit that minimizes token costs through intelligent payload reduction. Understanding how streaming request compression diverges from standard non-streaming transforms is critical for developers optimizing real-time chat and tool-output scenarios.
Standard Non-Streaming Compression Architecture
Standard compression operates on fully materialized request bodies before they reach the upstream provider. This pathway handles any payload where the complete byte stream is known upfront.
When you invoke the SDK-level Cave.compress method in packages/sdk/typescript/src/index.ts, the client sends the payload to the engine via the /sdk/v1/compress endpoint. The compress function (around line 1650) returns a compact representation together with a token-saving report. Because the engine receives the complete payload, it can perform byte-level replacement with full context awareness.
The CLI’s compress verb forwards identical requests for static files:
# Compress a fully materialized file and view token savings
caveman compress < my_prompt.txt
This approach applies to any request whose body is fully realized before the HTTP call—including file inputs, SDK API calls, or CLI commands without streaming transport enabled.
Streaming Request Compression Architecture
Streaming compression addresses LLM-chat or tool-output scenarios where the request body emits incrementally. Unlike standard compression, this pathway cannot re-encode bytes mid-stream and instead relies on context reconstruction.
According to the binary description in packages/cli/src/agents.generated.ts (line 2106), the caveman-mcp binary "lets streaming requests compress" by inserting a marker-only compression layer. When MCP is installed and active, the proxy records stream metadata—including tool calls and system messages—rather than the raw bytes. The proxy then regenerates the exact same byte stream on the recipient side using this stored context.
Without the MCP binary installed, the proxy treats streaming traffic as a black box. As noted in the runtime comments in packages/cli/src/index.ts (around line 3596), streaming turns simply "pass through uncompressed" when the engine MCP tools are not injected.
Key Differences Between Compression Modes
| Capability | Standard Compression | Streaming Compression |
|---|---|---|
| Payload handling | Full byte replacement | Marker-only metadata recording |
| Requirements | Always available | Requires caveman-mcp binary |
| Timing | Pre-request | Real-time during stream |
| Token metrics | compression_tokens_saved |
would_cut_stream_tokens |
The telemetry reporting in packages/cli/src/index.ts (lines 4710–4718) distinguishes these pathways through separate metric names. Non-streaming compression reports direct token reductions, while streaming scenarios report potential savings via would_cut_stream_tokens because the actual byte reduction happens through context deduplication across turns.
Implementation Examples
Standard Compression with the TypeScript SDK
Use the compress method when working with complete JSON payloads:
import { Cave } from "@caveman/sdk";
const cave = new Cave({ baseURL: "http://localhost:8080" });
const payload = JSON.stringify({ prompt: "Explain quantum computing" });
// Sends to /sdk/v1/compress endpoint
const result = await cave.compress(payload, { toon: true });
console.log("Compressed payload:", result.compressed);
console.log("Tokens saved:", result.tokens_saved);
Enabling Streaming Compression via CLI
First install the required MCP binary, then run streaming commands:
# Install the MCP recovery tool (one-time setup)
caveman mcp install
# Execute streaming chat with compression active
caveman chat --model gpt-4o-mini --stream
When active, the CLI displays context savings after the session ends:
cut context rides every later turn — worth ~123 tokens of the sent total
This output reflects the would_cut_stream_tokens metric. If MCP is missing, the warning "streaming turns ... pass through uncompressed" appears instead.
Summary
- Standard compression processes fully materialized payloads through the
/sdk/v1/compressendpoint, replacing original bytes with compact representations before the upstream request. - Streaming compression requires the
caveman-mcpbinary to record metadata markers and reconstruct streams on the far side, as raw byte replacement is impossible during incremental transmission. - Without MCP installed, streaming traffic passes through the proxy unchanged and uncompressed.
- Token savings metrics differ between modes:
compression_tokens_savedfor standard requests versuswould_cut_stream_tokensfor streaming scenarios.
Frequently Asked Questions
What happens to streaming requests if MCP is not installed?
When the caveman-mcp binary is missing or disabled, the proxy forwards streaming traffic without modification. As implemented in packages/cli/src/index.ts, the runtime emits a warning that "streaming turns ... pass through uncompressed" because the engine lacks the recovery tools needed to reconstruct context across stream increments.
What token savings metric applies to streaming compression?
Streaming compression reports savings through the would_cut_stream_tokens metric, visible in the telemetry handling code at line 4712 of packages/cli/src/index.ts. This differs from standard compression which reports compression_tokens_saved, reflecting the architectural distinction between context reconstruction and direct byte replacement.
Can I use standard compression on streaming chat responses?
No. Standard compression requires the entire payload to be materialized before the API call. Streaming chat responses emit data incrementally through Readable streams (imported in packages/cli/src/proxy-fetch.ts), making them incompatible with the synchronous /sdk/v1/compress endpoint used by the Cave.compress method.
Where is the streaming compression logic implemented?
The streaming infrastructure relies on the Readable stream handling in packages/cli/src/proxy-fetch.ts and the MCP binary registration in packages/cli/src/agents.generated.ts. The actual marker-only compression logic executes within the caveman-mcp proxy layer, which records stream metadata rather than compressing raw bytes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →