# How Streaming Request Compression Differs from Standard Non-Streaming Transforms in Caveman

> Discover how Caveman's streaming request compression differs from standard non-streaming transforms. Learn about MCP tool requirements and synchronous API calls for efficient data processing.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-09-04

---

**Streaming request compression in Caveman requires the optional MCP (Model-Cache-Proxy) tool to apply marker-only metadata recording, whereas standard compression processes the entire payload through a direct synchronous API call to the compression engine.**

Caveman is an LLM gateway optimization toolkit that minimizes token costs through intelligent payload reduction. Understanding how **streaming request compression** diverges from standard non-streaming transforms is critical for developers optimizing real-time chat and tool-output scenarios.

## Standard Non-Streaming Compression Architecture

Standard compression operates on fully materialized request bodies before they reach the upstream provider. This pathway handles any payload where the complete byte stream is known upfront.

When you invoke the SDK-level `Cave.compress` method in [`packages/sdk/typescript/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts), the client sends the payload to the engine via the `/sdk/v1/compress` endpoint. The `compress` function (around line 1650) returns a compact representation together with a token-saving report. Because the engine receives the complete payload, it can perform byte-level replacement with full context awareness.

The CLI’s `compress` verb forwards identical requests for static files:

```bash

# Compress a fully materialized file and view token savings

caveman compress < my_prompt.txt

```

This approach applies to any request whose body is fully realized before the HTTP call—including file inputs, SDK API calls, or CLI commands without streaming transport enabled.

## Streaming Request Compression Architecture

Streaming compression addresses LLM-chat or tool-output scenarios where the request body emits incrementally. Unlike standard compression, this pathway cannot re-encode bytes mid-stream and instead relies on context reconstruction.

According to the binary description in [`packages/cli/src/agents.generated.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/agents.generated.ts) (line 2106), the `caveman-mcp` binary "lets streaming requests compress" by inserting a **marker-only compression layer**. When MCP is installed and active, the proxy records stream metadata—including tool calls and system messages—rather than the raw bytes. The proxy then regenerates the exact same byte stream on the recipient side using this stored context.

Without the MCP binary installed, the proxy treats streaming traffic as a black box. As noted in the runtime comments in [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts) (around line 3596), streaming turns simply "pass through uncompressed" when the engine MCP tools are not injected.

## Key Differences Between Compression Modes

| Capability | Standard Compression | Streaming Compression |
|------------|---------------------|---------------------|
| **Payload handling** | Full byte replacement | Marker-only metadata recording |
| **Requirements** | Always available | Requires `caveman-mcp` binary |
| **Timing** | Pre-request | Real-time during stream |
| **Token metrics** | `compression_tokens_saved` | `would_cut_stream_tokens` |

The telemetry reporting in [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts) (lines 4710–4718) distinguishes these pathways through separate metric names. Non-streaming compression reports direct token reductions, while streaming scenarios report potential savings via `would_cut_stream_tokens` because the actual byte reduction happens through context deduplication across turns.

## Implementation Examples

### Standard Compression with the TypeScript SDK

Use the `compress` method when working with complete JSON payloads:

```typescript
import { Cave } from "@caveman/sdk";

const cave = new Cave({ baseURL: "http://localhost:8080" });
const payload = JSON.stringify({ prompt: "Explain quantum computing" });

// Sends to /sdk/v1/compress endpoint
const result = await cave.compress(payload, { toon: true });

console.log("Compressed payload:", result.compressed);
console.log("Tokens saved:", result.tokens_saved);

```

### Enabling Streaming Compression via CLI

First install the required MCP binary, then run streaming commands:

```bash

# Install the MCP recovery tool (one-time setup)

caveman mcp install

# Execute streaming chat with compression active

caveman chat --model gpt-4o-mini --stream

```

When active, the CLI displays context savings after the session ends:

```

cut context rides every later turn — worth ~123 tokens of the sent total

```

This output reflects the `would_cut_stream_tokens` metric. If MCP is missing, the warning "streaming turns ... pass through uncompressed" appears instead.

## Summary

- **Standard compression** processes fully materialized payloads through the `/sdk/v1/compress` endpoint, replacing original bytes with compact representations before the upstream request.
- **Streaming compression** requires the `caveman-mcp` binary to record metadata markers and reconstruct streams on the far side, as raw byte replacement is impossible during incremental transmission.
- Without MCP installed, streaming traffic passes through the proxy unchanged and uncompressed.
- Token savings metrics differ between modes: `compression_tokens_saved` for standard requests versus `would_cut_stream_tokens` for streaming scenarios.

## Frequently Asked Questions

### What happens to streaming requests if MCP is not installed?

When the `caveman-mcp` binary is missing or disabled, the proxy forwards streaming traffic without modification. As implemented in [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts), the runtime emits a warning that "streaming turns ... pass through uncompressed" because the engine lacks the recovery tools needed to reconstruct context across stream increments.

### What token savings metric applies to streaming compression?

Streaming compression reports savings through the `would_cut_stream_tokens` metric, visible in the telemetry handling code at line 4712 of [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts). This differs from standard compression which reports `compression_tokens_saved`, reflecting the architectural distinction between context reconstruction and direct byte replacement.

### Can I use standard compression on streaming chat responses?

No. Standard compression requires the entire payload to be materialized before the API call. Streaming chat responses emit data incrementally through `Readable` streams (imported in [`packages/cli/src/proxy-fetch.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/proxy-fetch.ts)), making them incompatible with the synchronous `/sdk/v1/compress` endpoint used by the `Cave.compress` method.

### Where is the streaming compression logic implemented?

The streaming infrastructure relies on the `Readable` stream handling in [`packages/cli/src/proxy-fetch.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/proxy-fetch.ts) and the MCP binary registration in [`packages/cli/src/agents.generated.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/agents.generated.ts). The actual marker-only compression logic executes within the `caveman-mcp` proxy layer, which records stream metadata rather than compressing raw bytes.