# How Cavecrew Presets Compress Subagent Behavior in Caveman

> Discover how Cavecrew presets like investigator, builder, and reviewer compress subagent behavior in Caveman, slashing token usage by 60-80% via the POST SDK v1 compress endpoint.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: internals
- Published: 2026-09-04

---

**Cavecrew presets automatically compress subagent tool outputs through the Caveman Engine's `POST /sdk/v1/compress` endpoint, reducing token consumption by 60-80% while preserving byte-safe, machine-parseable structure for the main conversation context.**

Cavecrew is a specialized preset system within the **JuliusBrussee/caveman** repository that bundles three sub-agents—investigator, builder, and reviewer—into a unified workflow. Understanding how Cavecrew presets compress subagent behavior is essential for maintaining efficient token budgets, as the system automatically routes all tool outputs through a compression pipeline before injecting them back into the main thread.

## The Cavecrew Compression Architecture

The Cavecrew preset system operates through a deterministic three-stage pipeline that processes sub-agent outputs before they reach the main conversation context.

### The Three Sub-Agent Presets

Cavecrew bundles specialized agents that each produce structured tool outputs:

- **cavecrew-investigator**: Returns multi-line lists of code exploration results
- **cavecrew-builder**: Returns one-line change descriptions for specific file ranges  
- **cavecrew-reviewer**: Returns structured diff-review lines with severity indicators

Despite differing output shapes, all three presets share an identical compression step through the Caveman runtime.

### Runtime Compression Mode

When Cavecrew activates, the Caveman runtime starts in **compress mode** (`think.mode = "compress"`), configured via the CLI flag in [[`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts)](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts#L166). This mode activates the Engine-side compression endpoint that intercepts sub-agent returns.

As implemented in the source, when a sub-agent's tool output returns to the main thread, the runtime automatically calls `POST /sdk/v1/compress`. The **TypeScript SDK** delegates this through the `Cave.compress` method defined in [[`packages/sdk/typescript/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts)](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts#L35-L47), which forwards the payload to the Engine and handles byte-safe fallback logic (lines 40-44 ensure that if compression fails, the original payload passes through unchanged).

## Preset-Specific Output Formats and Token Savings

Each Cavecrew preset produces a distinctive pre-compression format, but all undergo the same compression transformation. According to the [**SKILL.md**](https://github.com/JuliusBrussee/caveman/blob/main/skills/cavecrew/SKILL.md#L31-L41) documentation, subagent tool results get injected into main context verbatim after compression.

### cavecrew-investigator

The investigator returns multi-line lists using the format `path:line — symbol — note`. Before compression, these outputs typically consume approximately **2,000 tokens** of prose from a vanilla exploration. After compression through the Caveman pipeline, this reduces to roughly **700 tokens**, representing a **65% reduction** while maintaining grep-parseable path structures.

### cavecrew-builder  

The builder outputs one-line change descriptions following `path:line-range — change ≤10 words`. This compact format typically saves **~400 tokens per edit** compared to uncompressed change descriptions, allowing the main context to retain precise location data without verbose explanations.

### cavecrew-reviewer

The reviewer generates structured diff-review lines using the pattern `path:line: <emoji> <severity>: <problem>. <fix>.`. This format achieves approximately **500 tokens of savings** versus full diff commentary, preserving critical severity indicators and proposed fixes in a scannable format.

## Practical Implementation Examples

### Using the TypeScript SDK for Manual Compression

You can invoke the compression pipeline directly through the SDK when building custom workflows:

```typescript
import { Cave } from '@caveman/sdk';

const cave = new Cave({ baseURL: 'https://api.caveman.dev' });

async function compressInvestigatorOutput(payload: string) {
  // Delegates to POST /sdk/v1/compress via Cave.compress
  const result = await cave.compress(payload, { 
    /* optional compression options */ 
  });
  
  console.log('Compressed payload:', result.compressed);
  console.log('Tokens saved:', result.tokensSaved);
}

```

The `compress` method in [[`packages/sdk/typescript/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts)](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts#L35-L47) guarantees byte-safe delivery, automatically falling back to the original payload if the Engine cannot compress the specific content type.

### Running Presets via the CLI

Activate compression mode through the CLI runtime:

```bash

# Start CLI in compress mode (default for Cavecrew workflows)

caveman --mode compress

# Invoke investigator preset with automatic compression

caveman cavecrew-investigator "Where is function foo defined?"

```

The CLI sets `env.CAVE_ENGINE_TOON` based on the compress mode flag at line 166 of [[`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts)](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts#L166), ensuring every sub-agent routes through the compression endpoint before returning to the main thread.

### Chaining Multiple Presets

In typical workflows, you chain all three presets while maintaining token efficiency:

```typescript
// Each call automatically triggers compression
const investigation = await cave.run('cavecrew-investigator', query);
const build = await cave.run('cavecrew-builder', investigation.selectedPath);
const review = await cave.run('cavecrew-reviewer', build.diff);

```

Because each `run` call compresses its output before injection, total token usage equals the sum of three compressed payloads rather than three raw tool outputs, preserving the main context budget for subsequent reasoning steps.

## Summary

- **Cavecrew presets** bundle investigator, builder, and reviewer sub-agents that automatically route outputs through the Caveman compression pipeline.
- The **TypeScript SDK** implements compression via `Cave.compress` in [[`packages/sdk/typescript/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts)](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts#L35-L47), delegating to `POST /sdk/v1/compress` with byte-safe fallback logic.
- **Token reductions** range from 400 to 1,300 tokens per sub-agent call depending on the preset, achieving 60-80% compression rates.
- The **CLI** enables compression mode through the `--mode compress` flag defined in [[`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts)](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts#L166).
- Compressed payloads remain **verbatim-injectable** and machine-parseable, allowing the main thread to grep paths and line numbers without decoding custom formats.

## Frequently Asked Questions

### What is the Cavecrew compression pipeline?

The Cavecrew compression pipeline is an automated process within the JuliusBrussee/caveman repository that intercepts sub-agent tool outputs and routes them through the Engine's `POST /sdk/v1/compress` endpoint before injecting them into the main conversation context. This pipeline activates when the runtime operates in compress mode (`think.mode = "compress"`), reducing token usage while preserving the structured data formats that downstream processing requires.

### How does the SDK handle compression failures?

The TypeScript SDK implements **byte-safe compression** with automatic fallback behavior. According to the source code comments in [[`packages/sdk/typescript/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts)](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts#L40-L44), if the Engine cannot compress a specific payload, the SDK passes the original content through unchanged rather than throwing an error or returning malformed data. This ensures workflow continuity even when compression is incompatible with specific output types.

### Can I use compression without the Cavecrew presets?

Yes. While Cavecrew presets automatically enable compression mode, you can invoke the compression endpoint independently using the `Cave.compress` method from the TypeScript SDK or by setting the `--mode compress` flag when running the CLI. The compression endpoint accepts any string payload and returns a token-reduced version suitable for context injection, regardless of whether the content originated from Cavecrew sub-agents or custom tools.

### How much token savings does compression provide?

Token savings vary by preset and output complexity. Based on the Cavecrew SKILL documentation, **cavecrew-investigator** typically reduces outputs from ~2,000 tokens to ~700 tokens (65% reduction), **cavecrew-builder** saves approximately 400 tokens per edit, and **cavecrew-reviewer** achieves roughly 500 tokens of savings per review cycle. These reductions allow the main conversation context to retain more history and reasoning capacity across multi-step workflows.