How Cavecrew Presets Compress Subagent Behavior in Caveman
Cavecrew presets automatically compress subagent tool outputs through the Caveman Engine's POST /sdk/v1/compress endpoint, reducing token consumption by 60-80% while preserving byte-safe, machine-parseable structure for the main conversation context.
Cavecrew is a specialized preset system within the JuliusBrussee/caveman repository that bundles three sub-agents—investigator, builder, and reviewer—into a unified workflow. Understanding how Cavecrew presets compress subagent behavior is essential for maintaining efficient token budgets, as the system automatically routes all tool outputs through a compression pipeline before injecting them back into the main thread.
The Cavecrew Compression Architecture
The Cavecrew preset system operates through a deterministic three-stage pipeline that processes sub-agent outputs before they reach the main conversation context.
The Three Sub-Agent Presets
Cavecrew bundles specialized agents that each produce structured tool outputs:
- cavecrew-investigator: Returns multi-line lists of code exploration results
- cavecrew-builder: Returns one-line change descriptions for specific file ranges
- cavecrew-reviewer: Returns structured diff-review lines with severity indicators
Despite differing output shapes, all three presets share an identical compression step through the Caveman runtime.
Runtime Compression Mode
When Cavecrew activates, the Caveman runtime starts in compress mode (think.mode = "compress"), configured via the CLI flag in [packages/cli/src/index.ts](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts#L166). This mode activates the Engine-side compression endpoint that intercepts sub-agent returns.
As implemented in the source, when a sub-agent's tool output returns to the main thread, the runtime automatically calls POST /sdk/v1/compress. The TypeScript SDK delegates this through the Cave.compress method defined in [packages/sdk/typescript/src/index.ts](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts#L35-L47), which forwards the payload to the Engine and handles byte-safe fallback logic (lines 40-44 ensure that if compression fails, the original payload passes through unchanged).
Preset-Specific Output Formats and Token Savings
Each Cavecrew preset produces a distinctive pre-compression format, but all undergo the same compression transformation. According to the SKILL.md documentation, subagent tool results get injected into main context verbatim after compression.
cavecrew-investigator
The investigator returns multi-line lists using the format path:line — symbol — note. Before compression, these outputs typically consume approximately 2,000 tokens of prose from a vanilla exploration. After compression through the Caveman pipeline, this reduces to roughly 700 tokens, representing a 65% reduction while maintaining grep-parseable path structures.
cavecrew-builder
The builder outputs one-line change descriptions following path:line-range — change ≤10 words. This compact format typically saves ~400 tokens per edit compared to uncompressed change descriptions, allowing the main context to retain precise location data without verbose explanations.
cavecrew-reviewer
The reviewer generates structured diff-review lines using the pattern path:line: <emoji> <severity>: <problem>. <fix>.. This format achieves approximately 500 tokens of savings versus full diff commentary, preserving critical severity indicators and proposed fixes in a scannable format.
Practical Implementation Examples
Using the TypeScript SDK for Manual Compression
You can invoke the compression pipeline directly through the SDK when building custom workflows:
import { Cave } from '@caveman/sdk';
const cave = new Cave({ baseURL: 'https://api.caveman.dev' });
async function compressInvestigatorOutput(payload: string) {
// Delegates to POST /sdk/v1/compress via Cave.compress
const result = await cave.compress(payload, {
/* optional compression options */
});
console.log('Compressed payload:', result.compressed);
console.log('Tokens saved:', result.tokensSaved);
}
The compress method in [packages/sdk/typescript/src/index.ts](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts#L35-L47) guarantees byte-safe delivery, automatically falling back to the original payload if the Engine cannot compress the specific content type.
Running Presets via the CLI
Activate compression mode through the CLI runtime:
# Start CLI in compress mode (default for Cavecrew workflows)
caveman --mode compress
# Invoke investigator preset with automatic compression
caveman cavecrew-investigator "Where is function foo defined?"
The CLI sets env.CAVE_ENGINE_TOON based on the compress mode flag at line 166 of [packages/cli/src/index.ts](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts#L166), ensuring every sub-agent routes through the compression endpoint before returning to the main thread.
Chaining Multiple Presets
In typical workflows, you chain all three presets while maintaining token efficiency:
// Each call automatically triggers compression
const investigation = await cave.run('cavecrew-investigator', query);
const build = await cave.run('cavecrew-builder', investigation.selectedPath);
const review = await cave.run('cavecrew-reviewer', build.diff);
Because each run call compresses its output before injection, total token usage equals the sum of three compressed payloads rather than three raw tool outputs, preserving the main context budget for subsequent reasoning steps.
Summary
- Cavecrew presets bundle investigator, builder, and reviewer sub-agents that automatically route outputs through the Caveman compression pipeline.
- The TypeScript SDK implements compression via
Cave.compressin [packages/sdk/typescript/src/index.ts](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts#L35-L47), delegating toPOST /sdk/v1/compresswith byte-safe fallback logic. - Token reductions range from 400 to 1,300 tokens per sub-agent call depending on the preset, achieving 60-80% compression rates.
- The CLI enables compression mode through the
--mode compressflag defined in [packages/cli/src/index.ts](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts#L166). - Compressed payloads remain verbatim-injectable and machine-parseable, allowing the main thread to grep paths and line numbers without decoding custom formats.
Frequently Asked Questions
What is the Cavecrew compression pipeline?
The Cavecrew compression pipeline is an automated process within the JuliusBrussee/caveman repository that intercepts sub-agent tool outputs and routes them through the Engine's POST /sdk/v1/compress endpoint before injecting them into the main conversation context. This pipeline activates when the runtime operates in compress mode (think.mode = "compress"), reducing token usage while preserving the structured data formats that downstream processing requires.
How does the SDK handle compression failures?
The TypeScript SDK implements byte-safe compression with automatic fallback behavior. According to the source code comments in [packages/sdk/typescript/src/index.ts](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts#L40-L44), if the Engine cannot compress a specific payload, the SDK passes the original content through unchanged rather than throwing an error or returning malformed data. This ensures workflow continuity even when compression is incompatible with specific output types.
Can I use compression without the Cavecrew presets?
Yes. While Cavecrew presets automatically enable compression mode, you can invoke the compression endpoint independently using the Cave.compress method from the TypeScript SDK or by setting the --mode compress flag when running the CLI. The compression endpoint accepts any string payload and returns a token-reduced version suitable for context injection, regardless of whether the content originated from Cavecrew sub-agents or custom tools.
How much token savings does compression provide?
Token savings vary by preset and output complexity. Based on the Cavecrew SKILL documentation, cavecrew-investigator typically reduces outputs from ~2,000 tokens to ~700 tokens (65% reduction), cavecrew-builder saves approximately 400 tokens per edit, and cavecrew-reviewer achieves roughly 500 tokens of savings per review cycle. These reductions allow the main conversation context to retain more history and reasoning capacity across multi-step workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →