# The Difference Between Cavecrew and Vanilla Explore Agents: Context Usage Explained

> Understand the context usage difference between Cavecrew and vanilla Explore agents. Cavecrew agents use 65% less context, enabling complex workflows and preventing context window exhaustion.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-07-12

---

**Cavecrew agents consume approximately 65% less context than vanilla Explore agents by returning compressed, structured outputs (~700 tokens) instead of prose-heavy responses (~2,000 tokens), allowing complex multi-step workflows to complete without exhausting the model's context window.**

The **cavecrew** sub-agents in the `JuliusBrussee/caveman` repository offer a token-efficient alternative to the vanilla **Explore** agent when navigating codebases or performing structured tasks. While both leverage the same underlying Anthropic model, the critical difference between cavecrew and vanilla Explore agents lies in how they manage context usage—specifically, how much of the model's output is preserved in the main thread's conversation history. This distinction determines whether long-running sessions finish successfully or hit context limits after just a few delegations.

## Token Efficiency: The Core Difference

### Compressed Output Size

According to [`skills/cavecrew/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/cavecrew/SKILL.md) (lines 30-33), cavecrew implements **"caveman-compressed"** results that reduce token consumption by roughly 60% compared to raw model outputs. A typical `cavecrew-investigator` call returns approximately **700 tokens**, whereas the vanilla `Explore` agent generates around **2,000 tokens** for comparable tasks. This compression is achieved by stripping verbose prose and returning structured, concise data formats like file locations, function signatures, and minimal explanations.

### Context Budget Impact

The documentation explicitly states that this size difference represents **"the difference between context exhaustion and finishing the task"** ([`skills/cavecrew/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/cavecrew/SKILL.md), lines 30-33). When the main thread delegates to a sub-agent, the entire response is injected into the conversation context. Vanilla Explore's 2,000-token responses rapidly consume the available context window, limiting workflows to only a few delegation steps. Cavecrew's 700-token footprint preserves the main-thread budget, enabling chains of multiple sub-agent calls without truncation.

## When to Use Each Agent

### Structured Tasks: Use Cavecrew

Deploy cavecrew agents when you need concise, actionable results such as locating specific code, reviewing diffs, or making small edits. The [`agents/cavecrew-investigator.md`](https://github.com/JuliusBrussee/caveman/blob/main/agents/cavecrew-investigator.md) definition produces compressed outputs ideal for repetitive navigation tasks where prose explanations would waste tokens. Use this agent when building multi-step automation pipelines that require several context-preserving delegations.

### Explanatory Tasks: Use Vanilla Explore

Reserve the vanilla `Explore` agent for scenarios requiring rich natural-language commentary, architecture discussions, or detailed explanations of complex logic. When the goal is understanding rather than location—such as explaining authentication flows or design patterns—the full 2,000-token prose output provides necessary context that compressed formats cannot convey. The [`skills/cavecrew/README.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/cavecrew/README.md) (line 17) summarizes this distinction, recommending vanilla Explore for "richer natural-language explanation."

## Practical Implementation Examples

The following examples demonstrate the practical difference in context usage between the two approaches:

```javascript
// Using the vanilla Explore agent (prose output)
await caveman.run({
  agent: 'Explore',          // vanilla
  task: 'Explain how the login flow works.'
});
// → Returns a full-sentence explanation (~2k tokens)

```

```javascript
// Using cavecrew-investigator (compact, token-saving)
await caveman.run({
  agent: 'cavecrew-investigator',
  task: 'Where is the function `validateToken` defined?'
});
// → Returns a concise list:
//    src/auth.js:12 — `validateToken` — checks JWT signature
//    (≈ 700 tokens)

```

```javascript
// Chaining cavecrew for a refactor (still token-efficient)
await caveman.run({
  agent: 'cavecrew-investigator',
  task: 'Find all call-sites of `validateToken`.'
});

await caveman.run({
  agent: 'cavecrew-builder',
  task: 'Rename `validateToken` to `verifyJwt` in the two call-sites.'
});

await caveman.run({
  agent: 'cavecrew-reviewer',
  task: 'Review the diff for regressions.'
});
// Each step injects a compressed payload, preserving the main-thread budget.

```

## Technical Implementation Details

The token efficiency is enforced through model overrides defined in [`src/hooks/cavecrew-model-overrides.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/hooks/cavecrew-model-overrides.js), which configures the sub-agents to emit compressed structured outputs. This hook injects environment variables that control how the sub-agents format their responses, ensuring consistent ~700-token outputs that distinguish them from vanilla Explore.

## Summary

- **Cavecrew agents** return "caveman-compressed" outputs of approximately **700 tokens**, consuming roughly one-third of the context used by vanilla Explore.
- **Vanilla Explore** agents provide verbose prose responses of approximately **2,000 tokens**, suitable for explanatory tasks but costly for multi-step workflows.
- The **60% reduction** in token usage allows cavecrew to chain multiple sub-agent calls without exhausting the main thread's context window.
- Select **cavecrew** for code navigation, diff reviews, and structured edits; choose **vanilla Explore** for architecture commentary and detailed explanations.

## Frequently Asked Questions

### How much context does cavecrew actually save compared to vanilla Explore?

Cavecrew typically consumes **700 tokens** per call versus **2,000 tokens** for vanilla Explore, representing approximately **65% less context usage** per delegation. According to the source code in [`skills/cavecrew/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/cavecrew/SKILL.md) (lines 30-33), this compression prevents context exhaustion in long-running sessions that would otherwise fail after only a few vanilla Explore calls.

### Can I mix cavecrew and vanilla Explore agents in the same workflow?

Yes, but strategically. Use **vanilla Explore** for initial high-level architecture explanations where prose is valuable, then switch to **cavecrew** agents for subsequent implementation steps. This hybrid approach preserves context while maintaining code quality, as the compressed cavecrew outputs minimize the main-thread footprint during repetitive tasks.

### Does cavecrew sacrifice quality for token efficiency?

No. The [`agents/cavecrew-investigator.md`](https://github.com/JuliusBrussee/caveman/blob/main/agents/cavecrew-investigator.md) definition maintains accuracy by returning structured data (file paths, line numbers, function names) rather than natural language. The quality difference lies in format, not correctness—cavecrew provides actionable data while vanilla Explore provides explanatory prose.

### Which file controls the model overrides for cavecrew context compression?

The compression behavior is configured in [`src/hooks/cavecrew-model-overrides.js`](https://github.com/JuliusBrussee/caveman/blob/main/src/hooks/cavecrew-model-overrides.js), which injects environment variables that control how the sub-agents format their outputs. This hook ensures that cavecrew agents consistently return the compressed ~700-token responses that distinguish them from vanilla Explore.