Internal Architecture of Cavecrew Subagents: 3-Agent Pipeline for Context Efficiency

TLDR: The cavecrew meta-skill deploys three specialized subagents—investigator, builder, and reviewer—that use compressed, line-oriented output formats to reduce token consumption by approximately 60% compared to standard exploration agents.

The cavecrew subagents in the JuliusBrussee/caveman repository implement a delegation pattern that splits complex coding tasks across three purpose-built agents. By constraining each subagent to specific tools and enforcing a terse caveman-compressed output format, the architecture minimizes context window usage while maintaining high-precision code operations.

The Three Subagents and Their Specialized Roles

The cavecrew architecture divides work into three distinct pipelines, each defined in its own agent markdown file within the agents/ directory.

cavecrew-investigator

cavecrew-investigator is a read-only code locator defined in agents/cavecrew-investigator.md. It answers "where is X defined?" and "what calls Y?" using a restricted toolset:

  • Tools: Read, Grep, Glob, Bash
  • Model: haiku (specified in front-matter)
  • Output format: <path:line> — symbol — <≤ 6‑word note>

This compressed format yields approximately 60% fewer tokens than a plain Explore call, delivering results like src/pages/render.ts:12 — renderPage — renders the homepage instead of verbose prose.

cavecrew-builder

cavecrew-builder handles surgical edits on 1-2 files, as defined in agents/cavecrew-builder.md. It is architecturally constrained to small-scale modifications:

  • Tools: Read, Edit, Write, Grep, Glob
  • Model: Inherits the session default (no explicit override in front-matter)
  • Output format: <path:line‑range> — <change ≤ 10‑words> plus a verification line

Attempting to delegate a 5-file refactor to this agent wastes tokens because it returns too-big. per its contractual constraints.

cavecrew-reviewer

cavecrew-reviewer performs diff, branch, and file reviews according to agents/cavecrew-reviewer.md:

  • Tools: Read, Grep, Bash
  • Model: haiku
  • Output format: path:line: <emoji> <severity>: <problem>. <fix>.

This emits one-line findings with severity emojis (e.g., src/api/auth.ts:82: 🔴 bug: missing auth check before token refresh.), enabling rapid scanning without prose overhead.

How Context Reduction Works

The caveman-compressed front-matter declared in each agent file instructs the main thread to treat tool results as raw tokens. When a subagent finishes, its compressed result is injected verbatim into the main context.

Because the output is already a compact, line-oriented table rather than free-form prose, the main thread avoids allocating extra tokens for parsing. According to the source documentation in skills/cavecrew/SKILL.md, a 2,000-token Explore call costs approximately 2,000 tokens of context, while the same investigation from cavecrew-investigator costs only ~700 tokens. This 3× token reduction scales across multiple delegations, preventing context exhaustion during complex sessions (e.g., 20 delegations avoid window limits).

Configuration and Model Overrides

The architecture supports per-agent model selection without rebuilding the repository. The hook src/hooks/cavecrew-model-overrides.js injects environment variables into the agent front-matter at install time:

  • CAVECREW_INVESTIGATOR_MODEL
  • CAVECREW_BUILDER_MODEL
  • CAVECREW_REVIEWER_MODEL

This allows fine-tuning of model costs and capabilities for each specific pipeline stage while maintaining the compressed output contracts defined in the agent files.

Chaining Patterns for Complex Workflows

The meta-skill defined in skills/cavecrew/SKILL.md documents standard chaining patterns that maximize context efficiency:

  • Locate → Fix → Verify: The most common pattern spawns investigator to locate code, builder to apply the fix, and reviewer to verify the change.
  • Parallel Scout: Spawn 2-3 investigator calls simultaneously to gather definitions, callers, and tests, then aggregate results before proceeding to the next phase.

These patterns ensure that context-intensive work happens in parallel subagents rather than accumulating in the main thread's context window.

Practical Examples

Spawning an Investigator


# User prompt to main thread

Delegate to subagent: locate where the function `renderPage` is defined.

Injected result (~700 tokens instead of ~2,000):

Defs:
- src/pages/render.ts:12 — `renderPage` — renders the homepage
- src/components/header.ts:45 — `renderPage` — used in header component
2 defs, 0 callers.

This output follows the contract defined in agents/cavecrew-investigator.md (lines 38-44).

Surgical Edit with Builder


# Prompt to main thread

Spawn builder: fix typo in src/utils/helpers.ts line 27.

Compressed receipt:

src/utils/helpers.ts:27-27 — replace `recieve` with `receive`.
verified: re-read OK

Matches the format specified in agents/cavecrew-builder.md (lines 28-34).

Diff Review with Reviewer


# Prompt

Run reviewer on the latest PR diff.

One-line findings:

src/api/auth.ts:82: 🔴 bug: missing auth check before token refresh.
src/api/auth.ts:115: 🟡 risk: no timeout on external request.
totals: 1🔴 1🟡

Exactly the output contract in agents/cavecrew-reviewer.md (lines 24-30).

Summary

  • Three specialized subagents split work into read-only location (cavecrew-investigator), surgical editing (cavecrew-builder), and review (cavecrew-reviewer) phases.
  • Compressed output formats deliver approximately 60% token reduction versus vanilla exploration by using line-oriented tables instead of prose.
  • Model overrides via src/hooks/cavecrew-model-overrides.js enable per-agent configuration through environment variables without repository rebuilds.
  • Chaining patterns like Locate → Fix → Verify and Parallel Scout distribute context load across subagents to prevent window exhaustion.

Frequently Asked Questions

How much token reduction do cavecrew subagents provide?

The cavecrew subagents achieve approximately 60% token reduction compared to standard exploration agents, representing roughly 3× more efficient context usage. For example, a code location task consuming ~2,000 tokens in a vanilla Explore agent requires only ~700 tokens when handled by cavecrew-investigator.

What is the difference between cavecrew-investigator and cavecrew-builder?

cavecrew-investigator is strictly read-only, utilizing tools like Grep, Glob, and Bash to locate definitions and callers without modifying files. cavecrew-builder performs surgical edits using Edit and Write tools but is architecturally constrained to 1-2 file modifications; it returns too-big. if delegated larger refactoring tasks.

Can I configure different LLM models for each subagent?

Yes. The repository includes src/hooks/cavecrew-model-overrides.js, which injects environment variables (CAVECREW_INVESTIGATOR_MODEL, CAVECREW_BUILDER_MODEL, CAVECREW_REVIEWER_MODEL) into each agent's front-matter at install time. This allows you to assign cheaper models like haiku to investigators while using more capable models for builders, without modifying the core agent definitions.

What happens if I delegate a large refactoring task to cavecrew-builder?

The agent will return too-big. and refuse the task. According to the skill documentation in skills/cavecrew/SKILL.md, attempting to use the builder for 5-file refactors wastes tokens because the agent is explicitly designed for surgical edits on 1-2 files. Large refactors should be broken into smaller chunks or handled directly by the main thread.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →