# Caveman vs Caveman‑Code: Understanding the Difference Between Mouth‑Only and Full Agent Modes

> Understand the difference between caveman-code full agent mode and caveman mouth-only mode. Learn how each approach compresses AI agent outputs and interactions for efficient use.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-07-08

---

**`caveman`** is a lightweight skill that compresses only AI agent outputs (the "mouth"), while **`caveman‑code`** is a complete agent wrapper that compresses the entire interaction stack including prompts and tool calls.

The JuliusBrussee/caveman repository provides two distinct solutions for reducing token consumption in AI coding workflows. Understanding the difference between caveman-code and caveman determines whether you need a simple output filter for an existing agent or a comprehensive replacement that optimizes the full interaction lifecycle. Both tools implement the same compression algorithms, but they operate at fundamentally different layers of the AI stack.

## What Is Caveman? (The Mouth‑Only Skill)

### Scope and Functionality

`caveman` functions as a **skill/plugin** that intercepts an existing agent's natural‑language responses after generation completes. According to the source code in [`bin/lib/opencode-agent.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/opencode-agent.js), it hooks into the agent's output stream and applies compression solely to the conversational "mouth" — the explanatory text surrounding code, commands, and reasoning traces.

This surgical approach leaves underlying reasoning, executable code blocks, error messages, and tool outputs completely untouched. The agent retains full internal knowledge and functional capabilities; it simply communicates more efficiently through the compression hook defined in the repository settings.

### Token Savings and Configuration

The mouth‑only implementation delivers approximately **65% fewer output tokens** by rewriting verbose explanations into terse speech. Compression levels are defined in [`bin/lib/settings.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/settings.js), which supports multiple modes:

- **Default "full" compression**: Maximum token reduction for cost‑sensitive operations
- **"Lite" mode**: Minimal compression when readability takes priority over savings

Because `caveman` runs as a post‑processing hook rather than modifying the agent core, it adds negligible latency to existing workflows.

## What Is Caveman‑Code? (The Full Agent Wrapper)

### Complete Stack Integration

`caveman‑code` (distributed as the npm package `@juliusbrussee/caveman-code`) replaces the entire agent execution pipeline. Unlike the mouth‑only variant, this implementation wraps prompt handling, tool‑calling logic, planning modules, and output generation. The entry point resides in [`bin/install.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/install.js) of the separate `caveman-code` repository.

By sitting at the outermost layer of the interaction stack, this wrapper compresses not only final responses but also **internal prompts**, tool descriptions, system messages, and intermediate reasoning steps before they reach the language model.

### Dual Token Reduction

The full‑agent approach achieves roughly **2× fewer total tokens** compared to vanilla Codex‑style agents. This reduction spans both input (prompts, context, tool schemas) and output (responses, explanations), delivering superior cost savings and context‑window efficiency across the complete interaction cycle. For high‑volume automation or complex multi‑step coding tasks, this comprehensive compression prevents context window exhaustion while maintaining functional equivalence.

## Technical Implementation Details

### Caveman Core Files

The mouth‑only skill relies on two critical files in the JuliusBrussee/caveman repository:

- **[`bin/lib/opencode-agent.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/opencode-agent.js)**: Contains the post‑generation hook that applies compression to natural‑language outputs
- **[`bin/lib/settings.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/settings.js)**: Defines the compression algorithms, regex patterns, and level configurations used by the hook

### Caveman‑Code Architecture

The full agent wrapper operates from a distinct repository structure. The CLI script **[`bin/install.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/install.js)** serves as the entry point for the standalone package, initializing the complete wrapper that intercepts prompts before they reach the model and compresses responses before returning them to the user.

## Installation and Usage Patterns

### Installing the Mouth‑Only Skill

Add `caveman` to existing agents like Claude Code or OpenCode without changing your underlying infrastructure:

```bash

# Install the skill globally

npm install -g caveman

# Activate during any session

/caveman          # Enable full compression (default)

/caveman lite     # Enable minimal compression

```

### Installing the Full Agent

Replace your existing agent entirely with the token‑optimized wrapper:

```bash

# Install the complete agent package

npm install -g @juliusbrussee/caveman-code

# Launch with specific model and planning configurations

caveman-code run --model=gpt-4o --plan-mode

```

## When to Use Each Mode

### Use Caveman When...

You need a **drop‑in enhancement** for an existing agent without migrating workflows. This suits teams already using Claude Code, OpenCode, or similar tools who want immediate token savings on output costs. The mouth‑only mode requires zero changes to prompts, tool configurations, or deployment scripts.

### Use Caveman‑Code When...

You require **end‑to‑end efficiency** across both input and output token counts. This is essential for high‑volume automation, CI/CD pipelines, or complex refactoring tasks where context window limits constrain multi‑step operations. The full agent wrapper maximizes cost savings by compressing system prompts and tool descriptions that repeat across every interaction.

## Summary

- **`caveman`** is a post‑processing skill that compresses only natural‑language agent outputs, saving approximately **65% on output tokens** while leaving code and reasoning intact
- **`caveman‑code`** is a complete agent replacement that compresses prompts, tool calls, and outputs, reducing **total token usage by roughly 50%** (2× fewer tokens) across the full interaction lifecycle
- The mouth‑only implementation lives in [`bin/lib/opencode-agent.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/opencode-agent.js) and [`bin/lib/settings.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/settings.js) of the JuliusBrussee/caveman repository
- The full agent entry point is [`bin/install.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/install.js) in the `@juliusbrussee/caveman-code` package
- Install `caveman` for quick integration with existing agents; install `caveman‑code` for maximum token efficiency and context window management

## Frequently Asked Questions

### Can I use caveman and caveman‑code together?

No. `caveman` is designed as a plugin for existing agents, while `caveman‑code` is a standalone replacement. Running both would apply double compression and likely degrade output quality or cause parsing errors. Choose the mouth‑only skill if you want to keep your current agent infrastructure, or migrate to the full agent wrapper for comprehensive input and output savings.

### Does caveman affect code quality or execution?

No. The compression algorithm in [`bin/lib/opencode-agent.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/opencode-agent.js) specifically targets natural‑language explanations and conversational filler. It preserves code blocks, terminal commands, error traces, JSON tool outputs, and reasoning chains exactly as generated by the underlying agent. The functional capabilities and accuracy of the agent remain unchanged; only the verbosity of human‑readable explanations is reduced.

### Which option reduces API costs more effectively?

`caveman‑code` provides greater cost reduction because it compresses both input prompts and output responses. While `caveman` reduces output tokens by 65%, the full agent wrapper cuts total token usage approximately in half. For high‑volume operations where input tokens dominate costs — such as long context windows with extensive tool schemas — `caveman‑code` delivers superior economies across the complete billing cycle.

### Is there a performance overhead for either implementation?

Both implementations add minimal latency. `caveman` runs as a post‑processing hook after generation completes, adding milliseconds to render compressed text through the logic in [`bin/lib/opencode-agent.js`](https://github.com/JuliusBrussee/caveman/blob/main/bin/lib/opencode-agent.js). `caveman‑code` integrates compression into the generation pipeline itself, with optimizations in the wrapper logic that prevent significant slowdowns despite handling larger context windows and more complex prompt preprocessing.