Caveman Skill vs Engine/Proxy: Understanding the Architecture
The Caveman skill is a static, model-side instruction that instructs the LLM to speak in terse "caveman" style, while the Engine/Proxy is a standalone Go service that programmatically compresses arbitrary byte payloads before they reach the model.
The JuliusBrussee/caveman repository provides two complementary mechanisms for token optimization: a behavioral skill that alters how the model generates responses, and a compression engine that transforms external data flowing through the system. While both reduce token consumption, they operate at fundamentally different architectural layers—one inside the model's context window and one outside as a standalone service.
Architectural Overview
Understanding the distinction begins with recognizing where each component executes relative to the large language model.
The Caveman Skill: Model-Side Behavior Control
The Caveman skill is a declarative markdown file (skills/caveman/SKILL.md) that the Caveman dispatcher injects as a system-prompt style instruction at conversation start. When activated via the /caveman slash command, the skill instructs the LLM to drop filler words, articles, and emojis while preserving technical accuracy. This execution happens entirely inside the model—the LLM itself interprets the rules and generates the final terse output.
Because the skill is pure configuration with no runtime logic, it is completely stateless. It influences only the model's language generation, leaving code blocks, error messages, and JSON payloads untouched. The skill remains active for the entire conversation until explicitly disabled, affecting every response the model produces.
The Engine and Proxy: Runtime Compression Service
The Engine (caveman-engine) and its companion Proxy (caveman-proxy) constitute a compiled Go service that runs outside the model's execution environment. Implemented across engine/engine.go and proxy/protocol.go, this system programmatically analyzes, compresses, and recovers arbitrary byte payloads—including prompts, tool results, and memory records—before they reach the model.
Unlike the skill, the engine is stateful: when performing lossy compression, it writes the original uncompressed bytes to a CCR (Compressed Content Recovery) store, typically SQLite-backed via engine/ccr, and returns a recovery handle. The proxy exposes these capabilities via a JSON-RPC/STDIO interface defined in proxy/CLAUDE.md, allowing any client (CLI tools, MCP agents, or external services) to invoke compression without embedding the engine library directly.
Key Differences
Execution Location The skill runs inside the model's context as a behavioral instruction, while the engine runs as an external binary that preprocesses payloads.
Scope of Effect The Caveman skill affects only linguistic output—it rewrites sentences but never modifies structured data. The engine compresses any byte payload, including logs, JSON tool outputs, and non-text files.
State Management
The skill is stateless with no persistence mechanism. The engine maintains state through the CCR store, enabling exact reconstruction of original data via the Retrieve method in engine/engine.go.
Implementation Type
The skill is a static markdown file requiring no compilation. The engine is native Go code with a registry of compressors in engine/compressors/registry.go that must be compiled into the binary.
Implementation Details
Skill Implementation
The skill definition resides at skills/caveman/SKILL.md, which describes intensity levels, persistence rules, and stylistic constraints. The dispatcher parses this file and inserts its contents into the system prompt. No compiled code or runtime state exists—the model simply follows the behavioral rules described in the markdown.
Engine Implementation
The core compression logic lives in engine/engine.go, which exposes four primary methods: Compress, Simulate, Retrieve, and Stats. The engine routes payloads through appropriate compressors located in engine/compressors/*, each targeting specific data formats like JSON or markdown. Safety checks determine whether compression is safe, and the CCR subsystem handles recovery data storage. The proxy wraps this functionality in proxy/protocol.go, exposing it via JSON-RPC over stdio.
Practical Usage Examples
Activating the Caveman Skill
Use the CLI to install and activate the behavioral skill:
# Install the skill (once)
caveman tools skills install caveman
# Activate for current session
caveman wrap --caveman full
This injects skills/caveman/SKILL.md into the system prompt, causing the model to respond in the terse caveman style for the duration of the conversation.
Calling the Compression Engine from Go
Programmatically compress payloads using the Go SDK:
package main
import (
"fmt"
"github.com/JuliusBrussee/caveman/engine"
"github.com/JuliusBrussee/caveman/engine/ccr"
)
func main() {
// Initialize CCR store (nil disables lossy compression)
store, _ := ccr.Open("./recovery.db")
eng := engine.New(store, nil)
input := []byte("Connection pooling reuses open DB connections.")
res, err := eng.Compress(input, engine.Options{Mode: engine.ModeFull})
if err != nil {
panic(err)
}
fmt.Printf("Saved %.1f%% (%d → %d tokens)\n",
res.Ratio*100, res.TokensBefore, res.TokensAfter)
}
The Compress method in engine/engine.go automatically selects the optimal compressor and returns a recovery handle if CCR is enabled.
Using the Proxy Interface
Start the proxy and compress data via JSON-RPC:
# Launch the proxy (starts caveman-engine internally)
caveman start
# Compress via stdin
echo "Technical documentation content..." | caveman compress
The proxy forwards the request to the engine's Compress method and returns the optimized payload, making the service accessible to non-Go clients.
When to Use Which
Choose the Caveman skill when you want the model itself to generate concise, direct language without filler words. This reduces the token count of the model's natural language responses but preserves all technical content exactly as generated.
Choose the Engine/Proxy when you need to compress incoming context, tool outputs, or memory records before they enter the model's context window. This is essential for reducing input tokens in high-volume data scenarios or when storing conversation history.
Summary
- The Caveman skill is a static markdown instruction (
skills/caveman/SKILL.md) that runs inside the model and modifies linguistic output style only. - The Engine is a compiled Go binary (
engine/engine.go) that runs outside the model and performs algorithmic compression on arbitrary byte payloads. - The Proxy exposes the engine via JSON-RPC/STDIO (
proxy/protocol.go), enabling external services to compress data without direct library integration. - The skill is stateless and behavioral; the engine is stateful with CCR-backed recovery capabilities.
- The skill is triggered by slash commands; the engine is invoked programmatically through its API or CLI.
Frequently Asked Questions
Can I use the Caveman skill and Engine together?
Yes. The skill reduces the token count of the model's generated responses, while the engine compresses the input payloads and context data fed into the model. According to the JuliusBrussee/caveman source code, they serve complementary purposes and can be activated simultaneously via the CLI's wrap command with appropriate flags.
Does the Engine modify the model's behavior?
No. The Engine only transforms data payloads before they reach the model or after they are returned. It never influences how the model generates text. The Caveman skill, conversely, directly influences the model's output style by modifying the system prompt.
What happens if the Engine fails to compress data?
The Engine is designed to be fail-closed. If a compressor cannot parse the input or compression would not reduce token count, the engine passes the original data through unchanged with no CCR handle. This behavior ensures that data integrity is never compromised by failed compression attempts.
Is the Caveman skill available for custom LLM providers?
The skill is provider-agnostic because it is implemented as a markdown instruction file (skills/caveman/SKILL.md) that any dispatcher can inject into a system prompt. As long as the LLM supports system prompts and can follow behavioral instructions, the Caveman skill will function correctly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →