How Caveman Achieves Token Savings on Agent Outputs and Inputs: 2 Core Mechanisms Explained
Caveman reduces token costs through tool‑schema reduction and context compaction (summarization), trimming unused tool definitions and replacing long conversation histories with compact summaries.
The open‑source Caveman framework helps developers build LLM agents while actively minimizing token consumption. By implementing two complementary optimizations in its core runtime, Caveman lowers both input and output token bills. This guide examines how these mechanisms work according to the JuliusBrussee/caveman source code, where they live in the codebase, and how to leverage them in your own applications.
Tool‑Schema Reduction: Sending Only What You Need
Most agent frameworks transmit the complete catalog of available tools on every turn—even when only a subset is relevant. Caveman takes a smarter approach.
How Tool‑Schema Reduction Works
In packages/sdk/typescript/src/index.ts, the SDK computes a reduced set of tools actually required for the current request and forwards only those definitions to the model. The full and reduced catalogs are tracked separately:
tokensBefore: token count of the full tool catalogtokensAfter: token count of the reduced catalogtokensSaved: the computed delta between them
This logic appears in lines 22–31 of the TypeScript SDK index, where the RunResult interface exposes these fields to callers.
Why This Saves Money
Trimming unused tool definitions directly shrinks the payload sent to the LLM. Since providers bill by token count, every eliminated definition lowers the input cost. The client receives concrete visibility into these savings through the SDK's public API.
interface RunResult {
/** Tokens before the reduction (full catalog). */
tokensBefore: number;
/** Tokens after the reduction (reduced catalog). */
tokensAfter: number;
/** Tokens actually saved by the reduction. */
tokensSaved: number;
/** Inferred tokens saved by a compaction summary. */
tokensSavedInferred?: number;
}
Context Compaction: Summarizing to Stay Within Budget
When conversations grow long, agents risk hitting context window limits. Caveman solves this through context compaction—an intelligent summarization process that replaces raw message history with a compact representation.
The Compaction Engine Architecture
The compaction system lives in packages/agent/src/compaction.ts (lines 6–25 and 45–51). It implements a budget‑aware algorithm that:
- Monitors session token count against the model's context limit
- Evicts old messages when thresholds are exceeded
- Invokes a summarizer LLM to produce a concise state representation
- Reserves tokens for both the summarize call and subsequent working call
Token Savings Calculation
The engine models cost explicitly through CompactionOptions:
tokensBefore: tokens in the original historytokensAfter: tokens in the compacted summarytokensSavedInferred: estimated net reduction
Compaction only triggers when projected savings outweigh the cost of the extra summarize request. This guardrail prevents wasteful summarization of short contexts.
Why Summarization Beats Truncation
Raw conversation history often spans thousands of tokens. A well‑crafted summary captures essential state in a few hundred tokens. The net effect is a negative token delta for the turn—reported as tokensSavedInferred in run results. This mechanism appears in the code aggregation logic at packages/agent/src/code.ts (lines 604–821).
Practical Implementation: Using Token Savings Features
Inspecting Savings via the TypeScript SDK
import { Cave } from "@caveman/sdk";
async function demo() {
const cave = new Cave({
baseURL: "https://cave.example.com",
apiKey: "CAVE_API_KEY"
});
const result = await cave.run({
model: "gpt-5.5",
messages: [{ role: "user", content: "Explain quantum entanglement." }],
});
console.log(`Input tokens: ${result.tokensBefore}`);
console.log(`Reduced tokens: ${result.tokensAfter}`);
console.log(`Tokens saved by schema reduction: ${result.tokensSaved}`);
if (result.tokensSavedInferred) {
console.log(`Additional tokens saved by compaction: ${result.tokensSavedInferred}`);
}
}
The result object surfaces both schema savings (tokensSaved) and compaction savings (tokensSavedInferred).
Configuring Compaction Behavior
import { Cave, CompactionOptions } from "@caveman/sdk";
const compactionOpts: CompactionOptions = {
minTokensToCompact: 2048, // Trigger threshold
reserveTokens: 300, // Budget for summarizer call
};
await cave.run({
model: "gpt-5.5",
messages: longHistory,
compaction: compactionOpts,
});
Adjusting CompactionOptions lets you tune when summarization occurs and how aggressively tokens are reclaimed.
Monitoring via CLI
caveman run --model gpt-5.5 --prompt "Explain recursion" \
| grep "estimated tokens saved"
Sample output:
· 452 estimated tokens saved (inferred; counter basis estimated_bytes_div_4)
This formatting logic resides in packages/cli/src/index.ts (lines 10044–10045).
Budget Enforcement: Preventing Overruns While Saving
The packages/agent/src/budget.ts module implements token‑budget enforcement with built‑in reservation for compaction calls. This ensures agents:
- Never exceed allocated token quotas
- Still trigger compaction when savings are favorable
- Maintain predictable cost profiles
The budget system cooperates with the compaction engine to reserve tokens upfront, eliminating surprise overages from summarize operations.
Summary
- Tool‑schema reduction in
packages/sdk/typescript/src/index.tscomputes minimal tool sets, cutting input tokens before they reach the model - Context compaction in
packages/agent/src/compaction.tssummarizes long histories, replacing thousands of tokens with compact state representations - Budget reservation in
packages/agent/src/budget.tssafeguards against overruns while permitting aggressive optimization - Both mechanisms expose metrics through
RunResult(tokensSaved,tokensSavedInferred) for cost tracking and optimization
Frequently Asked Questions
How do I see exactly how many tokens Caveman saved on a request?
The SDK's RunResult object contains tokensSaved for schema reduction and tokensSavedInferred for compaction gains. Both fields populate automatically on every cave.run() call.
Does compaction always reduce token costs?
No. The compaction engine in packages/agent/src/compaction.ts only triggers when tokensSavedInferred is projected to be positive—that is, when summary cost is lower than the history it replaces. Short conversations skip compaction entirely.
Can I disable token‑saving features for debugging?
While the SDK doesn't expose a global "disable optimization" flag, you can effectively bypass compaction by setting minTokensToCompact to a very high value in CompactionOptions. Schema reduction is integral to the SDK's request pipeline.
What happens if the summarizer itself exceeds the reserved budget?
The budget system in packages/agent/src/budget.ts reserves tokens specifically for the compaction call. If a summarizer overruns, the error propagates to the caller—preventing silent cost explosions while maintaining transparency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →