How Caveman Learn Analyzes Agent History and Ranks Token Sinks for Optimization
Caveman Learn inspects an agent’s execution history, quantifies token consumption per segment, and ranks "token sinks" by cost impact to surface the biggest drivers of LLM spending.
Caveman Learn is the cost-analysis engine in the JuliusBrussee/caveman framework that transforms raw runtime telemetry into actionable optimization targets. By aggregating exact token counts from the agent’s memory store and calculating per-turn metrics, it produces a deterministic ranking of which history segments consume the most budget.
1. Gathering Raw History from the Runtime
Every turn, the Caveman runtime records a history segment—user messages, tool results, and system context—alongside a budget object that tracks precise token usage. This data is persisted by the Caveman Mem subsystem and serves as the ground truth for all downstream analysis.
In packages/agent/src/runtime.ts, the runtime captures three critical fields:
tokens_before: Token count entering the turntokens_after: Token count exiting the turnruntimeSegmentId: Unique identifier correlating budget entries with history segments
These fields record exact provider metrics including input, output, and cached tokens, ensuring the Learn phase never estimates what the runtime already measured.
2. Summarizing Per-Turn Metrics into LearnSink Objects
When caveman learn scan completes, the CLI reads the generated snapshot (caveman-learn.json) and walks the sessions array. For each session, it extracts tokens_before and tokens_after values and aggregates them into a LearnSink structure.
In packages/cli/src/index.ts (lines 10061‑10083), the constructor creates LearnSink instances that encapsulate:
- The raw token delta for the segment
- Session frequency metadata
- A unique
sink_idderived from the segment content
This structure normalizes disparate history events into a uniform schema ready for classification and sorting.
3. Classifying Sinks and Computing Token Rates
Before ranking, Caveman Learn classifies each sink into behavioral categories and computes standardized cost metrics. The classification happens in packages/cli/src/index.ts (lines 15383‑15401) via learnMeasuredPrefixSuffix and associated helpers.
Each sink receives a class label:
recurring_context: Repeated prompts carried across turnsload_bearing: Essential system prompts required for agent functioncache_efficiency: Segments exhibiting high cache hit ratios
The engine then calculates:
- tokens_per_turn: Average tokens added per interaction, derived from
(tokens_after - tokens_before)aggregated across sessions - tokens_per_day_rate: Projected daily cost based on observed session frequency and the
tokens_per_turnvalue
Helper functions counter and humanTokens format these numbers for both machine parsing and human readability.
4. Ranking Token Sinks by Cost Impact
The final stage sorts the LearnSink array descending by tokens_per_turn, placing the strongest cost drivers at the top of the report. This ranking loop in packages/cli/src/index.ts (lines 15384‑15401) produces the definitive list of optimization candidates.
For each ranked sink, the CLI outputs:
- title: Human-readable description of the history segment
- sink_id: Unique identifier for programmatic targeting
- class: The behavioral category (e.g.,
recurring_context) - tokens_per_turn and tokens_per_day_rate: Quantified cost metrics
A companion TUI in packages/cli/src/learn-tui.ts renders this data with progress bars and concise markdown views, allowing operators to scan the biggest sinks at a glance. The same ranking is exposable via caveman learn report --json for automated workflows.
Working with the Caveman Learn CLI
Generate a token sink report and inspect the ranking programmatically:
# 1. Run the learn scan to generate the JSON snapshot
$ caveman learn --json
# Creates ~/.caveman/reports/caveman-learn.json
// 2. Load and analyze the report programmatically
import { readFileSync } from "fs";
import { join } from "path";
const reportPath = join(
process.env.HOME ?? "",
".caveman",
"reports",
"caveman-learn.json"
);
const report = JSON.parse(readFileSync(reportPath, "utf8"));
// 3. Inspect ranked sinks (highest cost first)
report.sinks
.sort((a: any, b: any) => b.tokens_per_turn - a.tokens_per_turn)
.forEach((sink: any) => {
console.log(
`${sink.title} (ID=${sink.sink_id}) – ${sink.tokens_per_turn} t/turn [${sink.class}]`
);
});
Apply optimizations with dry-run verification:
# Generate an edit plan for a specific sink without modifying code
$ caveman learn apply tool_output_portfolio_abc --dry-run
Summary
- Raw capture:
packages/agent/src/runtime.tsrecords exact token counts (tokens_before,tokens_after) and history segments viaruntimeSegmentId - Aggregation:
packages/cli/src/index.ts(lines 10061‑10083) buildsLearnSinkobjects from thecaveman-learn.jsonsnapshot - Classification: Sinks are categorized as
recurring_context,load_bearing, orcache_efficiencyto distinguish trimmable fat from essential context - Metrics: The engine computes
tokens_per_turnandtokens_per_day_rateto project real cost impact - Ranking: Sinks are sorted descending by
tokens_per_turninpackages/cli/src/index.ts(lines 15384‑15401) and visualized inlearn-tui.ts
Frequently Asked Questions
How does Caveman Learn distinguish between essential and trimmable context?
Caveman Learn uses the classification field assigned during analysis. Sinks marked load_bearing represent essential system prompts that the engine avoids touching, while recurring_context or cache_efficiency sinks are flagged as safe candidates for compaction or truncation. This logic references the recoverable compression implementations in packages/agent/src/compaction.ts.
What file contains the logic for calculating tokens per turn?
The calculation occurs in packages/cli/src/index.ts within the learnMeasuredPrefixSuffix function (lines 15383‑15401). This code path aggregates the raw tokens_before and tokens_after values captured in packages/agent/src/runtime.ts and divides by session count to produce the tokens_per_turn average.
Can I automate fixes based on the token sink ranking?
Yes. The CLI supports caveman learn apply <sink_id> --dry-run, which consumes the same LearnSink metadata used for ranking to generate an edit plan. Because the ranking includes stable sink_id values derived from content hashing, automated scripts can target specific sinks without human-in-the-loop review.
Where does the runtime store the history segments analyzed by Learn?
History segments and their associated budget data are persisted by the Caveman Mem subsystem, referenced in packages/agent/src/runtime.ts. The caveman learn scan command later reads this persisted state from the agent’s memory store to produce the caveman-learn.json snapshot consumed by the CLI ranking logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →