How Token Savings Is Calculated in the /caveman-stats Command: A Complete Technical Breakdown

The /caveman-stats command calculates token savings by parsing the Claude session log to count output tokens, attributing them to specific Caveman modes, applying benchmark compression ratios to estimate saved tokens, and converting those savings into USD using model-specific pricing.

The /caveman-stats command in the JuliusBrussee/caveman repository provides detailed insights into how much token usage and cost the Caveman mode-saving tool prevents. Understanding how token savings is calculated in caveman-stats requires examining the session parsing, mode attribution, and benchmark-based estimation logic implemented in the source code.

Locate and Parse the Claude Session Log

The calculation begins by locating the active session file. The hook receives the path via the --session-file flag (injected by the mode tracker) or discovers the most recent session automatically using findRecentSession() at lines 55-75.

Once located, parseSession() processes the JSONL log file in src/hooks/caveman-stats.js. This function walks each line and aggregates critical metrics:

  • outputTokens from usage.output_tokens (line 95)
  • cacheReadTokens from usage.cache_read_input_tokens (line 96)
  • turns incremented per assistant message (line 97)
  • model from the first message.model encountered (line 98)
  • messages array with timestamps and token counts (lines 99-104)
// Excerpt from parseSession()
outputTokens    += usage.output_tokens           || 0;
cacheReadTokens += usage.cache_read_input_tokens || 0;

Attribute Tokens to Specific Caveman Modes

Caveman operates in different modes (full, lite, etc.), and the stats command must attribute token usage accurately rather than crediting the entire session to the current mode. The attributeByMode() function (lines 84-90) achieves this by:

  1. Reading the active mode flag (.caveman-active) and its modification time
  2. Loading the transition log (.caveman-mode-log.jsonl) via readModeLog()
  3. Matching each message timestamp to the most recent mode transition

The result is an attribution map that maps each mode to its token count, plus any unknownTokens that could not be matched to a specific mode.

const attribution = attributeByMode({
  messages: parsed.messages,
  modeLog,
  mode,
  flagMtimeMs,
  outputTokens: parsed.outputTokens,
});

Estimate Savings Using Compression Benchmarks

The core calculation happens in deriveSavings() (lines 44-50). The script references a COMPRESSION map at line 19 that defines benchmark ratios for each mode. Currently, only the full mode has a defined ratio of 0.65 (representing 65% token savings).

The formula reverses the compression to estimate how many tokens would have been generated without Caveman:

// From deriveSavings() - lines 44-47
estSavedTokens += Math.round(tokens / (1 - ratio)) - tokens;

For example, if 350 tokens were generated with a 65% savings ratio, the calculation estimates that approximately 1,000 tokens would have been required without Caveman, resulting in 650 saved tokens.

Convert Saved Tokens to USD

After calculating the estimated saved tokens, the script converts this value to dollars using the MODEL_OUTPUT_PRICE_PER_M table defined at lines 26-39. This hard-coded mapping associates model name prefixes with their USD cost per million output tokens.

The priceForModel() function selects the first matching prefix (most specific first), and deriveSavings() applies the pricing:

// Lines 48-50
const price = priceForModel(model);
const estSavedUsd = price !== null 
  ? (estSavedTokens / 1_000_000) * price 
  : 0;

The formatUsd() function (lines 49-53) then formats this value for human-readable display.

Detect Passive Memory Compression Savings

Beyond active token reduction, findCompressedPairs() (lines 113-144) scans the project for *.original.md / *.md file pairs created by the caveman-compress skill. It calculates:

  • Byte difference between original and compressed files
  • Token equivalent using bytesSaved / 4 (approximate bytes per token)
  • Count of compressed files

This provides a separate metric for "Memory compressed" tokens that supplements the runtime savings calculation.

Aggregate Lifetime Statistics

Every execution appends session totals to a hidden history file (.caveman-history.jsonl). The aggregateHistory() function (lines 63-84) maintains persistent metrics by:

  1. Reading the history file while keeping only the latest entry per session ID
  2. Summing output_tokens, est_saved_tokens, and est_saved_usd across all sessions

The humanizeTokens() function (lines 302-307) converts these large numbers into compact status-line suffixes like ⛏ 12.4k, which are written to .caveman-statusline-suffix for display in the Claude UI.

Rendering the Output

The formatStats() function (lines 447-462) assembles the final multi-line report displayed when users run /caveman-stats. For sharing purposes, formatShare() (lines 330-340) generates a condensed, tweet-ready summary.

Summary

  • Session parsing in parseSession() extracts raw token counts from Claude's JSONL logs at src/hooks/caveman-stats.js lines 95-104.
  • Mode attribution via attributeByMode() ensures savings are credited only to the active Caveman mode during specific time periods.
  • Benchmark ratios defined in the COMPRESSION map (line 19) estimate how many tokens would have been used without Caveman.
  • Savings formula Math.round(tokens / (1 - ratio)) - tokens calculates the difference between actual and estimated token usage.
  • USD conversion uses the MODEL_OUTPUT_PRICE_PER_M pricing table and priceForModel() to monetize the savings.
  • File compression detection via findCompressedPairs() adds passive memory savings to the total.
  • Lifetime aggregation in aggregateHistory() maintains running totals across all sessions for status-line display.

Frequently Asked Questions

How does caveman-stats know which mode was active for each message?

The attributeByMode() function cross-references message timestamps against the .caveman-mode-log.jsonl transition log. It matches each message to the most recent mode change before that timestamp, ensuring tokens are attributed to the correct mode rather than the current active flag.

What compression ratio does caveman-stats use for calculations?

According to the COMPRESSION map at line 19, the full mode uses a 0.65 ratio (65% savings). This means for every token generated, the tool estimates that 2.86 tokens would have been required without Caveman (calculated as 1 / (1 - 0.65)).

Why does the token savings calculation only consider output tokens?

The deriveSavings() function focuses exclusively on output_tokens because Caveman specifically optimizes Claude's response generation through context compression and mode-specific prompting. Input tokens (cache reads and writes) are tracked for reporting but not included in the savings estimate, as the optimization primarily reduces response length rather than prompt size.

Where does the pricing data come from in caveman-stats?

The pricing data resides in the hard-coded MODEL_OUTPUT_PRICE_PER_M object at lines 26-39 of src/hooks/caveman-stats.js. This table maps model name prefixes (like claude-3-5-sonnet) to their USD cost per million output tokens, allowing the script to calculate monetary savings without external API calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →