How RTK Identifies Missed Token Savings Opportunities in Claude Code History

TLDR: RTK's discover command scans local Claude Code session logs in ~/.claude/projects, extracts Bash commands using ClaudeProvider::extract_commands, classifies them via registry::classify_command, estimates potential token savings using output-length heuristics, and generates a report highlighting where RTK filtering could have reduced LLM token usage.

The RTK CLI tool from the rtk-ai/rtk repository includes a powerful analytics feature that retrospectively analyzes your Claude Code workflow. By processing historical session data stored locally on your machine, RTK identifies specific commands where token filtering would have saved money and context window space. The entire pipeline is orchestrated through src/discover/mod.rs and operates in five distinct stages.

The Five-Stage Discovery Pipeline

Stage 1: Session Discovery

The process begins in src/discover/provider.rs (lines 71-94) with ClaudeProvider::discover_sessions. This function walks the ~/.claude/projects directory to locate session files.

  • Recursively searches for *.jsonl files containing Claude Code conversation history
  • Optionally filters by project name or age (sessions from the last N days)
  • Returns a vector of file paths to be processed

Stage 2: Command Extraction

For each discovered session file, ClaudeProvider::extract_commands in src/discover/provider.rs (lines 38-63) parses the JSONL structure to identify executable commands.

  • Matches assistant tool_use blocks containing Bash commands
  • Pairs each command with the subsequent user tool_result block
  • Records the command string, output length in bytes, and error status
  • Stores results in ExtractedCommand structs for further analysis

Stage 3: Classification and Token Estimation

Each extracted command is fed to registry::classify_command in src/discover/registry.rs (lines 88-112). This applies the same rule engine RTK uses at runtime.

  • Regex-based classification: Categorizes commands as Supported, Unsupported, or Ignored
  • Special status detection (e.g., Passthrough for partial filtering)
  • Token estimation:
    • If output_len is known: uses len / 4 (approximate tokens per byte)
    • Otherwise: falls back to category_avg_tokens from historical averages
  • Savings calculation: output_tokens × estimated_savings_pct / 100

Stage 4: Bypass Detection

RTK specifically flags commands where filtering was explicitly disabled. The has_rtk_disabled_prefix function in src/discover/registry.rs (lines 381-398) detects commands starting with the environment flag RTK_DISABLED=1.

These commands are counted as missed opportunities because the user disabled RTK on a command that the tool could have filtered. This same detection logic is reused in src/analytics/gain.rs (lines 629-652) for consistency across analytics pathways.

Stage 5: Report Generation

Finally, discover::report::format_text and format_json in src/discover/report.rs (lines 74-118) aggregate the processed data into human-readable summaries or machine-readable JSON.

  • Calculates totals: commands analyzed, already-wrapped commands, supported-but-unfiltered commands, and RTK_DISABLED bypasses
  • Groups results by command category
  • Displays estimated token savings per group

Using the RTK Discover Command

To analyze your recent Claude Code usage, run the discover command with appropriate filters:


# Scan the current project for the last 30 days, showing top 10 missed savings

rtk discover --limit 10

For automation and dashboards, export to JSON:


# Machine-readable report for the last 7 days (CI dashboard compatible)

rtk discover --format json --since 7 > rtk-discover-report.json

To audit where you explicitly disabled RTK:


# Show only commands run with RTK_DISABLED=1

rtk discover --all --format text | grep RTK_DISABLED

Programmatic Integration

You can also invoke the discovery pipeline programmatically using the RTK Rust library. The run function handles the complete workflow:

use rtk::discover::run;

fn main() -> anyhow::Result<()> {
    // Find missed savings for all projects, last 14 days, text format
    run(
        None,          // no specific project filter
        true,          // scan all projects in ~/.claude/projects
        14,            // days of history to analyze
        20,            // limit results per section
        "text",        // output format: "text" or "json"
        0,             // verbose level
    )
}

Summary

  • Session scanning: ClaudeProvider::discover_sessions locates JSONL files in ~/.claude/projects with optional date filtering
  • Command extraction: Parses tool_use and tool_result blocks to capture Bash commands and their output sizes
  • Token estimation: Uses len / 4 heuristics or category averages in registry::classify_command to estimate potential savings
  • Bypass detection: has_rtk_disabled_prefix flags RTK_DISABLED=1 commands as explicit missed opportunities
  • Reporting: format_text and format_json in src/discover/report.rs convert aggregated data into actionable reports

Frequently Asked Questions

How accurate is RTK's token savings estimation?

RTK uses a conservative approximation of 4 bytes per token (dividing output length by 4) when actual token counts are unavailable. According to the source code in src/discover/registry.rs, this heuristic provides a reliable lower-bound estimate for budgeting purposes. When historical output lengths are missing, RTK falls back to category_avg_tokens based on command type.

What does RTK consider a "bypass"?

Any command prefixed with the environment variable RTK_DISABLED=1 is detected by the has_rtk_disabled_prefix function in src/discover/registry.rs (lines 381-398). These represent intentional disabling of RTK on commands that would otherwise have been supported and filtered, making them explicit missed savings opportunities rather than unsupported commands.

Can I export discovery data for CI dashboards?

Yes. The --format json flag outputs machine-readable JSON through format_json in src/discover/report.rs. This format includes structured data on estimated tokens, command categories, and savings percentages suitable for ingestion by analytics platforms, CI pipeline reports, or custom dashboards.

Does RTK discover sessions across all projects?

By default, RTK filters sessions by the current project context, but the --all flag instructs ClaudeProvider::discover_sessions to scan the entire ~/.claude/projects directory. You can also use --project <name> to target a specific project's history without changing directories.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →