# How RTK Identifies Missed Token Savings Opportunities in Claude Code History

> Learn how RTK identifies missed token savings in Claude Code history. RTK scans logs, classifies commands, and estimates savings to reduce LLM token usage.

- Repository: [rtk-ai/rtk](https://github.com/rtk-ai/rtk)
- Tags: how-to-guide
- Published: 2026-04-24

---

**TLDR:** RTK's `discover` command scans local Claude Code session logs in `~/.claude/projects`, extracts Bash commands using `ClaudeProvider::extract_commands`, classifies them via `registry::classify_command`, estimates potential token savings using output-length heuristics, and generates a report highlighting where RTK filtering could have reduced LLM token usage.

The RTK CLI tool from the **rtk-ai/rtk** repository includes a powerful analytics feature that retrospectively analyzes your Claude Code workflow. By processing historical session data stored locally on your machine, RTK identifies specific commands where token filtering would have saved money and context window space. The entire pipeline is orchestrated through [`src/discover/mod.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/mod.rs) and operates in five distinct stages.

## The Five-Stage Discovery Pipeline

### Stage 1: Session Discovery

The process begins in [`src/discover/provider.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/provider.rs) (lines 71-94) with `ClaudeProvider::discover_sessions`. This function walks the `~/.claude/projects` directory to locate session files.

- Recursively searches for `*.jsonl` files containing Claude Code conversation history
- Optionally filters by **project name** or **age** (sessions from the last *N* days)
- Returns a vector of file paths to be processed

### Stage 2: Command Extraction

For each discovered session file, `ClaudeProvider::extract_commands` in [`src/discover/provider.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/provider.rs) (lines 38-63) parses the JSONL structure to identify executable commands.

- Matches **assistant** `tool_use` blocks containing Bash commands
- Pairs each command with the subsequent **user** `tool_result` block
- Records the command string, output length in bytes, and error status
- Stores results in `ExtractedCommand` structs for further analysis

### Stage 3: Classification and Token Estimation

Each extracted command is fed to `registry::classify_command` in [`src/discover/registry.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/registry.rs) (lines 88-112). This applies the same rule engine RTK uses at runtime.

- **Regex-based classification**: Categorizes commands as **Supported**, **Unsupported**, or **Ignored**
- Special status detection (e.g., `Passthrough` for partial filtering)
- **Token estimation**:
  - If `output_len` is known: uses `len / 4` (approximate tokens per byte)
  - Otherwise: falls back to `category_avg_tokens` from historical averages
- **Savings calculation**: `output_tokens × estimated_savings_pct / 100`

### Stage 4: Bypass Detection

RTK specifically flags commands where filtering was explicitly disabled. The `has_rtk_disabled_prefix` function in [`src/discover/registry.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/registry.rs) (lines 381-398) detects commands starting with the environment flag `RTK_DISABLED=1`.

These commands are counted as *missed opportunities* because the user disabled RTK on a command that the tool could have filtered. This same detection logic is reused in [`src/analytics/gain.rs`](https://github.com/rtk-ai/rtk/blob/main/src/analytics/gain.rs) (lines 629-652) for consistency across analytics pathways.

### Stage 5: Report Generation

Finally, `discover::report::format_text` and `format_json` in [`src/discover/report.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/report.rs) (lines 74-118) aggregate the processed data into human-readable summaries or machine-readable JSON.

- Calculates totals: commands analyzed, already-wrapped commands, supported-but-unfiltered commands, and RTK_DISABLED bypasses
- Groups results by command category
- Displays estimated token savings per group

## Using the RTK Discover Command

To analyze your recent Claude Code usage, run the `discover` command with appropriate filters:

```bash

# Scan the current project for the last 30 days, showing top 10 missed savings

rtk discover --limit 10

```

For automation and dashboards, export to JSON:

```bash

# Machine-readable report for the last 7 days (CI dashboard compatible)

rtk discover --format json --since 7 > rtk-discover-report.json

```

To audit where you explicitly disabled RTK:

```bash

# Show only commands run with RTK_DISABLED=1

rtk discover --all --format text | grep RTK_DISABLED

```

## Programmatic Integration

You can also invoke the discovery pipeline programmatically using the RTK Rust library. The `run` function handles the complete workflow:

```rust
use rtk::discover::run;

fn main() -> anyhow::Result<()> {
    // Find missed savings for all projects, last 14 days, text format
    run(
        None,          // no specific project filter
        true,          // scan all projects in ~/.claude/projects
        14,            // days of history to analyze
        20,            // limit results per section
        "text",        // output format: "text" or "json"
        0,             // verbose level
    )
}

```

## Summary

- **Session scanning**: `ClaudeProvider::discover_sessions` locates JSONL files in `~/.claude/projects` with optional date filtering
- **Command extraction**: Parses `tool_use` and `tool_result` blocks to capture Bash commands and their output sizes
- **Token estimation**: Uses `len / 4` heuristics or category averages in `registry::classify_command` to estimate potential savings
- **Bypass detection**: `has_rtk_disabled_prefix` flags `RTK_DISABLED=1` commands as explicit missed opportunities
- **Reporting**: `format_text` and `format_json` in [`src/discover/report.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/report.rs) convert aggregated data into actionable reports

## Frequently Asked Questions

### How accurate is RTK's token savings estimation?

RTK uses a conservative approximation of **4 bytes per token** (dividing output length by 4) when actual token counts are unavailable. According to the source code in [`src/discover/registry.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/registry.rs), this heuristic provides a reliable lower-bound estimate for budgeting purposes. When historical output lengths are missing, RTK falls back to `category_avg_tokens` based on command type.

### What does RTK consider a "bypass"?

Any command prefixed with the environment variable `RTK_DISABLED=1` is detected by the `has_rtk_disabled_prefix` function in [`src/discover/registry.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/registry.rs) (lines 381-398). These represent intentional disabling of RTK on commands that would otherwise have been supported and filtered, making them explicit missed savings opportunities rather than unsupported commands.

### Can I export discovery data for CI dashboards?

Yes. The `--format json` flag outputs machine-readable JSON through `format_json` in [`src/discover/report.rs`](https://github.com/rtk-ai/rtk/blob/main/src/discover/report.rs). This format includes structured data on estimated tokens, command categories, and savings percentages suitable for ingestion by analytics platforms, CI pipeline reports, or custom dashboards.

### Does RTK discover sessions across all projects?

By default, RTK filters sessions by the current project context, but the `--all` flag instructs `ClaudeProvider::discover_sessions` to scan the entire `~/.claude/projects` directory. You can also use `--project <name>` to target a specific project's history without changing directories.