How Token Counting and Cost Calculation Work in Roo Code

Roo Code standardizes API cost estimation across providers using a three-layer architecture that combines a tiktoken-based encoder with a 1.5x fudge factor, provider-agnostic model metadata, and separate calculation paths for Anthropic and OpenAI pricing models.

Roo Code (RooCodeInc/Roo-Code) abstracts token accounting and price calculation so the same UI can work seamlessly with Anthropic-compatible and OpenAI-compatible backends. The system normalizes every request payload into token counts using a shared encoder, then applies provider-specific pricing logic to generate accurate dollar estimates. This technical deep dive examines the implementation from the low-level tokenizer to the final cost computation.

The Three-Layer Architecture

The cost estimation pipeline splits into three distinct layers, each handling a specific responsibility while remaining provider-agnostic where possible.

Tokenization Layer

The tokenization layer converts request and response payloads—including text, images, and tool calls—into raw token counts. The core implementation lives in src/utils/tiktoken.ts, which uses the tiktoken library with the o200k_base encoder and applies a fudge factor (TOKEN_FUDGE_FACTOR = 1.5) to align estimates with real-world provider counts.

Images use a heuristic based on the square root of the base64 string length, while tool calls are serialized to plain text using serializeToolUse and serializeToolResult before tokenization. The public utility countTokens() in src/utils/countTokens.ts optionally runs this encoder in a Web Worker (src/workers/countTokens.ts) to keep the UI responsive during heavy computations.

Model Pricing Metadata

Per-model price rates and optional long-context multipliers are defined in src/api/transform/cache-strategy/types.ts within the ModelInfo interface. This metadata includes:

  • inputPrice and outputPrice (USD per 1M tokens)
  • cacheWritesPrice and cacheReadsPrice (for OpenAI-compatible caching)
  • longContextPricing object with thresholdTokens and multipliers (OpenAI-specific)

Anthropic models expose separate input and output prices without cache-specific fields, while OpenAI models include cache-write and cache-read pricing alongside potential long-context tiers.

Cost Calculation Layer

The final layer computes dollar figures from token counts and metadata. Located in src/shared/cost.ts, this layer exports two public helpers: calculateApiCostAnthropic() and calculateApiCostOpenAI(). Both internally call calculateApiCostInternal(), but handle token accounting differently based on provider conventions.

Token Counting Implementation Details

When you send a request containing mixed content types, Roo Code processes each component before feeding it to the encoder.

import { countTokens } from "@roo/utils/countTokens"
import type { Anthropic } from "@anthropic-ai/sdk"

const payload: Anthropic.Messages.ContentBlockParam[] = [
  { type: "text", text: "Explain quantum computing." },
  { type: "image", source: { data: "base64-..." } },
  { type: "tool_use", name: "search", input: { query: "qubits" } }
]

const tokenCount = await countTokens(payload)
// Returns: ~150 tokens (including image heuristic and tool serialization)

The countTokens() function delegates to the worker thread, which iterates through content blocks, encodes text, estimates image tokens using the square root heuristic, and serializes tool calls to JSON strings before passing everything to the cached Tiktoken encoder. The resulting count is multiplied by TOKEN_FUDGE_FACTOR (1.5) to account for discrepancies between client-side estimation and actual provider tokenization.

Provider-Specific Cost Calculation

The critical difference between providers lies in how cached tokens factor into the total cost calculation.

Anthropic-Compatible Calculation

For Anthropic models, input tokens do not include cached tokens in the reported count. The calculateApiCostAnthropic() function in src/shared/cost.ts explicitly adds cache creation and cache read tokens to the input total before applying rates:

const anthCost = calculateApiCostAnthropic(
  modelInfo,
  inputTokens,      // Regular prompt tokens only
  outputTokens,
  cacheWriteTokens, // Cached prompt writes
  cacheReadTokens   // Cached prompt reads
)

The function computes totalInputTokens = inputTokens + cacheWriteTokens + cacheReadTokens, then applies the formula:

Total Cost = (inputRate × inputTokens) + (outputRate × outputTokens) + (cacheWriteRate × cacheWriteTokens) + (cacheReadRate × cacheReadTokens)

OpenAI-Compatible Calculation

OpenAI-compatible APIs report input tokens that already include cached tokens. The calculateApiCostOpenAI() function must therefore extract the non-cached portion for accurate pricing:

const openCost = calculateApiCostOpenAI(
  modelInfo,
  inputTokens,      // Already includes cache reads/writes
  outputTokens,
  cacheWriteTokens,
  cacheReadTokens
)

Internally, it calculates nonCachedInputTokens = inputTokens - cacheWriteTokens - cacheReadTokens to ensure you don't double-pay for cached content, while still charging the specific cache-write and cache-read rates for those operations.

Long-Context Pricing Adjustments

OpenAI models occasionally offer tiered pricing where context windows exceeding a threshold trigger different rates. When a model defines longContextPricing with a thresholdTokens value, calculateApiCostOpenAI() automatically invokes applyLongContextPricing() to adjust the input and output rates:

// Inside src/shared/cost.ts
if (inputTokens > modelInfo.longContextPricing.thresholdTokens) {
  effectiveInputRate *= modelInfo.longContextPricing.inputPriceMultiplier
  effectiveOutputRate *= modelInfo.longContextPricing.outputPriceMultiplier
}

Anthropic models do not expose this tiered pricing structure, so the calculation uses flat rates throughout.

Price Parsing and Model Configuration

Price values originate from the model catalog as strings (e.g., "0.01" for $0.01 per 1K tokens) and are parsed using the parseApiPrice() helper in src/shared/cost.ts:

export const parseApiPrice = (price: string | number | undefined): number => {
  return price ? parseFloat(String(price)) * 1_000_000 : 0
}

This normalization converts per-1K or per-1M token rates into consistent multipliers used by the calculation functions.

Summary

  • Roo Code uses a unified tokenization pipeline in src/utils/tiktoken.ts with the o200k_base encoder and a 1.5x fudge factor to estimate token counts across all providers.
  • The countTokens() utility handles text, images (via square-root-of-base64-length heuristic), and tool calls (via JSON serialization) in a background Web Worker at src/workers/countTokens.ts.
  • Anthropic cost calculation (calculateApiCostAnthropic) treats cached tokens as additive to the input total, while OpenAI calculation (calculateApiCostOpenAI) assumes cached tokens are already included in the input count.
  • OpenAI models support long-context pricing tiers via applyLongContextPricing(), automatically applying multipliers when token counts exceed defined thresholds.
  • All pricing metadata lives in the ModelInfo interface defined in src/api/transform/cache-strategy/types.ts, with actual computation centralized in src/shared/cost.ts.

Frequently Asked Questions

Why does Roo Code multiply token counts by a 1.5 fudge factor?

The TOKEN_FUDGE_FACTOR = 1.5 constant in src/utils/tiktoken.ts compensates for discrepancies between client-side tiktoken estimation and the actual tokenization performed by API providers. Because Roo Code cannot access provider-specific tokenizers server-side, this multiplier provides a conservative upper bound that ensures cost estimates match or slightly exceed actual billing, preventing budget surprises.

How does Roo Code calculate tokens for images and tool calls?

Images are estimated using the square root of the base64 string length heuristic, while tool calls are serialized to JSON strings using serializeToolUse and serializeToolResult functions before being tokenized as standard text. This approach allows the o200k_base encoder to process non-text content without provider-specific vision tokenization APIs.

What is the difference between Anthropic and OpenAI cache handling?

Anthropic-compatible backends report input tokens excluding cache hits, requiring Roo Code to manually add cacheWriteTokens and cacheReadTokens to the input total. OpenAI-compatible backends include cached tokens in the reported inputTokens count, requiring the calculator to extract nonCachedInputTokens = inputTokens - cacheWriteTokens - cacheReadTokens to avoid double-billing while still charging specific cache read/write rates.

When does long-context pricing apply to OpenAI models?

Long-context pricing activates when the total input tokens exceed the thresholdTokens value defined in the model's longContextPricing configuration. The applyLongContextPricing() function in src/shared/cost.ts automatically multiplies the input and output rates by the specified multipliers (typically 1.2x or higher) when this threshold is crossed, reflecting the provider's tiered pricing structure for large context windows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →