# How Token Counting and Cost Calculation Work in Roo Code

> Discover how Roo Code standardizes API token counting and cost calculation across providers using its unique three-layer architecture and provider-agnostic model metadata.

- Repository: [Roo Code/Roo-Code](https://github.com/RooCodeInc/Roo-Code)
- Tags: internals
- Published: 2026-04-26

---

**Roo Code standardizes API cost estimation across providers using a three-layer architecture that combines a tiktoken-based encoder with a 1.5x fudge factor, provider-agnostic model metadata, and separate calculation paths for Anthropic and OpenAI pricing models.**

Roo Code (RooCodeInc/Roo-Code) abstracts **token accounting** and **price calculation** so the same UI can work seamlessly with Anthropic-compatible and OpenAI-compatible backends. The system normalizes every request payload into token counts using a shared encoder, then applies provider-specific pricing logic to generate accurate dollar estimates. This technical deep dive examines the implementation from the low-level tokenizer to the final cost computation.

## The Three-Layer Architecture

The cost estimation pipeline splits into three distinct layers, each handling a specific responsibility while remaining provider-agnostic where possible.

### Tokenization Layer

The tokenization layer converts request and response payloads—including text, images, and tool calls—into raw token counts. The core implementation lives in [`src/utils/tiktoken.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/utils/tiktoken.ts), which uses the **tiktoken** library with the `o200k_base` encoder and applies a **fudge factor** (`TOKEN_FUDGE_FACTOR = 1.5`) to align estimates with real-world provider counts.

Images use a heuristic based on the square root of the base64 string length, while tool calls are serialized to plain text using `serializeToolUse` and `serializeToolResult` before tokenization. The public utility `countTokens()` in [`src/utils/countTokens.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/utils/countTokens.ts) optionally runs this encoder in a Web Worker ([`src/workers/countTokens.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/workers/countTokens.ts)) to keep the UI responsive during heavy computations.

### Model Pricing Metadata

Per-model price rates and optional long-context multipliers are defined in [`src/api/transform/cache-strategy/types.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/api/transform/cache-strategy/types.ts) within the `ModelInfo` interface. This metadata includes:

- `inputPrice` and `outputPrice` (USD per 1M tokens)
- `cacheWritesPrice` and `cacheReadsPrice` (for OpenAI-compatible caching)
- `longContextPricing` object with `thresholdTokens` and multipliers (OpenAI-specific)

Anthropic models expose separate input and output prices without cache-specific fields, while OpenAI models include cache-write and cache-read pricing alongside potential long-context tiers.

### Cost Calculation Layer

The final layer computes dollar figures from token counts and metadata. Located in [`src/shared/cost.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/shared/cost.ts), this layer exports two public helpers: `calculateApiCostAnthropic()` and `calculateApiCostOpenAI()`. Both internally call `calculateApiCostInternal()`, but handle token accounting differently based on provider conventions.

## Token Counting Implementation Details

When you send a request containing mixed content types, Roo Code processes each component before feeding it to the encoder.

```typescript
import { countTokens } from "@roo/utils/countTokens"
import type { Anthropic } from "@anthropic-ai/sdk"

const payload: Anthropic.Messages.ContentBlockParam[] = [
  { type: "text", text: "Explain quantum computing." },
  { type: "image", source: { data: "base64-..." } },
  { type: "tool_use", name: "search", input: { query: "qubits" } }
]

const tokenCount = await countTokens(payload)
// Returns: ~150 tokens (including image heuristic and tool serialization)

```

The `countTokens()` function delegates to the worker thread, which iterates through content blocks, encodes text, estimates image tokens using the square root heuristic, and serializes tool calls to JSON strings before passing everything to the cached `Tiktoken` encoder. The resulting count is multiplied by `TOKEN_FUDGE_FACTOR` (1.5) to account for discrepancies between client-side estimation and actual provider tokenization.

## Provider-Specific Cost Calculation

The critical difference between providers lies in how cached tokens factor into the total cost calculation.

### Anthropic-Compatible Calculation

For Anthropic models, **input tokens do not include cached tokens** in the reported count. The `calculateApiCostAnthropic()` function in [`src/shared/cost.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/shared/cost.ts) explicitly adds cache creation and cache read tokens to the input total before applying rates:

```typescript
const anthCost = calculateApiCostAnthropic(
  modelInfo,
  inputTokens,      // Regular prompt tokens only
  outputTokens,
  cacheWriteTokens, // Cached prompt writes
  cacheReadTokens   // Cached prompt reads
)

```

The function computes `totalInputTokens = inputTokens + cacheWriteTokens + cacheReadTokens`, then applies the formula:

**Total Cost = (inputRate × inputTokens) + (outputRate × outputTokens) + (cacheWriteRate × cacheWriteTokens) + (cacheReadRate × cacheReadTokens)**

### OpenAI-Compatible Calculation

OpenAI-compatible APIs report **input tokens that already include cached tokens**. The `calculateApiCostOpenAI()` function must therefore extract the non-cached portion for accurate pricing:

```typescript
const openCost = calculateApiCostOpenAI(
  modelInfo,
  inputTokens,      // Already includes cache reads/writes
  outputTokens,
  cacheWriteTokens,
  cacheReadTokens
)

```

Internally, it calculates `nonCachedInputTokens = inputTokens - cacheWriteTokens - cacheReadTokens` to ensure you don't double-pay for cached content, while still charging the specific cache-write and cache-read rates for those operations.

### Long-Context Pricing Adjustments

OpenAI models occasionally offer tiered pricing where context windows exceeding a threshold trigger different rates. When a model defines `longContextPricing` with a `thresholdTokens` value, `calculateApiCostOpenAI()` automatically invokes `applyLongContextPricing()` to adjust the input and output rates:

```typescript
// Inside src/shared/cost.ts
if (inputTokens > modelInfo.longContextPricing.thresholdTokens) {
  effectiveInputRate *= modelInfo.longContextPricing.inputPriceMultiplier
  effectiveOutputRate *= modelInfo.longContextPricing.outputPriceMultiplier
}

```

Anthropic models do not expose this tiered pricing structure, so the calculation uses flat rates throughout.

## Price Parsing and Model Configuration

Price values originate from the model catalog as strings (e.g., "0.01" for $0.01 per 1K tokens) and are parsed using the `parseApiPrice()` helper in [`src/shared/cost.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/shared/cost.ts):

```typescript
export const parseApiPrice = (price: string | number | undefined): number => {
  return price ? parseFloat(String(price)) * 1_000_000 : 0
}

```

This normalization converts per-1K or per-1M token rates into consistent multipliers used by the calculation functions.

## Summary

- **Roo Code** uses a unified tokenization pipeline in [`src/utils/tiktoken.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/utils/tiktoken.ts) with the `o200k_base` encoder and a 1.5x fudge factor to estimate token counts across all providers.
- The `countTokens()` utility handles text, images (via square-root-of-base64-length heuristic), and tool calls (via JSON serialization) in a background Web Worker at [`src/workers/countTokens.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/workers/countTokens.ts).
- **Anthropic** cost calculation (`calculateApiCostAnthropic`) treats cached tokens as additive to the input total, while **OpenAI** calculation (`calculateApiCostOpenAI`) assumes cached tokens are already included in the input count.
- OpenAI models support **long-context pricing tiers** via `applyLongContextPricing()`, automatically applying multipliers when token counts exceed defined thresholds.
- All pricing metadata lives in the `ModelInfo` interface defined in [`src/api/transform/cache-strategy/types.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/api/transform/cache-strategy/types.ts), with actual computation centralized in [`src/shared/cost.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/shared/cost.ts).

## Frequently Asked Questions

### Why does Roo Code multiply token counts by a 1.5 fudge factor?

The `TOKEN_FUDGE_FACTOR = 1.5` constant in [`src/utils/tiktoken.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/utils/tiktoken.ts) compensates for discrepancies between client-side tiktoken estimation and the actual tokenization performed by API providers. Because Roo Code cannot access provider-specific tokenizers server-side, this multiplier provides a conservative upper bound that ensures cost estimates match or slightly exceed actual billing, preventing budget surprises.

### How does Roo Code calculate tokens for images and tool calls?

Images are estimated using the square root of the base64 string length heuristic, while tool calls are serialized to JSON strings using `serializeToolUse` and `serializeToolResult` functions before being tokenized as standard text. This approach allows the `o200k_base` encoder to process non-text content without provider-specific vision tokenization APIs.

### What is the difference between Anthropic and OpenAI cache handling?

**Anthropic-compatible** backends report input tokens excluding cache hits, requiring Roo Code to manually add `cacheWriteTokens` and `cacheReadTokens` to the input total. **OpenAI-compatible** backends include cached tokens in the reported `inputTokens` count, requiring the calculator to extract `nonCachedInputTokens = inputTokens - cacheWriteTokens - cacheReadTokens` to avoid double-billing while still charging specific cache read/write rates.

### When does long-context pricing apply to OpenAI models?

Long-context pricing activates when the total input tokens exceed the `thresholdTokens` value defined in the model's `longContextPricing` configuration. The `applyLongContextPricing()` function in [`src/shared/cost.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/shared/cost.ts) automatically multiplies the input and output rates by the specified multipliers (typically 1.2x or higher) when this threshold is crossed, reflecting the provider's tiered pricing structure for large context windows.