LiteLLM Token Usage and Cost Tracking Across Providers: A Complete Technical Guide

LiteLLM normalizes token counts and calculates monetary costs for any supported LLM provider by combining provider-specific token counters with a static pricing map and extensible custom pricing overrides.

LiteLLM (BerriAI/litellm) provides a unified interface for dozens of LLM providers, and its token usage and cost tracking across providers works seamlessly whether you are calling OpenAI, Anthropic, Vertex AI, or any other supported endpoint. The system aggregates prompt and completion tokens, handles multimodal inputs, and computes real-time costs using a comprehensive pricing dictionary that you can override at runtime.

How LiteLLM Counts Tokens Across Providers

When you invoke a completion, LiteLLM dispatches to a provider-specific token counter that understands the target model's tokenization scheme.

The entry point is the token_counter function in litellm/litellm_core_utils/token_counter.py (line 349). This function inspects the model name and routes to specialized implementations located in provider subdirectories such as litellm/llms/openai/count_tokens.py or litellm/llms/anthropic/count_tokens.py.

These counters return structured counts including:

  • Prompt tokens and completion tokens
  • Multimodal tokens (images, audio, video durations)
  • Reasoning tokens (for models like Claude 3.7 Sonnet or Gemini that separate thinking steps from output)

The Usage Object and Cost Attribution

After counting, LiteLLM wraps raw counts in the Usage class defined in litellm/types/utils.py (line 1500). This model extends OpenAI's CompletionUsage with additional fields required for modern LLM billing:

  • reasoning_tokens and text_tokens (auto-calculated totals)
  • cache_creation_input_tokens and cache_read_input_tokens (for Anthropic and DeepSeek prompt caching)
  • image_tokens, video_tokens, character_count (multimodal billing units)
  • server_tool_use and web_search_requests (for Anthropic tool-use pricing)

The Usage object also carries a cost attribute when cost calculation is enabled.

Cost Calculation Mechanism

LiteLLM computes cost through a pipeline that merges static pricing data with optional runtime overrides:

  1. Price Map Lookup – The library loads model_prices_and_context_window.json, which defines per-token, per-character, per-second, and per-modal-unit rates (e.g., input_cost_per_token, output_cost_per_video_per_second).

  2. Custom Pricing Application – Callers can supply a custom_pricing dictionary that supersedes the built-in map. The helper _cost_per_token_custom_pricing_helper in litellm/cost_calculator.py (line 71) merges these values.

  3. Response Cost Calculation – After the LLM returns, the _response_cost_calculator method in litellm/litellm_core_utils/litellm_logging.py (line 1429) extracts the model name, token counts, and caching metadata, then calls the public litellm.response_cost_calculator function to produce a float response_cost.

The final ModelResponse object contains the usage attribute and a hidden _hidden_params["response_cost"] field for downstream logging.

Provider-Specific Handling for Advanced Features

LiteLLM's cost tracking handles provider-specific billing nuances automatically:

  • Reasoning tokens – When only reasoning_tokens are supplied (Anthropic, Gemini), the Usage constructor auto-calculates text_tokens to ensure accurate total counts.

  • Prompt caching – For providers supporting cache creation and cache read (Anthropic, DeepSeek), the Usage model stores cache_creation_input_tokens and cache_read_input_tokens. The cost calculator applies specific cache pricing rates when present in the JSON map.

  • Tool use and web search – Fields like server_tool_use and web_search_requests capture Anthropic-specific tool call billing.

  • Multimodal inputs – The cost calculator pulls matching per-unit prices for images (output_cost_per_image), video duration (output_cost_per_video_per_second), and audio from the static map based on counts populated by the provider-specific token counter.

Practical Implementation Examples

Basic Usage and Cost Retrieval

import litellm

# Standard completion with automatic usage tracking

response = litellm.completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Summarize the plot of Inception"}],
)

# Inspect normalized token counts

print(f"Prompt tokens: {response.usage.prompt_tokens}")
print(f"Completion tokens: {response.usage.completion_tokens}")
print(f"Total tokens: {response.usage.total_tokens}")

# Calculate dollar cost

cost = litellm.response_cost_calculator(
    response_object=response,
    model="gpt-4o-mini",
)
print(f"Request cost: ${cost:.6f}")

Custom Pricing Overrides


# Override default pricing for internal model routing or budget forecasting

custom_rates = {
    "input_cost_per_token": 0.0000002,
    "output_cost_per_token": 0.0000004,
}

custom_cost = litellm.response_cost_calculator(
    response_object=response,
    model="gpt-4o-mini",
    custom_pricing=custom_rates,
)
print(f"Custom pricing cost: ${custom_cost:.6f}")

Streaming Cost Calculation


# Cost is available once the final chunk arrives with usage metadata

stream = litellm.completion(
    model="claude-3-5-sonnet-20240620",
    messages=[{"role": "user", "content": "Explain quantum entanglement"}],
    stream=True,
)

final_chunk = None
for chunk in stream:
    final_chunk = chunk

# The last chunk contains usage data for cost calculation

if final_chunk and getattr(final_chunk, "usage", None):
    streaming_cost = litellm.response_cost_calculator(
        response_object=final_chunk,
        model="claude-3-5-sonnet-20240620",
    )
    print(f"Streaming request cost: ${streaming_cost:.6f}")

Key Source Files and Architecture

File Purpose
litellm/litellm_core_utils/token_counter.py Generic token counter entry point that dispatches to provider-specific implementations
litellm/llms/{provider}/count_tokens.py Provider-specific tokenization logic (OpenAI, Anthropic, Vertex AI, etc.)
litellm/types/utils.py Usage class definition extending OpenAI's schema with cost and multimodal fields
litellm/cost_calculator.py Public response_cost_calculator and _cost_per_token_custom_pricing_helper
litellm/litellm_core_utils/litellm_logging.py _response_cost_calculator method integrating usage data with pricing
model_prices_and_context_window.json Static pricing dictionary with per-token and per-modal rates
tests/test_litellm/test_cost_calculator.py Unit tests verifying usage-to-cost conversion accuracy

Summary

  • LiteLLM token usage and cost tracking across providers relies on provider-specific token counters that feed a normalized Usage object.
  • The Usage class in litellm/types/utils.py extends OpenAI's standard to include reasoning tokens, caching metrics, and multimodal counts.
  • Cost calculation combines the static model_prices_and_context_window.json map with optional custom_pricing overrides via litellm/cost_calculator.py.
  • The _response_cost_calculator in litellm_logging.py bridges usage data and pricing to produce the final response_cost float.
  • This architecture works identically across OpenAI, Anthropic, Vertex AI, and other providers, handling streaming and batch requests uniformly.

Frequently Asked Questions

How does LiteLLM handle token counting for multimodal inputs like images and video?

LiteLLM's provider-specific token counters populate fields like image_tokens, video_tokens, and character_count in the Usage object. The cost calculator then pulls the corresponding rates from model_prices_and_context_window.json (e.g., output_cost_per_image or output_cost_per_video_per_second) to compute the final cost.

Can I override the default pricing for a model in LiteLLM?

Yes. Pass a custom_pricing dictionary to litellm.response_cost_calculator with keys like input_cost_per_token and output_cost_per_token. The _cost_per_token_custom_pricing_helper function in litellm/cost_calculator.py merges these values, overriding the static JSON map for that specific calculation.

Does LiteLLM calculate costs during streaming responses?

LiteLLM calculates costs once the stream completes and the final chunk contains a usage object. You can then pass that final chunk to litellm.response_cost_calculator to retrieve the exact dollar cost of the streaming request.

How does LiteLLM track prompt caching costs for Anthropic models?

The Usage model captures cache_creation_input_tokens and cache_read_input_tokens when present in the provider response. The cost calculator checks model_prices_and_context_window.json for cache-specific pricing entries (e.g., per-1k cache tokens) and factors these into the total response_cost returned in _hidden_params.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →