LiteLLM Token Usage and Cost Tracking Across Providers: A Complete Technical Guide
LiteLLM normalizes token counts and calculates monetary costs for any supported LLM provider by combining provider-specific token counters with a static pricing map and extensible custom pricing overrides.
LiteLLM (BerriAI/litellm) provides a unified interface for dozens of LLM providers, and its token usage and cost tracking across providers works seamlessly whether you are calling OpenAI, Anthropic, Vertex AI, or any other supported endpoint. The system aggregates prompt and completion tokens, handles multimodal inputs, and computes real-time costs using a comprehensive pricing dictionary that you can override at runtime.
How LiteLLM Counts Tokens Across Providers
When you invoke a completion, LiteLLM dispatches to a provider-specific token counter that understands the target model's tokenization scheme.
The entry point is the token_counter function in litellm/litellm_core_utils/token_counter.py (line 349). This function inspects the model name and routes to specialized implementations located in provider subdirectories such as litellm/llms/openai/count_tokens.py or litellm/llms/anthropic/count_tokens.py.
These counters return structured counts including:
- Prompt tokens and completion tokens
- Multimodal tokens (images, audio, video durations)
- Reasoning tokens (for models like Claude 3.7 Sonnet or Gemini that separate thinking steps from output)
The Usage Object and Cost Attribution
After counting, LiteLLM wraps raw counts in the Usage class defined in litellm/types/utils.py (line 1500). This model extends OpenAI's CompletionUsage with additional fields required for modern LLM billing:
reasoning_tokensandtext_tokens(auto-calculated totals)cache_creation_input_tokensandcache_read_input_tokens(for Anthropic and DeepSeek prompt caching)image_tokens,video_tokens,character_count(multimodal billing units)server_tool_useandweb_search_requests(for Anthropic tool-use pricing)
The Usage object also carries a cost attribute when cost calculation is enabled.
Cost Calculation Mechanism
LiteLLM computes cost through a pipeline that merges static pricing data with optional runtime overrides:
-
Price Map Lookup – The library loads
model_prices_and_context_window.json, which defines per-token, per-character, per-second, and per-modal-unit rates (e.g.,input_cost_per_token,output_cost_per_video_per_second). -
Custom Pricing Application – Callers can supply a
custom_pricingdictionary that supersedes the built-in map. The helper_cost_per_token_custom_pricing_helperinlitellm/cost_calculator.py(line 71) merges these values. -
Response Cost Calculation – After the LLM returns, the
_response_cost_calculatormethod inlitellm/litellm_core_utils/litellm_logging.py(line 1429) extracts the model name, token counts, and caching metadata, then calls the publiclitellm.response_cost_calculatorfunction to produce a floatresponse_cost.
The final ModelResponse object contains the usage attribute and a hidden _hidden_params["response_cost"] field for downstream logging.
Provider-Specific Handling for Advanced Features
LiteLLM's cost tracking handles provider-specific billing nuances automatically:
-
Reasoning tokens – When only
reasoning_tokensare supplied (Anthropic, Gemini), theUsageconstructor auto-calculatestext_tokensto ensure accurate total counts. -
Prompt caching – For providers supporting cache creation and cache read (Anthropic, DeepSeek), the
Usagemodel storescache_creation_input_tokensandcache_read_input_tokens. The cost calculator applies specific cache pricing rates when present in the JSON map. -
Tool use and web search – Fields like
server_tool_useandweb_search_requestscapture Anthropic-specific tool call billing. -
Multimodal inputs – The cost calculator pulls matching per-unit prices for images (
output_cost_per_image), video duration (output_cost_per_video_per_second), and audio from the static map based on counts populated by the provider-specific token counter.
Practical Implementation Examples
Basic Usage and Cost Retrieval
import litellm
# Standard completion with automatic usage tracking
response = litellm.completion(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Summarize the plot of Inception"}],
)
# Inspect normalized token counts
print(f"Prompt tokens: {response.usage.prompt_tokens}")
print(f"Completion tokens: {response.usage.completion_tokens}")
print(f"Total tokens: {response.usage.total_tokens}")
# Calculate dollar cost
cost = litellm.response_cost_calculator(
response_object=response,
model="gpt-4o-mini",
)
print(f"Request cost: ${cost:.6f}")
Custom Pricing Overrides
# Override default pricing for internal model routing or budget forecasting
custom_rates = {
"input_cost_per_token": 0.0000002,
"output_cost_per_token": 0.0000004,
}
custom_cost = litellm.response_cost_calculator(
response_object=response,
model="gpt-4o-mini",
custom_pricing=custom_rates,
)
print(f"Custom pricing cost: ${custom_cost:.6f}")
Streaming Cost Calculation
# Cost is available once the final chunk arrives with usage metadata
stream = litellm.completion(
model="claude-3-5-sonnet-20240620",
messages=[{"role": "user", "content": "Explain quantum entanglement"}],
stream=True,
)
final_chunk = None
for chunk in stream:
final_chunk = chunk
# The last chunk contains usage data for cost calculation
if final_chunk and getattr(final_chunk, "usage", None):
streaming_cost = litellm.response_cost_calculator(
response_object=final_chunk,
model="claude-3-5-sonnet-20240620",
)
print(f"Streaming request cost: ${streaming_cost:.6f}")
Key Source Files and Architecture
| File | Purpose |
|---|---|
litellm/litellm_core_utils/token_counter.py |
Generic token counter entry point that dispatches to provider-specific implementations |
litellm/llms/{provider}/count_tokens.py |
Provider-specific tokenization logic (OpenAI, Anthropic, Vertex AI, etc.) |
litellm/types/utils.py |
Usage class definition extending OpenAI's schema with cost and multimodal fields |
litellm/cost_calculator.py |
Public response_cost_calculator and _cost_per_token_custom_pricing_helper |
litellm/litellm_core_utils/litellm_logging.py |
_response_cost_calculator method integrating usage data with pricing |
model_prices_and_context_window.json |
Static pricing dictionary with per-token and per-modal rates |
tests/test_litellm/test_cost_calculator.py |
Unit tests verifying usage-to-cost conversion accuracy |
Summary
- LiteLLM token usage and cost tracking across providers relies on provider-specific token counters that feed a normalized
Usageobject. - The
Usageclass inlitellm/types/utils.pyextends OpenAI's standard to include reasoning tokens, caching metrics, and multimodal counts. - Cost calculation combines the static
model_prices_and_context_window.jsonmap with optionalcustom_pricingoverrides vialitellm/cost_calculator.py. - The
_response_cost_calculatorinlitellm_logging.pybridges usage data and pricing to produce the finalresponse_costfloat. - This architecture works identically across OpenAI, Anthropic, Vertex AI, and other providers, handling streaming and batch requests uniformly.
Frequently Asked Questions
How does LiteLLM handle token counting for multimodal inputs like images and video?
LiteLLM's provider-specific token counters populate fields like image_tokens, video_tokens, and character_count in the Usage object. The cost calculator then pulls the corresponding rates from model_prices_and_context_window.json (e.g., output_cost_per_image or output_cost_per_video_per_second) to compute the final cost.
Can I override the default pricing for a model in LiteLLM?
Yes. Pass a custom_pricing dictionary to litellm.response_cost_calculator with keys like input_cost_per_token and output_cost_per_token. The _cost_per_token_custom_pricing_helper function in litellm/cost_calculator.py merges these values, overriding the static JSON map for that specific calculation.
Does LiteLLM calculate costs during streaming responses?
LiteLLM calculates costs once the stream completes and the final chunk contains a usage object. You can then pass that final chunk to litellm.response_cost_calculator to retrieve the exact dollar cost of the streaming request.
How does LiteLLM track prompt caching costs for Anthropic models?
The Usage model captures cache_creation_input_tokens and cache_read_input_tokens when present in the provider response. The cost calculator checks model_prices_and_context_window.json for cache-specific pricing entries (e.g., per-1k cache tokens) and factors these into the total response_cost returned in _hidden_params.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →