# LiteLLM Token Usage and Cost Tracking Across Providers: A Complete Technical Guide

> Master LiteLLM token usage and cost tracking across providers. Our guide shows how LiteLLM normalizes counts and calculates costs for seamless LLM management.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: deep-dive
- Published: 2026-03-26

---

**LiteLLM normalizes token counts and calculates monetary costs for any supported LLM provider by combining provider-specific token counters with a static pricing map and extensible custom pricing overrides.**

LiteLLM (BerriAI/litellm) provides a unified interface for dozens of LLM providers, and its token usage and cost tracking across providers works seamlessly whether you are calling OpenAI, Anthropic, Vertex AI, or any other supported endpoint. The system aggregates prompt and completion tokens, handles multimodal inputs, and computes real-time costs using a comprehensive pricing dictionary that you can override at runtime.

## How LiteLLM Counts Tokens Across Providers

When you invoke a completion, LiteLLM dispatches to a **provider-specific token counter** that understands the target model's tokenization scheme.

The entry point is the `token_counter` function in [`litellm/litellm_core_utils/token_counter.py`](https://github.com/BerriAI/litellm/blob/main/litellm/litellm_core_utils/token_counter.py) (line 349). This function inspects the model name and routes to specialized implementations located in provider subdirectories such as [`litellm/llms/openai/count_tokens.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/openai/count_tokens.py) or [`litellm/llms/anthropic/count_tokens.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/anthropic/count_tokens.py).

These counters return structured counts including:
- **Prompt tokens** and **completion tokens**
- **Multimodal tokens** (images, audio, video durations)
- **Reasoning tokens** (for models like Claude 3.7 Sonnet or Gemini that separate thinking steps from output)

## The Usage Object and Cost Attribution

After counting, LiteLLM wraps raw counts in the **`Usage`** class defined in [`litellm/types/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/utils.py) (line 1500). This model extends OpenAI's `CompletionUsage` with additional fields required for modern LLM billing:

- `reasoning_tokens` and `text_tokens` (auto-calculated totals)
- `cache_creation_input_tokens` and `cache_read_input_tokens` (for Anthropic and DeepSeek prompt caching)
- `image_tokens`, `video_tokens`, `character_count` (multimodal billing units)
- `server_tool_use` and `web_search_requests` (for Anthropic tool-use pricing)

The `Usage` object also carries a `cost` attribute when cost calculation is enabled.

## Cost Calculation Mechanism

LiteLLM computes cost through a pipeline that merges static pricing data with optional runtime overrides:

1. **Price Map Lookup** – The library loads [`model_prices_and_context_window.json`](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json), which defines per-token, per-character, per-second, and per-modal-unit rates (e.g., `input_cost_per_token`, `output_cost_per_video_per_second`).

2. **Custom Pricing Application** – Callers can supply a `custom_pricing` dictionary that supersedes the built-in map. The helper `_cost_per_token_custom_pricing_helper` in [`litellm/cost_calculator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/cost_calculator.py) (line 71) merges these values.

3. **Response Cost Calculation** – After the LLM returns, the `_response_cost_calculator` method in [`litellm/litellm_core_utils/litellm_logging.py`](https://github.com/BerriAI/litellm/blob/main/litellm/litellm_core_utils/litellm_logging.py) (line 1429) extracts the model name, token counts, and caching metadata, then calls the public `litellm.response_cost_calculator` function to produce a float `response_cost`.

The final `ModelResponse` object contains the `usage` attribute and a hidden `_hidden_params["response_cost"]` field for downstream logging.

## Provider-Specific Handling for Advanced Features

LiteLLM's cost tracking handles provider-specific billing nuances automatically:

- **Reasoning tokens** – When only `reasoning_tokens` are supplied (Anthropic, Gemini), the `Usage` constructor auto-calculates `text_tokens` to ensure accurate total counts.

- **Prompt caching** – For providers supporting cache creation and cache read (Anthropic, DeepSeek), the `Usage` model stores `cache_creation_input_tokens` and `cache_read_input_tokens`. The cost calculator applies specific cache pricing rates when present in the JSON map.

- **Tool use and web search** – Fields like `server_tool_use` and `web_search_requests` capture Anthropic-specific tool call billing.

- **Multimodal inputs** – The cost calculator pulls matching per-unit prices for images (`output_cost_per_image`), video duration (`output_cost_per_video_per_second`), and audio from the static map based on counts populated by the provider-specific token counter.

## Practical Implementation Examples

### Basic Usage and Cost Retrieval

```python
import litellm

# Standard completion with automatic usage tracking

response = litellm.completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Summarize the plot of Inception"}],
)

# Inspect normalized token counts

print(f"Prompt tokens: {response.usage.prompt_tokens}")
print(f"Completion tokens: {response.usage.completion_tokens}")
print(f"Total tokens: {response.usage.total_tokens}")

# Calculate dollar cost

cost = litellm.response_cost_calculator(
    response_object=response,
    model="gpt-4o-mini",
)
print(f"Request cost: ${cost:.6f}")

```

### Custom Pricing Overrides

```python

# Override default pricing for internal model routing or budget forecasting

custom_rates = {
    "input_cost_per_token": 0.0000002,
    "output_cost_per_token": 0.0000004,
}

custom_cost = litellm.response_cost_calculator(
    response_object=response,
    model="gpt-4o-mini",
    custom_pricing=custom_rates,
)
print(f"Custom pricing cost: ${custom_cost:.6f}")

```

### Streaming Cost Calculation

```python

# Cost is available once the final chunk arrives with usage metadata

stream = litellm.completion(
    model="claude-3-5-sonnet-20240620",
    messages=[{"role": "user", "content": "Explain quantum entanglement"}],
    stream=True,
)

final_chunk = None
for chunk in stream:
    final_chunk = chunk

# The last chunk contains usage data for cost calculation

if final_chunk and getattr(final_chunk, "usage", None):
    streaming_cost = litellm.response_cost_calculator(
        response_object=final_chunk,
        model="claude-3-5-sonnet-20240620",
    )
    print(f"Streaming request cost: ${streaming_cost:.6f}")

```

## Key Source Files and Architecture

| File | Purpose |
|------|---------|
| [`litellm/litellm_core_utils/token_counter.py`](https://github.com/BerriAI/litellm/blob/main/litellm/litellm_core_utils/token_counter.py) | Generic token counter entry point that dispatches to provider-specific implementations |
| `litellm/llms/{provider}/count_tokens.py` | Provider-specific tokenization logic (OpenAI, Anthropic, Vertex AI, etc.) |
| [`litellm/types/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/utils.py) | `Usage` class definition extending OpenAI's schema with cost and multimodal fields |
| [`litellm/cost_calculator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/cost_calculator.py) | Public `response_cost_calculator` and `_cost_per_token_custom_pricing_helper` |
| [`litellm/litellm_core_utils/litellm_logging.py`](https://github.com/BerriAI/litellm/blob/main/litellm/litellm_core_utils/litellm_logging.py) | `_response_cost_calculator` method integrating usage data with pricing |
| [`model_prices_and_context_window.json`](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json) | Static pricing dictionary with per-token and per-modal rates |
| [`tests/test_litellm/test_cost_calculator.py`](https://github.com/BerriAI/litellm/blob/main/tests/test_litellm/test_cost_calculator.py) | Unit tests verifying usage-to-cost conversion accuracy |

## Summary

- **LiteLLM token usage and cost tracking across providers** relies on provider-specific token counters that feed a normalized `Usage` object.
- The `Usage` class in [`litellm/types/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/utils.py) extends OpenAI's standard to include reasoning tokens, caching metrics, and multimodal counts.
- Cost calculation combines the static [`model_prices_and_context_window.json`](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json) map with optional `custom_pricing` overrides via [`litellm/cost_calculator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/cost_calculator.py).
- The `_response_cost_calculator` in [`litellm_logging.py`](https://github.com/BerriAI/litellm/blob/main/litellm_logging.py) bridges usage data and pricing to produce the final `response_cost` float.
- This architecture works identically across OpenAI, Anthropic, Vertex AI, and other providers, handling streaming and batch requests uniformly.

## Frequently Asked Questions

### How does LiteLLM handle token counting for multimodal inputs like images and video?

LiteLLM's provider-specific token counters populate fields like `image_tokens`, `video_tokens`, and `character_count` in the `Usage` object. The cost calculator then pulls the corresponding rates from [`model_prices_and_context_window.json`](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json) (e.g., `output_cost_per_image` or `output_cost_per_video_per_second`) to compute the final cost.

### Can I override the default pricing for a model in LiteLLM?

Yes. Pass a `custom_pricing` dictionary to `litellm.response_cost_calculator` with keys like `input_cost_per_token` and `output_cost_per_token`. The `_cost_per_token_custom_pricing_helper` function in [`litellm/cost_calculator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/cost_calculator.py) merges these values, overriding the static JSON map for that specific calculation.

### Does LiteLLM calculate costs during streaming responses?

LiteLLM calculates costs once the stream completes and the final chunk contains a `usage` object. You can then pass that final chunk to `litellm.response_cost_calculator` to retrieve the exact dollar cost of the streaming request.

### How does LiteLLM track prompt caching costs for Anthropic models?

The `Usage` model captures `cache_creation_input_tokens` and `cache_read_input_tokens` when present in the provider response. The cost calculator checks [`model_prices_and_context_window.json`](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json) for cache-specific pricing entries (e.g., per-1k cache tokens) and factors these into the total `response_cost` returned in `_hidden_params`.