# LiteLLM Router Strategies Comparison: Cost, Latency, Load, and Rate Limits Explained

> Compare LiteLLM router strategies lowest_cost, lowest_latency, least_busy, and lowest_tpm_rpm to optimize LLM deployments for cost, speed, load, and rate limits.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: performance
- Published: 2026-03-26

---

**LiteLLM provides four distinct routing strategies—`cost-based-routing`, `latency-based-routing`, `least-busy`, and `usage-based-routing`—that optimize deployment selection by minimizing monetary cost, response time, concurrent load, or token-rate limits respectively.**

LiteLLM's router intelligently distributes LLM requests across multiple model deployments using real-time telemetry. Each routing strategy implements a specialized logging handler that records specific metrics in a shared `DualCache`, enabling the router to select the optimal deployment for every request while respecting provider TPM/RPM constraints.

## Architectural Overview

The routing system centers on the `Router` class defined in [`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py). During initialization, the `routing_strategy` parameter determines which logging handler is instantiated (lines 781–846). All strategies inherit from `BaseRoutingStrategy` in [`litellm/router_strategy/base_routing_strategy.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_strategy/base_routing_strategy.py), which provides shared utilities for batched Redis writes, TTL handling, and asynchronous cache synchronization.

When a request arrives, the router invokes the strategy's `async_get_available_deployments()` method, which receives the `model_group`, `healthy_deployments`, and message content, then returns a single deployment dictionary based on the strategy's optimization criteria.

## Cost-Based Routing (lowest_cost)

**`cost-based-routing`** minimizes estimated spend per request by calculating input and output token costs using rates from `litellm.model_cost`.

The `LowestCostLoggingHandler` in [`litellm/router_strategy/lowest_cost.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_strategy/lowest_cost.py) maintains a per-minute map (`<model_group>_map`) tracking token counts and request rates. For each healthy deployment, the handler computes `item_cost = input_cost + output_cost`, filters out deployments that would exceed TPM/RPM limits, and returns the deployment with the lowest calculated cost.

This strategy is ideal for budget-conscious applications using multiple OpenAI models with varying per-token pricing tiers.

## Latency-Based Routing (lowest_latency)

**`latency-based-routing`** optimizes for fastest response time by tracking time-to-first-token (TTFT) for streaming calls or total response duration for standard requests.

Implemented in [`litellm/router_strategy/lowest_latency.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_strategy/lowest_latency.py), the `LowestLatencyLoggingHandler` stores recent latency values in a bounded list (`max_latency_list_size`) within the cache entry `latency_key = f"{model_group}_map"`. During routing, it computes the average latency for each deployment, discards any exceeding rate limits, and selects the deployment with the lowest historical average.

Use this strategy for real-time chat interfaces or streaming applications where user-perceived speed is critical.

## Least-Busy Routing

**`least-busy`** distributes traffic evenly by tracking the number of in-flight requests per deployment.

The `LeastBusyLoggingHandler` in [`litellm/router_strategy/least_busy.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_strategy/least_busy.py) increments a counter (`<model_group>_request_count`) in `log_pre_api_call()` before sending a request and decrements it in `log_success_event()` or `log_failure_event()` upon completion. The router selects the deployment with the smallest active request count, breaking ties randomly.

This approach prevents individual deployments from becoming overwhelmed during high-concurrency batch processing workloads.

## Usage-Based Routing (lowest_tpm_rpm)

**`usage-based-routing`** (and its v2 variant) enforces strict compliance with provider token-per-minute (TPM) and request-per-minute (RPM) limits.

The handlers in [`litellm/router_strategy/lowest_tpm_rpm.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_strategy/lowest_tpm_rpm.py) and [`lowest_tpm_rpm_v2.py`](https://github.com/BerriAI/litellm/blob/main/lowest_tpm_rpm_v2.py) maintain separate cache entries for per-minute TPM (`<model_group>:tpm:<minute>`) and RPM (`<model_group>:rpm:<minute>`). When routing, the handler calculates projected usage (current count + new request tokens) and returns the first deployment that remains below its configured limits, or the one with the lowest projected TPM when multiple options exist.

Version 2 includes refactored TTL handling for improved cache expiration management. Use this strategy when working with strict Azure OpenAI quotas or other providers imposing hard rate limits.

## Implementation Examples

### Basic Cost Optimization

```python
import litellm

model_list = [
    {
        "model": "gpt-4",
        "api_key": "sk-xxxx",
        "metadata": {"model_group": "openai"},
    },
    {
        "model": "gpt-4-32k", 
        "api_key": "sk-yyyy",
        "metadata": {"model_group": "openai"},
    },
]

response = litellm.completion(
    model="gpt-4",
    messages=[{"role": "user", "content": "Explain machine learning"}],
    model_list=model_list,
    routing_strategy="cost-based-routing",
)

```

The `LowestCostLoggingHandler.async_get_available_deployments()` automatically selects the cheaper deployment when both satisfy TPM/RPM constraints.

### Latency-Aware Streaming

```python
response = litellm.completion(
    model="gpt-4",
    messages=[{"role": "user", "content": "Explain quantum computing"}],
    model_list=model_list,
    routing_strategy="latency-based-routing",
    routing_strategy_args={"lowest_latency_buffer": 0.2},
    stream=True,
)

for chunk in response:
    print(chunk.choices[0].delta.content, end="")

```

For streaming calls, the handler records TTFT in `latency_key` and averages recent values to predict the fastest deployment.

### Rate Limit Protection

```python
response = litellm.completion(
    model="azure-gpt-3.5",
    messages=[{"role": "user", "content": "Write a haiku"}],
    model_list=model_list,
    routing_strategy="usage-based-routing",
)

```

The handler automatically reads `tpm` and `rpm` limits from the provider's model info and avoids deployments that would exceed per-minute quotas.

### High-Throughput Load Balancing

```python
response = litellm.completion(
    model="gpt-4",
    messages=batch_messages,
    model_list=model_list,
    routing_strategy="least-busy",
    num_retries=2,
)

```

The router tracks in-flight requests via `<model_group>_request_count` and distributes load across the least utilized deployments.

## Summary

- **Cost-based-routing** minimizes spend by calculating per-token costs from `litellm.model_cost` and selecting the cheapest eligible deployment in [`lowest_cost.py`](https://github.com/BerriAI/litellm/blob/main/lowest_cost.py).
- **Latency-based-routing** reduces response time by averaging historical TTFT values stored in a bounded list within [`lowest_latency.py`](https://github.com/BerriAI/litellm/blob/main/lowest_latency.py).
- **Least-busy** prevents overload by tracking active request counts incremented in `log_pre_api_call()` and selecting deployments with minimal concurrency in [`least_busy.py`](https://github.com/BerriAI/litellm/blob/main/least_busy.py).
- **Usage-based-routing** enforces provider limits by monitoring per-minute TPM/RPM counters in [`lowest_tpm_rpm.py`](https://github.com/BerriAI/litellm/blob/main/lowest_tpm_rpm.py) (v1) or [`lowest_tpm_rpm_v2.py`](https://github.com/BerriAI/litellm/blob/main/lowest_tpm_rpm_v2.py) (v2), returning only deployments with sufficient remaining quota.

## Frequently Asked Questions

### How do I configure multiple routing strategies simultaneously?

LiteLLM selects one primary strategy via the `routing_strategy` parameter, but you can combine strategies by implementing custom logic in your application layer or using tag-based routing to filter deployments before applying metric-based selection. The router's fallback mechanism will retry with the next-best deployment according to the chosen strategy's ranking.

### What happens if all deployments exceed TPM/RPM limits?

When using `usage-based-routing`, if no deployment has sufficient remaining quota for the projected request size, the router raises a `RateLimitError`. For other strategies, the router filters out over-limit deployments; if none remain eligible, it either retries with exponential backoff or raises an exception depending on your `num_retries` configuration.

### Does latency-based routing work for non-streaming requests?

Yes. The `LowestLatencyLoggingHandler` in [`lowest_latency.py`](https://github.com/BerriAI/litellm/blob/main/lowest_latency.py) records total response duration for standard requests and time-to-first-token (TTFT) for streaming calls. Both metrics are stored in the same bounded list and averaged to determine the fastest deployment regardless of the `stream` parameter.

### Where does LiteLLM store routing metrics?

All strategies use the `DualCache` instance (in-memory with optional Redis backing) defined in [`base_routing_strategy.py`](https://github.com/BerriAI/litellm/blob/main/base_routing_strategy.py). Cost data resides in `<model_group>_map`, latency in `latency_key`, busy counts in `<model_group>_request_count`, and rate limits in `<model_group>:tpm:<minute>` and `<model_group>:rpm:<minute>` entries.