LiteLLM Router Strategies Comparison: Cost, Latency, Load, and Rate Limits Explained
LiteLLM provides four distinct routing strategies—cost-based-routing, latency-based-routing, least-busy, and usage-based-routing—that optimize deployment selection by minimizing monetary cost, response time, concurrent load, or token-rate limits respectively.
LiteLLM's router intelligently distributes LLM requests across multiple model deployments using real-time telemetry. Each routing strategy implements a specialized logging handler that records specific metrics in a shared DualCache, enabling the router to select the optimal deployment for every request while respecting provider TPM/RPM constraints.
Architectural Overview
The routing system centers on the Router class defined in litellm/router.py. During initialization, the routing_strategy parameter determines which logging handler is instantiated (lines 781–846). All strategies inherit from BaseRoutingStrategy in litellm/router_strategy/base_routing_strategy.py, which provides shared utilities for batched Redis writes, TTL handling, and asynchronous cache synchronization.
When a request arrives, the router invokes the strategy's async_get_available_deployments() method, which receives the model_group, healthy_deployments, and message content, then returns a single deployment dictionary based on the strategy's optimization criteria.
Cost-Based Routing (lowest_cost)
cost-based-routing minimizes estimated spend per request by calculating input and output token costs using rates from litellm.model_cost.
The LowestCostLoggingHandler in litellm/router_strategy/lowest_cost.py maintains a per-minute map (<model_group>_map) tracking token counts and request rates. For each healthy deployment, the handler computes item_cost = input_cost + output_cost, filters out deployments that would exceed TPM/RPM limits, and returns the deployment with the lowest calculated cost.
This strategy is ideal for budget-conscious applications using multiple OpenAI models with varying per-token pricing tiers.
Latency-Based Routing (lowest_latency)
latency-based-routing optimizes for fastest response time by tracking time-to-first-token (TTFT) for streaming calls or total response duration for standard requests.
Implemented in litellm/router_strategy/lowest_latency.py, the LowestLatencyLoggingHandler stores recent latency values in a bounded list (max_latency_list_size) within the cache entry latency_key = f"{model_group}_map". During routing, it computes the average latency for each deployment, discards any exceeding rate limits, and selects the deployment with the lowest historical average.
Use this strategy for real-time chat interfaces or streaming applications where user-perceived speed is critical.
Least-Busy Routing
least-busy distributes traffic evenly by tracking the number of in-flight requests per deployment.
The LeastBusyLoggingHandler in litellm/router_strategy/least_busy.py increments a counter (<model_group>_request_count) in log_pre_api_call() before sending a request and decrements it in log_success_event() or log_failure_event() upon completion. The router selects the deployment with the smallest active request count, breaking ties randomly.
This approach prevents individual deployments from becoming overwhelmed during high-concurrency batch processing workloads.
Usage-Based Routing (lowest_tpm_rpm)
usage-based-routing (and its v2 variant) enforces strict compliance with provider token-per-minute (TPM) and request-per-minute (RPM) limits.
The handlers in litellm/router_strategy/lowest_tpm_rpm.py and lowest_tpm_rpm_v2.py maintain separate cache entries for per-minute TPM (<model_group>:tpm:<minute>) and RPM (<model_group>:rpm:<minute>). When routing, the handler calculates projected usage (current count + new request tokens) and returns the first deployment that remains below its configured limits, or the one with the lowest projected TPM when multiple options exist.
Version 2 includes refactored TTL handling for improved cache expiration management. Use this strategy when working with strict Azure OpenAI quotas or other providers imposing hard rate limits.
Implementation Examples
Basic Cost Optimization
import litellm
model_list = [
{
"model": "gpt-4",
"api_key": "sk-xxxx",
"metadata": {"model_group": "openai"},
},
{
"model": "gpt-4-32k",
"api_key": "sk-yyyy",
"metadata": {"model_group": "openai"},
},
]
response = litellm.completion(
model="gpt-4",
messages=[{"role": "user", "content": "Explain machine learning"}],
model_list=model_list,
routing_strategy="cost-based-routing",
)
The LowestCostLoggingHandler.async_get_available_deployments() automatically selects the cheaper deployment when both satisfy TPM/RPM constraints.
Latency-Aware Streaming
response = litellm.completion(
model="gpt-4",
messages=[{"role": "user", "content": "Explain quantum computing"}],
model_list=model_list,
routing_strategy="latency-based-routing",
routing_strategy_args={"lowest_latency_buffer": 0.2},
stream=True,
)
for chunk in response:
print(chunk.choices[0].delta.content, end="")
For streaming calls, the handler records TTFT in latency_key and averages recent values to predict the fastest deployment.
Rate Limit Protection
response = litellm.completion(
model="azure-gpt-3.5",
messages=[{"role": "user", "content": "Write a haiku"}],
model_list=model_list,
routing_strategy="usage-based-routing",
)
The handler automatically reads tpm and rpm limits from the provider's model info and avoids deployments that would exceed per-minute quotas.
High-Throughput Load Balancing
response = litellm.completion(
model="gpt-4",
messages=batch_messages,
model_list=model_list,
routing_strategy="least-busy",
num_retries=2,
)
The router tracks in-flight requests via <model_group>_request_count and distributes load across the least utilized deployments.
Summary
- Cost-based-routing minimizes spend by calculating per-token costs from
litellm.model_costand selecting the cheapest eligible deployment inlowest_cost.py. - Latency-based-routing reduces response time by averaging historical TTFT values stored in a bounded list within
lowest_latency.py. - Least-busy prevents overload by tracking active request counts incremented in
log_pre_api_call()and selecting deployments with minimal concurrency inleast_busy.py. - Usage-based-routing enforces provider limits by monitoring per-minute TPM/RPM counters in
lowest_tpm_rpm.py(v1) orlowest_tpm_rpm_v2.py(v2), returning only deployments with sufficient remaining quota.
Frequently Asked Questions
How do I configure multiple routing strategies simultaneously?
LiteLLM selects one primary strategy via the routing_strategy parameter, but you can combine strategies by implementing custom logic in your application layer or using tag-based routing to filter deployments before applying metric-based selection. The router's fallback mechanism will retry with the next-best deployment according to the chosen strategy's ranking.
What happens if all deployments exceed TPM/RPM limits?
When using usage-based-routing, if no deployment has sufficient remaining quota for the projected request size, the router raises a RateLimitError. For other strategies, the router filters out over-limit deployments; if none remain eligible, it either retries with exponential backoff or raises an exception depending on your num_retries configuration.
Does latency-based routing work for non-streaming requests?
Yes. The LowestLatencyLoggingHandler in lowest_latency.py records total response duration for standard requests and time-to-first-token (TTFT) for streaming calls. Both metrics are stored in the same bounded list and averaged to determine the fastest deployment regardless of the stream parameter.
Where does LiteLLM store routing metrics?
All strategies use the DualCache instance (in-memory with optional Redis backing) defined in base_routing_strategy.py. Cost data resides in <model_group>_map, latency in latency_key, busy counts in <model_group>_request_count, and rate limits in <model_group>:tpm:<minute> and <model_group>:rpm:<minute> entries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →