LiteLLM Router Failover and Retry Logic: A Deep Dive into Resilient LLM Routing
LiteLLM’s router automatically reroutes failed requests to backup models using configurable fallback chains, ensuring high availability across multiple LLM providers without changing your application code.
The BerriAI/litellm repository provides a unified interface for calling hundreds of LLM providers through OpenAI-compatible APIs. At its core, the Router class (litellm/router.py) implements sophisticated failover and retry logic that transparently handles rate limits, service outages, and context window overflows by switching to alternative deployments.
How the LiteLLM Router Handles Failover
Router Configuration and Entry Points
All failover behavior originates in RouterConfig (litellm/types/router.py), which defines the schema for fallbacks, context_window_fallbacks, content_policy_fallbacks, and max_fallbacks. When you invoke litellm.completion() or the proxy’s /v1/chat/completions endpoint, the router first checks for fallback eligibility.
In litellm/utils.py, the router detects fallback requirements via the flag is_completion_with_fallbacks = kwargs.get("fallbacks") is not None (around line 1800). This boolean determines whether the request enters the resilient execution path or proceeds as a standard single-model call.
Pre-Fallback Metadata Injection
Before attempting any model calls, the router executes _update_kwargs_before_fallbacks (lines 1345-1355 in litellm/router.py). This method injects critical metadata—including tracing IDs and retry counters—into the request kwargs. By preserving this context across every attempt in the fallback chain, LiteLLM ensures observability tools can correlate primary and fallback requests as a single logical operation.
Error Detection and Fallback Triggers
The router distinguishes between fatal errors and recoverable conditions that warrant failover. When the primary model raises a MidStreamFallbackError—commonly triggered by HTTP 429 rate limits, authentication failures, or content policy violations—the router immediately enters the fallback logic at line 1974 of litellm/router.py.
Non-streaming errors are caught in the main execution loop, while streaming responses require special handling to ensure connection cleanup before attempting the next model in the chain.
The Fallback Loop Mechanism
The core resilience logic resides in the fallback iteration loop (lines 2000-2080 in litellm/router.py). When triggered, the router:
- Retrieves the ordered fallback list for the primary model (up to
max_fallbacksitems) - Updates
kwargs["model"]to the next fallback candidate - Invokes
_ageneric_api_call_with_fallbacks_helperwith the modified parameters - Returns immediately on success, or continues iterating through the chain
If all configured fallbacks exhaust without success, the router returns the final error to the client, preserving the complete error history in the response metadata.
Specialized Fallback Types
LiteLLM provides domain-specific fallback configurations for distinct failure modes:
Context Window Fallbacks: When a request exceeds a model’s token limit (e.g., sending 8k tokens to a 4k-context model), the router consults context_window_fallbacks (defined in litellm/types/router.py lines 162-168). The system automatically rewrites the prompt and retries with a model possessing a larger context window, such as routing from gpt-3.5-turbo to gpt-4.
Content Policy Fallbacks: For requests rejected by provider-specific safety filters, content_policy_fallbacks route to alternative models with different moderation thresholds, ensuring business-critical prompts can still be processed through "soft-filter" endpoints.
Streaming Failover Handling
Streaming endpoints require graceful connection management to prevent resource leaks. The stream_with_fallbacks wrapper monitors the async generator for mid-stream errors. If a MidStreamFallbackError occurs during token generation, the router:
- Closes the original stream iterator cleanly
- Initiates a new streaming connection to the fallback model
- Preserves chunk ordering so the client receives a seamless token stream
The test suite in tests/test_litellm/test_streaming_connection_cleanup.py (lines 170-238) validates that connections are properly terminated before fallback initiation, preventing socket exhaustion under high retry volumes.
Configuration Examples for Production Use
Basic Model Fallback Configuration
Configure provider-agnostic failover for any completion call:
import litellm
# Route from premium to standard model on any failure
litellm.router_fallbacks = {
"gpt-4o": ["gpt-4o-mini", "gpt-3.5-turbo", "claude-3-haiku-20240307"]
}
response = litellm.completion(
model="gpt-4o",
messages=[{"role": "user", "content": "Analyze this financial report"}],
max_tokens=500,
request_timeout=30
)
The router attempts gpt-4o first; if it returns a 429 or 5xx error, it automatically retries with gpt-4o-mini, continuing through the chain until success or exhaustion.
Context Window Overflow Protection
Prevent token limit errors by preemptively routing large requests:
litellm.context_window_fallbacks = {
"gpt-3.5-turbo": ["gpt-4", "gpt-4o"],
"claude-3-haiku-20240307": ["claude-3-sonnet-20240229"]
}
response = litellm.completion(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "Summarize this 50-page document..."}],
max_tokens=4000 # Exceeds gpt-3.5-turbo's 4k limit
)
When the router detects the input exceeds the configured context window, it transparently switches to the fallback model without raising an error to your application.
Streaming with Automatic Failover
Handle mid-stream failures for real-time applications:
import asyncio
import litellm
async def resilient_stream():
async for chunk in litellm.acompletion(
model="azure/gpt-4o",
messages=[{"role": "user", "content": "Generate a markdown table"}],
stream=True,
fallbacks=["openai/gpt-4o", "anthropic/claude-3-sonnet-20240229"],
max_fallbacks=2
):
content = chunk.choices[0].delta.get("content", "")
print(content, end="", flush=True)
asyncio.run(resilient_stream())
If Azure returns a rate limit mid-generation, the router closes the Azure stream and resumes generation from OpenAI’s GPT-4o, maintaining output continuity.
Observability and Monitoring
Every fallback attempt increments Prometheus counters defined in litellm/types/integrations/prometheus.py (lines 181-206). Monitor deployment health using these metrics:
litellm_deployment_successful_fallbacks: Count of requests saved by failoverlitellm_deployment_failed_fallbacks: Count of exhausted fallback chainslitellm_deployment_latency: Per-attempt latency for bottleneck identification
Integrate these with Grafana or Slack webhooks to alert when primary model failure rates exceed thresholds, indicating provider instability or quota exhaustion.
Summary
- RouterConfig (
litellm/types/router.py) defines fallback chains throughfallbacks,context_window_fallbacks, andcontent_policy_fallbacksparameters. - The
_update_kwargs_before_fallbacksmethod preserves request metadata across all retry attempts in the fallback loop. - MidStreamFallbackError detection triggers the resilient execution path in
litellm/router.pylines 1974-2080. - Streaming failover requires explicit connection cleanup to prevent resource leaks, validated in the streaming cleanup test suite.
- Prometheus metrics provide real-time visibility into fallback success rates and deployment health.
Frequently Asked Questions
How does LiteLLM determine when to trigger a fallback?
LiteLLM checks for the is_completion_with_fallbacks flag in litellm/utils.py (line 1800) by inspecting kwargs.get("fallbacks"). During execution, specific exceptions like MidStreamFallbackError or standard HTTP 429/5xx errors trigger entry into the fallback loop at line 1974 of litellm/router.py. The router differentiates between retryable errors (rate limits, timeouts) and fatal errors (invalid auth, malformed requests) to avoid unnecessary failover attempts.
What happens to streaming connections during a failover?
When a mid-stream error occurs, the router’s stream_with_fallbacks wrapper immediately closes the original async iterator to release the HTTP connection, then initializes a new streaming request to the next model in the fallback chain. The test file tests/test_litellm/test_streaming_connection_cleanup.py validates that sockets are properly terminated before the fallback attempt, preventing connection pool exhaustion during high-error scenarios.
Can I configure different fallback chains for different error types?
Yes. LiteLLM supports three distinct fallback configurations in RouterConfig: fallbacks for general errors, context_window_fallbacks for token limit overflows, and content_policy_fallbacks for moderation violations. You can define independent routing logic for each scenario, such as falling back to cheaper models for rate limits but larger models for context window errors.
How can I monitor fallback rates in production?
The router emits Prometheus metrics defined in litellm/types/integrations/prometheus.py, specifically litellm_deployment_successful_fallbacks and litellm_deployment_failed_fallbacks. These counters increment inside the fallback loop in litellm/router.py, allowing you to track resilience efficiency per deployment. Export these metrics to your observability stack to set alerts when fallback rates exceed normal operational thresholds.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →