# LiteLLM Router Failover and Retry Logic: A Deep Dive into Resilient LLM Routing

> Explore LiteLLM router failover and retry logic. Automatically reroute failed LLM requests to backup models enhancing application resilience without code changes.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: deep-dive
- Published: 2026-03-26

---

**LiteLLM’s router automatically reroutes failed requests to backup models using configurable fallback chains, ensuring high availability across multiple LLM providers without changing your application code.**

The **BerriAI/litellm** repository provides a unified interface for calling hundreds of LLM providers through OpenAI-compatible APIs. At its core, the **Router** class ([`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)) implements sophisticated failover and retry logic that transparently handles rate limits, service outages, and context window overflows by switching to alternative deployments.

## How the LiteLLM Router Handles Failover

### Router Configuration and Entry Points

All failover behavior originates in **RouterConfig** ([`litellm/types/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/router.py)), which defines the schema for `fallbacks`, `context_window_fallbacks`, `content_policy_fallbacks`, and `max_fallbacks`. When you invoke `litellm.completion()` or the proxy’s `/v1/chat/completions` endpoint, the router first checks for fallback eligibility.

In [`litellm/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/utils.py), the router detects fallback requirements via the flag `is_completion_with_fallbacks = kwargs.get("fallbacks") is not None` (around line 1800). This boolean determines whether the request enters the resilient execution path or proceeds as a standard single-model call.

### Pre-Fallback Metadata Injection

Before attempting any model calls, the router executes `_update_kwargs_before_fallbacks` (lines 1345-1355 in [`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)). This method injects critical metadata—including tracing IDs and retry counters—into the request kwargs. By preserving this context across every attempt in the fallback chain, LiteLLM ensures observability tools can correlate primary and fallback requests as a single logical operation.

### Error Detection and Fallback Triggers

The router distinguishes between fatal errors and recoverable conditions that warrant failover. When the primary model raises a **MidStreamFallbackError**—commonly triggered by HTTP 429 rate limits, authentication failures, or content policy violations—the router immediately enters the fallback logic at line 1974 of [`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py).

Non-streaming errors are caught in the main execution loop, while streaming responses require special handling to ensure connection cleanup before attempting the next model in the chain.

### The Fallback Loop Mechanism

The core resilience logic resides in the fallback iteration loop (lines 2000-2080 in [`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)). When triggered, the router:

1. Retrieves the ordered fallback list for the primary model (up to `max_fallbacks` items)
2. Updates `kwargs["model"]` to the next fallback candidate
3. Invokes `_ageneric_api_call_with_fallbacks_helper` with the modified parameters
4. Returns immediately on success, or continues iterating through the chain

If all configured fallbacks exhaust without success, the router returns the final error to the client, preserving the complete error history in the response metadata.

### Specialized Fallback Types

LiteLLM provides domain-specific fallback configurations for distinct failure modes:

**Context Window Fallbacks**: When a request exceeds a model’s token limit (e.g., sending 8k tokens to a 4k-context model), the router consults `context_window_fallbacks` (defined in [`litellm/types/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/router.py) lines 162-168). The system automatically rewrites the prompt and retries with a model possessing a larger context window, such as routing from `gpt-3.5-turbo` to `gpt-4`.

**Content Policy Fallbacks**: For requests rejected by provider-specific safety filters, `content_policy_fallbacks` route to alternative models with different moderation thresholds, ensuring business-critical prompts can still be processed through "soft-filter" endpoints.

### Streaming Failover Handling

Streaming endpoints require graceful connection management to prevent resource leaks. The `stream_with_fallbacks` wrapper monitors the async generator for mid-stream errors. If a **MidStreamFallbackError** occurs during token generation, the router:

- Closes the original stream iterator cleanly
- Initiates a new streaming connection to the fallback model
- Preserves chunk ordering so the client receives a seamless token stream

The test suite in [`tests/test_litellm/test_streaming_connection_cleanup.py`](https://github.com/BerriAI/litellm/blob/main/tests/test_litellm/test_streaming_connection_cleanup.py) (lines 170-238) validates that connections are properly terminated before fallback initiation, preventing socket exhaustion under high retry volumes.

## Configuration Examples for Production Use

### Basic Model Fallback Configuration

Configure provider-agnostic failover for any completion call:

```python
import litellm

# Route from premium to standard model on any failure

litellm.router_fallbacks = {
    "gpt-4o": ["gpt-4o-mini", "gpt-3.5-turbo", "claude-3-haiku-20240307"]
}

response = litellm.completion(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Analyze this financial report"}],
    max_tokens=500,
    request_timeout=30
)

```

The router attempts `gpt-4o` first; if it returns a 429 or 5xx error, it automatically retries with `gpt-4o-mini`, continuing through the chain until success or exhaustion.

### Context Window Overflow Protection

Prevent token limit errors by preemptively routing large requests:

```python
litellm.context_window_fallbacks = {
    "gpt-3.5-turbo": ["gpt-4", "gpt-4o"],
    "claude-3-haiku-20240307": ["claude-3-sonnet-20240229"]
}

response = litellm.completion(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "Summarize this 50-page document..."}],
    max_tokens=4000  # Exceeds gpt-3.5-turbo's 4k limit

)

```

When the router detects the input exceeds the configured context window, it transparently switches to the fallback model without raising an error to your application.

### Streaming with Automatic Failover

Handle mid-stream failures for real-time applications:

```python
import asyncio
import litellm

async def resilient_stream():
    async for chunk in litellm.acompletion(
        model="azure/gpt-4o",
        messages=[{"role": "user", "content": "Generate a markdown table"}],
        stream=True,
        fallbacks=["openai/gpt-4o", "anthropic/claude-3-sonnet-20240229"],
        max_fallbacks=2
    ):
        content = chunk.choices[0].delta.get("content", "")
        print(content, end="", flush=True)

asyncio.run(resilient_stream())

```

If Azure returns a rate limit mid-generation, the router closes the Azure stream and resumes generation from OpenAI’s GPT-4o, maintaining output continuity.

## Observability and Monitoring

Every fallback attempt increments Prometheus counters defined in [`litellm/types/integrations/prometheus.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/integrations/prometheus.py) (lines 181-206). Monitor deployment health using these metrics:

- `litellm_deployment_successful_fallbacks`: Count of requests saved by failover
- `litellm_deployment_failed_fallbacks`: Count of exhausted fallback chains
- `litellm_deployment_latency`: Per-attempt latency for bottleneck identification

Integrate these with Grafana or Slack webhooks to alert when primary model failure rates exceed thresholds, indicating provider instability or quota exhaustion.

## Summary

- **RouterConfig** ([`litellm/types/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/router.py)) defines fallback chains through `fallbacks`, `context_window_fallbacks`, and `content_policy_fallbacks` parameters.
- The `_update_kwargs_before_fallbacks` method preserves request metadata across all retry attempts in the fallback loop.
- **MidStreamFallbackError** detection triggers the resilient execution path in [`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py) lines 1974-2080.
- Streaming failover requires explicit connection cleanup to prevent resource leaks, validated in the streaming cleanup test suite.
- Prometheus metrics provide real-time visibility into fallback success rates and deployment health.

## Frequently Asked Questions

### How does LiteLLM determine when to trigger a fallback?

LiteLLM checks for the `is_completion_with_fallbacks` flag in [`litellm/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/utils.py) (line 1800) by inspecting `kwargs.get("fallbacks")`. During execution, specific exceptions like **MidStreamFallbackError** or standard HTTP 429/5xx errors trigger entry into the fallback loop at line 1974 of [`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py). The router differentiates between retryable errors (rate limits, timeouts) and fatal errors (invalid auth, malformed requests) to avoid unnecessary failover attempts.

### What happens to streaming connections during a failover?

When a mid-stream error occurs, the router’s `stream_with_fallbacks` wrapper immediately closes the original async iterator to release the HTTP connection, then initializes a new streaming request to the next model in the fallback chain. The test file [`tests/test_litellm/test_streaming_connection_cleanup.py`](https://github.com/BerriAI/litellm/blob/main/tests/test_litellm/test_streaming_connection_cleanup.py) validates that sockets are properly terminated before the fallback attempt, preventing connection pool exhaustion during high-error scenarios.

### Can I configure different fallback chains for different error types?

Yes. LiteLLM supports three distinct fallback configurations in `RouterConfig`: `fallbacks` for general errors, `context_window_fallbacks` for token limit overflows, and `content_policy_fallbacks` for moderation violations. You can define independent routing logic for each scenario, such as falling back to cheaper models for rate limits but larger models for context window errors.

### How can I monitor fallback rates in production?

The router emits Prometheus metrics defined in [`litellm/types/integrations/prometheus.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/integrations/prometheus.py), specifically `litellm_deployment_successful_fallbacks` and `litellm_deployment_failed_fallbacks`. These counters increment inside the fallback loop in [`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py), allowing you to track resilience efficiency per deployment. Export these metrics to your observability stack to set alerts when fallback rates exceed normal operational thresholds.