# High Availability and Load Balancing LiteLLM Proxy: A Production Deployment Guide

> Achieve high availability and load balancing for your LiteLLM proxy. Learn how to deploy stateless, horizontally scaled, and automatically load-balanced model endpoints for production readiness.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: how-to-guide
- Published: 2026-03-26

---

**LiteLLM Proxy achieves high availability through stateless horizontal scaling and native load balancing across model deployments using a round-robin router that automatically skips unhealthy endpoints.**

The LiteLLM Proxy from BerriAI/litellm aggregates dozens of Large Language Model (LLM) backends behind a single OpenAI-compatible API. For production workloads requiring fault tolerance and scalability, the proxy supports both high availability (HA) deployment patterns and intelligent load balancing across multiple deployments of the same model.

## How Load Balancing Works in LiteLLM

LiteLLM implements load balancing through the **`Router`** class located in [`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py). The Router maintains a registry of **deployments** (model + endpoint combinations) and organizes them into **model groups** based on the `model_name` key.

When you designate a model group as load-balanced, the Router distributes incoming requests across all healthy deployments in that group using a round-robin strategy. The selection logic lives in `Router.get_available_deployment()`, which maintains an internal `_rr_counter` that increments with each request to cycle through available endpoints.

The Router evaluates deployment health before selection. If a deployment fails health checks, it is marked as unhealthy and excluded from the rotation until it recovers. This ensures that traffic only flows to responsive backends.

## Configuring Load-Balanced Model Groups

To enable load balancing, set **`load_balanced: true`** in your [`config.yaml`](https://github.com/BerriAI/litellm/blob/main/config.yaml) (or equivalent JSON configuration passed via `WORKER_CONFIG`). The configuration parser in [`litellm/proxy/config_utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/config_utils.py) extracts this flag during initialization and registers the deployments as a pooled group.

```yaml
model_list:
  - model_name: "gpt-3.5-turbo"
    load_balanced: true                # Enables load balancing

    deployments:
      - litellm_params:
          model: "azure/gpt-35-turbo-1"
          api_key: "AZURE_KEY_1"
      - litellm_params:
          model: "azure/gpt-35-turbo-2"
          api_key: "AZURE_KEY_2"

```

When a request arrives for `gpt-3.5-turbo`, the Router selects the next deployment in the sequence using the modulo operator on the round-robin counter. This simple distribution ensures even traffic spread across all configured endpoints.

## Health Checks and Automatic Failover

The proxy implements proactive health monitoring through **[`litellm/proxy/health_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/health_check.py)**. A background coroutine periodically invokes health endpoints for each deployment using HTTP or gRPC probes.

If a deployment fails to respond or returns an error status, the health check mechanism updates the Router's internal state to mark that specific deployment as **unhealthy**. Subsequent requests automatically bypass failed deployments until the health probe detects recovery. This provides circuit-breaker functionality without manual intervention.

The proxy also optimizes connection pooling through a **shared AIOHTTP session**. The function `_initialize_shared_aiohttp_session` in [`litellm/proxy/proxy_server.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py) (lines 40–64) creates a reusable TCP connection pool that reduces latency and improves scalability when communicating with multiple backend providers.

## Guardrail Load Balancing

LiteLLM extends load balancing beyond model inference to **guardrails** (pre-call moderation and post-call filters). You can distribute guardrail workloads across multiple service instances by defining multiple deployments with the same `guardrail_name`.

The helper `_should_use_guardrail_load_balancing` in [`litellm/proxy/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/utils.py) (lines 872–883) detects when multiple guardrail definitions share a name. When processing requests, `_execute_guardrail_with_load_balancing` (lines 945–960) iterates through available guardrail deployments, distributing moderation workloads across the pool.

```yaml
guardrails:
  - guardrail_name: "toxicity_check"
    type: "moderation"
    litellm_params:
      model: "openai/moderation-1"
      api_key: "sk-..."
  - guardrail_name: "toxicity_check"
    type: "moderation"
    litellm_params:
      model: "openai/moderation-2"
      api_key: "sk-..."

```

## Deploying for High Availability

The LiteLLM Proxy is **stateless** by design, storing only ephemeral data (like round-robin counters) in memory. This architecture makes horizontal scaling straightforward: you can run multiple proxy replicas behind any standard TCP/HTTP load balancer (NGINX, Envoy, Kubernetes Ingress, or cloud load balancers).

For true high availability:

- **Run multiple replicas**: Deploy at least two proxy pods using Kubernetes Deployments or Docker Swarm replicated mode
- **Externalize shared state**: Use Redis Sentinel or managed Redis for shared rate-limiting and spend-tracking data (the proxy caches these in Redis rather than local memory)
- **Graceful shutdowns**: FastAPI lifespan hooks in [`proxy_server.py`](https://github.com/BerriAI/litellm/blob/main/proxy_server.py) ensure the shared AIOHTTP session and database connections close cleanly on pod termination
- **Configuration updates**: Mount your config as a ConfigMap or volume; updating `WORKER_CONFIG` triggers a configuration reload

## Implementation Examples

### YAML Configuration with Load Balancing

Define your model group with the `load_balanced` flag and multiple deployments:

```yaml

# /etc/litellm/config.yaml

model_list:
  - model_name: "anthropic.claude-3"
    load_balanced: true
    deployments:
      - litellm_params:
          model: "anthropic/claude-3-opus-1"
          api_key: "${ANTHROPIC_KEY_1}"
      - litellm_params:
          model: "anthropic/claude-3-opus-2"
          api_key: "${ANTHROPIC_KEY_2}"
      - litellm_params:
          model: "bedrock/anthropic.claude-3-opus"
          aws_access_key_id: "${AWS_KEY}"
          aws_secret_access_key: "${AWS_SECRET}"

```

### Docker Compose HA Setup

Deploy three stateless replicas behind your own load balancer:

```yaml
version: "3.8"
services:
  litellm-proxy:
    image: litellm/litellm:latest
    environment:
      - WORKER_CONFIG=/app/config.yaml
      - DATABASE_URL=postgresql://user:pw@db:5432/litellm
      - REDIS_HOST=redis
    volumes:
      - ./config.yaml:/app/config.yaml:ro
    deploy:
      mode: replicated
      replicas: 3
      restart_policy:
        condition: on-failure
    ports:
      - "4000"

```

### Python SDK Router Usage

Use the Router programmatically for custom applications:

```python
import litellm

router = litellm.Router(
    model_list=[
        {
            "model_name": "gpt-4",
            "load_balanced": True,
            "deployments": [
                {
                    "litellm_params": {
                        "model": "openai/gpt-4",
                        "api_key": "sk-..."
                    }
                },
                {
                    "litellm_params": {
                        "model": "azure/gpt-4",
                        "api_key": "AZURE_KEY",
                        "api_base": "https://my-resource.openai.azure.com"
                    }
                }
            ],
        }
    ]
)

# Automatically load-balanced across both deployments

response = router.completion(
    model="gpt-4",
    messages=[{"role": "user", "content": "Explain quantum computing"}]
)

```

## Summary

- **Stateless Architecture**: The LiteLLM Proxy requires no cross-pod communication, enabling simple horizontal scaling behind standard load balancers.
- **Native Load Balancing**: Set `load_balanced: true` in [`config.yaml`](https://github.com/BerriAI/litellm/blob/main/config.yaml) to enable round-robin distribution across model deployments via the `Router` class in [`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py).
- **Automatic Health Failover**: The health check module in [`litellm/proxy/health_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/health_check.py) continuously monitors deployments and removes unhealthy endpoints from rotation.
- **Guardrail Distribution**: Load balance moderation tasks across multiple guardrail instances using the helpers in [`litellm/proxy/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/utils.py).
- **Connection Optimization**: Shared AIOHTTP sessions reduce TCP overhead when managing dozens of backend connections.

## Frequently Asked Questions

### What is the difference between high availability and load balancing in LiteLLM?

**High availability** refers to the proxy's deployment topology: running multiple stateless replicas behind a load balancer to eliminate single points of failure. **Load balancing** is the traffic distribution mechanism within the application logic, where the `Router` class spreads requests across multiple model deployments. You can have load balancing without HA (single proxy instance), but production environments typically implement both.

### How does LiteLLM detect unhealthy model deployments?

The proxy runs periodic health probes defined in [`litellm/proxy/health_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/health_check.py). These HTTP/gRPC checks test each deployment's availability. When a deployment fails consecutive checks, the Router marks it as unhealthy in its internal state map, causing `get_available_deployment()` to skip that endpoint until health checks pass again.

### Can I use load balancing strategies other than round-robin?

Currently, the Router implements **round-robin** selection using a modulo counter (`_rr_counter % len(healthy_deployments)`). The source code indicates that **latency-based weighting** is planned for future versions. For now, you can implement custom strategies by subclassing the `Router` class or using the LiteLLM Python SDK to build your own selection logic on top of the deployment health data.

### How do I share rate limit data across multiple LiteLLM proxy instances?

While the proxy itself is stateless, rate limits and spend tracking require shared storage. Configure the `DATABASE_URL` environment variable to point to a clustered PostgreSQL instance, and set `REDIS_HOST` to a highly available Redis cluster (Redis Sentinel or managed cloud Redis). This ensures that all proxy replicas share synchronized state for rate-limiting counters and usage tracking.