High Availability and Load Balancing LiteLLM Proxy: A Production Deployment Guide
LiteLLM Proxy achieves high availability through stateless horizontal scaling and native load balancing across model deployments using a round-robin router that automatically skips unhealthy endpoints.
The LiteLLM Proxy from BerriAI/litellm aggregates dozens of Large Language Model (LLM) backends behind a single OpenAI-compatible API. For production workloads requiring fault tolerance and scalability, the proxy supports both high availability (HA) deployment patterns and intelligent load balancing across multiple deployments of the same model.
How Load Balancing Works in LiteLLM
LiteLLM implements load balancing through the Router class located in litellm/router.py. The Router maintains a registry of deployments (model + endpoint combinations) and organizes them into model groups based on the model_name key.
When you designate a model group as load-balanced, the Router distributes incoming requests across all healthy deployments in that group using a round-robin strategy. The selection logic lives in Router.get_available_deployment(), which maintains an internal _rr_counter that increments with each request to cycle through available endpoints.
The Router evaluates deployment health before selection. If a deployment fails health checks, it is marked as unhealthy and excluded from the rotation until it recovers. This ensures that traffic only flows to responsive backends.
Configuring Load-Balanced Model Groups
To enable load balancing, set load_balanced: true in your config.yaml (or equivalent JSON configuration passed via WORKER_CONFIG). The configuration parser in litellm/proxy/config_utils.py extracts this flag during initialization and registers the deployments as a pooled group.
model_list:
- model_name: "gpt-3.5-turbo"
load_balanced: true # Enables load balancing
deployments:
- litellm_params:
model: "azure/gpt-35-turbo-1"
api_key: "AZURE_KEY_1"
- litellm_params:
model: "azure/gpt-35-turbo-2"
api_key: "AZURE_KEY_2"
When a request arrives for gpt-3.5-turbo, the Router selects the next deployment in the sequence using the modulo operator on the round-robin counter. This simple distribution ensures even traffic spread across all configured endpoints.
Health Checks and Automatic Failover
The proxy implements proactive health monitoring through litellm/proxy/health_check.py. A background coroutine periodically invokes health endpoints for each deployment using HTTP or gRPC probes.
If a deployment fails to respond or returns an error status, the health check mechanism updates the Router's internal state to mark that specific deployment as unhealthy. Subsequent requests automatically bypass failed deployments until the health probe detects recovery. This provides circuit-breaker functionality without manual intervention.
The proxy also optimizes connection pooling through a shared AIOHTTP session. The function _initialize_shared_aiohttp_session in litellm/proxy/proxy_server.py (lines 40–64) creates a reusable TCP connection pool that reduces latency and improves scalability when communicating with multiple backend providers.
Guardrail Load Balancing
LiteLLM extends load balancing beyond model inference to guardrails (pre-call moderation and post-call filters). You can distribute guardrail workloads across multiple service instances by defining multiple deployments with the same guardrail_name.
The helper _should_use_guardrail_load_balancing in litellm/proxy/utils.py (lines 872–883) detects when multiple guardrail definitions share a name. When processing requests, _execute_guardrail_with_load_balancing (lines 945–960) iterates through available guardrail deployments, distributing moderation workloads across the pool.
guardrails:
- guardrail_name: "toxicity_check"
type: "moderation"
litellm_params:
model: "openai/moderation-1"
api_key: "sk-..."
- guardrail_name: "toxicity_check"
type: "moderation"
litellm_params:
model: "openai/moderation-2"
api_key: "sk-..."
Deploying for High Availability
The LiteLLM Proxy is stateless by design, storing only ephemeral data (like round-robin counters) in memory. This architecture makes horizontal scaling straightforward: you can run multiple proxy replicas behind any standard TCP/HTTP load balancer (NGINX, Envoy, Kubernetes Ingress, or cloud load balancers).
For true high availability:
- Run multiple replicas: Deploy at least two proxy pods using Kubernetes Deployments or Docker Swarm replicated mode
- Externalize shared state: Use Redis Sentinel or managed Redis for shared rate-limiting and spend-tracking data (the proxy caches these in Redis rather than local memory)
- Graceful shutdowns: FastAPI lifespan hooks in
proxy_server.pyensure the shared AIOHTTP session and database connections close cleanly on pod termination - Configuration updates: Mount your config as a ConfigMap or volume; updating
WORKER_CONFIGtriggers a configuration reload
Implementation Examples
YAML Configuration with Load Balancing
Define your model group with the load_balanced flag and multiple deployments:
# /etc/litellm/config.yaml
model_list:
- model_name: "anthropic.claude-3"
load_balanced: true
deployments:
- litellm_params:
model: "anthropic/claude-3-opus-1"
api_key: "${ANTHROPIC_KEY_1}"
- litellm_params:
model: "anthropic/claude-3-opus-2"
api_key: "${ANTHROPIC_KEY_2}"
- litellm_params:
model: "bedrock/anthropic.claude-3-opus"
aws_access_key_id: "${AWS_KEY}"
aws_secret_access_key: "${AWS_SECRET}"
Docker Compose HA Setup
Deploy three stateless replicas behind your own load balancer:
version: "3.8"
services:
litellm-proxy:
image: litellm/litellm:latest
environment:
- WORKER_CONFIG=/app/config.yaml
- DATABASE_URL=postgresql://user:pw@db:5432/litellm
- REDIS_HOST=redis
volumes:
- ./config.yaml:/app/config.yaml:ro
deploy:
mode: replicated
replicas: 3
restart_policy:
condition: on-failure
ports:
- "4000"
Python SDK Router Usage
Use the Router programmatically for custom applications:
import litellm
router = litellm.Router(
model_list=[
{
"model_name": "gpt-4",
"load_balanced": True,
"deployments": [
{
"litellm_params": {
"model": "openai/gpt-4",
"api_key": "sk-..."
}
},
{
"litellm_params": {
"model": "azure/gpt-4",
"api_key": "AZURE_KEY",
"api_base": "https://my-resource.openai.azure.com"
}
}
],
}
]
)
# Automatically load-balanced across both deployments
response = router.completion(
model="gpt-4",
messages=[{"role": "user", "content": "Explain quantum computing"}]
)
Summary
- Stateless Architecture: The LiteLLM Proxy requires no cross-pod communication, enabling simple horizontal scaling behind standard load balancers.
- Native Load Balancing: Set
load_balanced: trueinconfig.yamlto enable round-robin distribution across model deployments via theRouterclass inlitellm/router.py. - Automatic Health Failover: The health check module in
litellm/proxy/health_check.pycontinuously monitors deployments and removes unhealthy endpoints from rotation. - Guardrail Distribution: Load balance moderation tasks across multiple guardrail instances using the helpers in
litellm/proxy/utils.py. - Connection Optimization: Shared AIOHTTP sessions reduce TCP overhead when managing dozens of backend connections.
Frequently Asked Questions
What is the difference between high availability and load balancing in LiteLLM?
High availability refers to the proxy's deployment topology: running multiple stateless replicas behind a load balancer to eliminate single points of failure. Load balancing is the traffic distribution mechanism within the application logic, where the Router class spreads requests across multiple model deployments. You can have load balancing without HA (single proxy instance), but production environments typically implement both.
How does LiteLLM detect unhealthy model deployments?
The proxy runs periodic health probes defined in litellm/proxy/health_check.py. These HTTP/gRPC checks test each deployment's availability. When a deployment fails consecutive checks, the Router marks it as unhealthy in its internal state map, causing get_available_deployment() to skip that endpoint until health checks pass again.
Can I use load balancing strategies other than round-robin?
Currently, the Router implements round-robin selection using a modulo counter (_rr_counter % len(healthy_deployments)). The source code indicates that latency-based weighting is planned for future versions. For now, you can implement custom strategies by subclassing the Router class or using the LiteLLM Python SDK to build your own selection logic on top of the deployment health data.
How do I share rate limit data across multiple LiteLLM proxy instances?
While the proxy itself is stateless, rate limits and spend tracking require shared storage. Configure the DATABASE_URL environment variable to point to a clustered PostgreSQL instance, and set REDIS_HOST to a highly available Redis cluster (Redis Sentinel or managed cloud Redis). This ensures that all proxy replicas share synchronized state for rate-limiting counters and usage tracking.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →