How LiteLLM Handles Rate Limiting and Cooldown Periods: A Complete Technical Guide
LiteLLM protects inference workloads through three integrated mechanisms: per-deployment TPM/RPM limits enforced via pre-call checks, saturation-aware priority reservations for global capacity management, and automatic cooldown periods that remove unhealthy backends from rotation.
LiteLLM, the popular open-source LLM gateway by BerriAI, implements sophisticated rate limiting and cooldown period management to prevent individual model deployments from overwhelming upstream providers while maintaining high availability. The architecture combines model-level token and request throttling with dynamic, priority-aware global limits and automated failure isolation. This article examines the implementation details found in the litellm source code, including specific file paths and configuration patterns used in production environments.
Model-Level TPM and RPM Enforcement
When enforce_model_rate_limits=True is configured, LiteLLM instantiates the ModelRateLimitingCheck class from litellm/router_utils/pre_call_checks/model_rate_limit_check.py to intercept requests before they reach the provider.
Deployment-Specific Limit Checks
The pre-call check extracts TPM (tokens-per-minute) and RPM (requests-per-minute) limits from the deployment configuration—searching top-level keys, litellm_params, or model_info dictionaries (lines 52-78). For each request, LiteLLM constructs cache keys scoped to the model ID, deployment name, and current minute using the format model_id:deployment_name:tpm:HH-MM (lines 81-89).
TPM enforcement operates via local cache reads only. If the stored token count meets or exceeds the limit, the system immediately raises a RateLimitError (HTTP 429) without incrementing counters (lines 112-122). RPM enforcement uses DualCache.increment_cache for atomic increments; if the new count exceeds the configured limit, the request is rejected (lines 132-140).
Successful requests update counters asynchronously via async_log_success_event, which adds actual token usage to the TPM counter (lines 46-84).
Dynamic Priority-Aware Rate Limiting
The DynamicRateLimiter in litellm/proxy/hooks/dynamic_rate_limiter_v3.py provides global capacity management with saturation awareness and priority reservations. This enterprise feature operates as a proxy hook (async_pre_call_hook) using a three-phase validation strategy.
Three-Phase Atomic Validation
Phase 1: Read-Only Saturation Check
The _check_model_saturation method queries Redis-backed counters to calculate a saturation ratio (0 = empty, 1 = at capacity) without mutating state (lines 30-96).
Phase 2: Priority Evaluation
When saturation exceeds the configured saturation_threshold (default requires Enterprise configuration), the system enforces priority-based limits. High-priority callers receive reserved capacity percentages defined in priority_reservation mappings (lines 40-45).
Phase 3: Atomic Increment
Using Lua-scripted operations, the limiter increments model-wide counters always, and priority counters only if the request passes all limits. This separation prevents "phantom usage" where blocked requests would still consume capacity (lines 20-31, 63-84).
Automatic Cooldown Periods for Failing Deployments
LiteLLM uses the CooldownCache mechanism in litellm/router_utils/cooldown_cache.py to temporarily remove unhealthy deployments from the routing pool after repeated failures.
CooldownCache Implementation
When a deployment fails more than allowed_fails times (default 1), the router calls add_deployment_to_cooldown with a key formatted as deployment:<model_id>:cooldown (line 9). The cached value contains:
- Masked exception messages processed by
SensitiveDataMaskerto prevent secret leakage (lines 36-41) - HTTP status codes
- Unix timestamps
- Configurable TTL (default 60 seconds via
DEFAULT_COOLDOWN_TIME_SECONDS)
During routing, _filter_cooldown_deployments excludes any deployment with an active cooldown entry, ensuring traffic routes only to healthy backends.
Configuration Parameters
| Setting | Description | Default |
|---|---|---|
allowed_fails |
Number of consecutive failures before cooldown | 1 |
cooldown_time |
Duration in seconds to withhold traffic | 60 |
enforce_model_rate_limits |
Enable TPM/RPM pre-call checks | False |
Practical Configuration Examples
Enabling Model-Level Rate Limits
from litellm import Router
router = Router(
model_list=[
{
"model_name": "openai-gpt-4",
"litellm_params": {"model": "gpt-4"},
"tpm": 300_000, # 300k tokens per minute
"rpm": 60, # 60 requests per minute
},
],
enforce_model_rate_limits=True, # Activate pre-call checks
allowed_fails=2, # Cooldown after 2 failures
cooldown_time=120, # 2-minute cooldown period
)
Configuring Priority-Aware Dynamic Limits
# config.yaml
litellm_settings:
dynamic_rate_limiter: true
priority_reservation:
high: 0.7 # Reserve 70% for high priority when saturated
medium: 0.2
low: 0.1
saturation_threshold: 0.8
saturation_check_cache_ttl: 5
Automatic Cooldown Configuration
router = Router(
model_list=[{
"model_name": "azure-gpt-35",
"litellm_params": {
"model": "azure/deployment-1",
"api_key": "<key>",
"api_base": "https://example.openai.azure.com/",
},
}],
allowed_fails=1, # Immediate cooldown on first failure
cooldown_time=90, # 90-second exclusion from rotation
)
On failure, the system stores a CooldownCacheValue with masked error details and a 90-second TTL, automatically reintegrating the deployment once healthy.
Summary
- Per-deployment protection: The
ModelRateLimitingCheckclass enforces TPM/RPM limits using minute-scoped cache keys before requests leave the proxy, preventing individual providers from receiving excessive traffic. - Global capacity management: The dynamic rate limiter v3 implements three-phase atomic checks with priority reservations, ensuring high-priority workloads receive guaranteed capacity during saturation.
- Failure isolation: The
CooldownCacheautomatically removes failing deployments using masked error tracking and configurable TTLs, preventing cascade failures across the model pool. - Flexible configuration: All mechanisms support granular tuning via Router constructor parameters or
litellm_settingsin YAML configuration files.
Frequently Asked Questions
How does LiteLLM track rate limits for individual model deployments?
LiteLLM tracks rate limits using scoped cache keys in the format model_id:deployment_name:tpm:HH-MM or model_id:deployment_name:rpm:HH-MM. TPM counters are read locally before incrementing, while RPM counters use DualCache.increment_cache for atomic updates. The system checks these counters in the pre_call_check method of ModelRateLimitingCheck before forwarding requests to providers.
What happens when a deployment exceeds its configured rate limits?
When limits are exceeded, LiteLLM raises a RateLimitError with HTTP status 429 immediately during the pre-call phase. For TPM violations, this occurs after reading the local cache (lines 112-122 in model_rate_limit_check.py). For RPM violations, the error triggers after the atomic increment reveals the count exceeds the threshold (lines 132-140). The request never reaches the upstream provider.
How does LiteLLM prioritize requests when the system approaches capacity?
The dynamic rate limiter v3 calculates a saturation ratio by querying current usage counters. When this ratio exceeds saturation_threshold (e.g., 0.8), the system enforces priority reservations configured in priority_reservation. High-priority requests proceed if capacity exists within their reserved percentage, while lower-priority requests may be rate-limited. The three-phase atomic increment ensures accurate capacity tracking without phantom usage.
Can I customize how long a deployment stays in cooldown after failures?
Yes, configure the cooldown_time parameter in the Router constructor (in seconds) or set the LITELLM_COOLDOWN_TIME environment variable. The default is 60 seconds (DEFAULT_COOLDOWN_TIME_SECONDS). You can also adjust allowed_fails to control how many consecutive errors trigger cooldown. The CooldownCache stores entries with your specified TTL and automatically expires them, reintegrating the deployment into the routing pool.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →