How LiteLLM Handles Rate Limiting and Cooldown Periods: A Complete Technical Guide

LiteLLM protects inference workloads through three integrated mechanisms: per-deployment TPM/RPM limits enforced via pre-call checks, saturation-aware priority reservations for global capacity management, and automatic cooldown periods that remove unhealthy backends from rotation.

LiteLLM, the popular open-source LLM gateway by BerriAI, implements sophisticated rate limiting and cooldown period management to prevent individual model deployments from overwhelming upstream providers while maintaining high availability. The architecture combines model-level token and request throttling with dynamic, priority-aware global limits and automated failure isolation. This article examines the implementation details found in the litellm source code, including specific file paths and configuration patterns used in production environments.

Model-Level TPM and RPM Enforcement

When enforce_model_rate_limits=True is configured, LiteLLM instantiates the ModelRateLimitingCheck class from litellm/router_utils/pre_call_checks/model_rate_limit_check.py to intercept requests before they reach the provider.

Deployment-Specific Limit Checks

The pre-call check extracts TPM (tokens-per-minute) and RPM (requests-per-minute) limits from the deployment configuration—searching top-level keys, litellm_params, or model_info dictionaries (lines 52-78). For each request, LiteLLM constructs cache keys scoped to the model ID, deployment name, and current minute using the format model_id:deployment_name:tpm:HH-MM (lines 81-89).

TPM enforcement operates via local cache reads only. If the stored token count meets or exceeds the limit, the system immediately raises a RateLimitError (HTTP 429) without incrementing counters (lines 112-122). RPM enforcement uses DualCache.increment_cache for atomic increments; if the new count exceeds the configured limit, the request is rejected (lines 132-140).

Successful requests update counters asynchronously via async_log_success_event, which adds actual token usage to the TPM counter (lines 46-84).

Dynamic Priority-Aware Rate Limiting

The DynamicRateLimiter in litellm/proxy/hooks/dynamic_rate_limiter_v3.py provides global capacity management with saturation awareness and priority reservations. This enterprise feature operates as a proxy hook (async_pre_call_hook) using a three-phase validation strategy.

Three-Phase Atomic Validation

Phase 1: Read-Only Saturation Check
The _check_model_saturation method queries Redis-backed counters to calculate a saturation ratio (0 = empty, 1 = at capacity) without mutating state (lines 30-96).

Phase 2: Priority Evaluation
When saturation exceeds the configured saturation_threshold (default requires Enterprise configuration), the system enforces priority-based limits. High-priority callers receive reserved capacity percentages defined in priority_reservation mappings (lines 40-45).

Phase 3: Atomic Increment
Using Lua-scripted operations, the limiter increments model-wide counters always, and priority counters only if the request passes all limits. This separation prevents "phantom usage" where blocked requests would still consume capacity (lines 20-31, 63-84).

Automatic Cooldown Periods for Failing Deployments

LiteLLM uses the CooldownCache mechanism in litellm/router_utils/cooldown_cache.py to temporarily remove unhealthy deployments from the routing pool after repeated failures.

CooldownCache Implementation

When a deployment fails more than allowed_fails times (default 1), the router calls add_deployment_to_cooldown with a key formatted as deployment:<model_id>:cooldown (line 9). The cached value contains:

  • Masked exception messages processed by SensitiveDataMasker to prevent secret leakage (lines 36-41)
  • HTTP status codes
  • Unix timestamps
  • Configurable TTL (default 60 seconds via DEFAULT_COOLDOWN_TIME_SECONDS)

During routing, _filter_cooldown_deployments excludes any deployment with an active cooldown entry, ensuring traffic routes only to healthy backends.

Configuration Parameters

Setting Description Default
allowed_fails Number of consecutive failures before cooldown 1
cooldown_time Duration in seconds to withhold traffic 60
enforce_model_rate_limits Enable TPM/RPM pre-call checks False

Practical Configuration Examples

Enabling Model-Level Rate Limits

from litellm import Router

router = Router(
    model_list=[
        {
            "model_name": "openai-gpt-4",
            "litellm_params": {"model": "gpt-4"},
            "tpm": 300_000,            # 300k tokens per minute

            "rpm": 60,                 # 60 requests per minute

        },
    ],
    enforce_model_rate_limits=True,   # Activate pre-call checks

    allowed_fails=2,                  # Cooldown after 2 failures

    cooldown_time=120,                # 2-minute cooldown period

)

Configuring Priority-Aware Dynamic Limits


# config.yaml

litellm_settings:
  dynamic_rate_limiter: true
  
  priority_reservation:
    high: 0.7      # Reserve 70% for high priority when saturated

    medium: 0.2
    low: 0.1
  
  saturation_threshold: 0.8
  saturation_check_cache_ttl: 5

Automatic Cooldown Configuration

router = Router(
    model_list=[{
        "model_name": "azure-gpt-35",
        "litellm_params": {
            "model": "azure/deployment-1",
            "api_key": "<key>",
            "api_base": "https://example.openai.azure.com/",
        },
    }],
    allowed_fails=1,          # Immediate cooldown on first failure

    cooldown_time=90,         # 90-second exclusion from rotation

)

On failure, the system stores a CooldownCacheValue with masked error details and a 90-second TTL, automatically reintegrating the deployment once healthy.

Summary

  • Per-deployment protection: The ModelRateLimitingCheck class enforces TPM/RPM limits using minute-scoped cache keys before requests leave the proxy, preventing individual providers from receiving excessive traffic.
  • Global capacity management: The dynamic rate limiter v3 implements three-phase atomic checks with priority reservations, ensuring high-priority workloads receive guaranteed capacity during saturation.
  • Failure isolation: The CooldownCache automatically removes failing deployments using masked error tracking and configurable TTLs, preventing cascade failures across the model pool.
  • Flexible configuration: All mechanisms support granular tuning via Router constructor parameters or litellm_settings in YAML configuration files.

Frequently Asked Questions

How does LiteLLM track rate limits for individual model deployments?

LiteLLM tracks rate limits using scoped cache keys in the format model_id:deployment_name:tpm:HH-MM or model_id:deployment_name:rpm:HH-MM. TPM counters are read locally before incrementing, while RPM counters use DualCache.increment_cache for atomic updates. The system checks these counters in the pre_call_check method of ModelRateLimitingCheck before forwarding requests to providers.

What happens when a deployment exceeds its configured rate limits?

When limits are exceeded, LiteLLM raises a RateLimitError with HTTP status 429 immediately during the pre-call phase. For TPM violations, this occurs after reading the local cache (lines 112-122 in model_rate_limit_check.py). For RPM violations, the error triggers after the atomic increment reveals the count exceeds the threshold (lines 132-140). The request never reaches the upstream provider.

How does LiteLLM prioritize requests when the system approaches capacity?

The dynamic rate limiter v3 calculates a saturation ratio by querying current usage counters. When this ratio exceeds saturation_threshold (e.g., 0.8), the system enforces priority reservations configured in priority_reservation. High-priority requests proceed if capacity exists within their reserved percentage, while lower-priority requests may be rate-limited. The three-phase atomic increment ensures accurate capacity tracking without phantom usage.

Can I customize how long a deployment stays in cooldown after failures?

Yes, configure the cooldown_time parameter in the Router constructor (in seconds) or set the LITELLM_COOLDOWN_TIME environment variable. The default is 60 seconds (DEFAULT_COOLDOWN_TIME_SECONDS). You can also adjust allowed_fails to control how many consecutive errors trigger cooldown. The CooldownCache stores entries with your specified TTL and automatically expires them, reintegrating the deployment into the routing pool.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →