How FreeLLMAPI Handles Rate Limiting Across Different LLM Providers
FreeLLMAPI enforces multi-layered rate limits at the provider, key, and model levels using sliding-window counters, provider-wide daily caps, and escalating cooldowns to ensure reliable LLM routing without exceeding upstream quotas.
FreeLLMAPI is an open-source LLM routing layer that aggregates multiple providers into a single API endpoint. To prevent individual API keys from exhausting their contractual limits while maintaining high availability, the system implements a sophisticated FreeLLMAPI rate limiting strategy that tracks usage across five distinct layers.
The Five-Layer Rate Limiting Architecture
FreeLLMAPI applies granular controls to every provider → key → model combination. According to the source code in server/src/services/ratelimit.ts, the system tracks both temporal request volumes and token consumption while respecting provider-specific constraints.
Per-Minute and Per-Day Request Caps (RPM/RPD)
The core rate limiting mechanism uses sliding-window counters stored in memory and persisted to SQLite. For each request, the router checks against configured rpm_limit and rpd_limit values:
- In-memory tracking: The
windowsmap maintains active counters for quick lookups - Database persistence: Usage history is stored in the
rate_limit_usagetable to survive restarts - Check function:
canMakeRequest(platform, model_id, keyId, limits)validates against both per-minute and per-day request quotas
Token-Based Limits (TPM/TPD)
Beyond raw request counts, FreeLLMAPI tracks token consumption (prompt + completion) to prevent quota exhaustion from high-token workloads:
import { canUseTokens, recordTokens } from './ratelimit';
// Before routing: validate token budget
const limits = {
rpm: entry.rpm_limit,
rpd: entry.rpd_limit,
tpm: entry.tpm_limit,
tpd: entry.tpd_limit,
};
if (!canUseTokens(entry.platform, entry.model_id, keyId, estimatedTokens, limits)) {
continue; // Skip this key, try next
}
// After successful response: record actual usage
recordTokens(platform, modelId, keyId, actualTokenCount);
Provider-Wide Daily Caps
Some providers (notably OpenRouter's free tier) enforce a single daily quota shared across all models on an account. FreeLLMAPI handles this via environment variables:
# Configure provider-wide daily request cap
PROVIDER_DAILY_REQUEST_CAP_OPENROUTER=50
PROVIDER_DAILY_REQUEST_CAP_OPENAI=1000
# Disable the cap entirely
PROVIDER_DAILY_REQUEST_CAP_OPENROUTER=0
The canUseProvider(platform, keyId) function in server/src/services/ratelimit.ts checks the providerDailyRequestCount aggregate, which sums usage across every model for that specific key.
Automatic Cooldown Management
When upstream providers return specific HTTP status codes, FreeLLMAPI temporarily bans the key/model combination to prevent hammering exhausted endpoints:
- Trigger codes: 429 (rate limit), 402 (payment required), 403 (model forbidden)
- Escalation schedule: 2 minutes → 10 minutes → 1 hour → 24 hours
- Persistence: Cooldowns are stored in the
rate_limit_cooldownstable and the in-memorycooldownsmap
Use isOnCooldown(platform, model_id, keyId) to check status before routing, and getNextCooldownDuration() to calculate the appropriate penalty based on recent throttle history.
Dynamic Quota Learning from Headers
FreeLLMAPI observes actual provider behavior to refine its limits. The server/src/services/provider-quota.ts module parses quota headers (e.g., x-ratelimit-limit-requests) via parseQuotaObservationsFromResponse(), storing findings in provider_quota_state and provider_quota_observations tables. This empirical data feeds back into the router's decision-making without manual configuration.
Routing Logic and Fallback Behavior
The router (server/src/services/router.ts) implements a validation chain inside selectKeyForModel that sequentially tests each available key:
import {
canMakeRequest,
canUseTokens,
canUseProvider,
isOnCooldown,
} from './ratelimit';
// Validation chain for each candidate key
if (isOnCooldown(entry.platform, entry.model_id, kid)) continue;
if (!canUseProvider(entry.platform, kid)) continue; // Provider-wide daily cap
if (!canMakeRequest(entry.platform, entry.model_id, kid, limits)) continue; // RPM/RPD
if (!canUseTokens(entry.platform, entry.model_id, kid, estimatedTokens, limits)) continue; // TPM/TPD
If a key fails any check, the router immediately tries the next available key. When all keys for a model are exhausted, FreeLLMAPI proceeds to the next model in the fallback chain. Only when the entire chain is depleted does it throw a RouteError (HTTP 429), signaling that the client must wait for limits to reset.
Configuring Rate Limits for Different Providers
FreeLLMAPI stores default provider caps in DEFAULT_PROVIDER_DAILY_REQUEST_CAPS but allows granular overrides:
- Environment variables: Set
PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>for provider-wide limits - Model-specific limits: Define
rpm_limit,rpd_limit,tpm_limit, andtpd_limitin the model configuration entries - SQLite persistence: The
rate_limit_usageandrate_limit_cooldownstables ensure state survives process restarts
After each successful request, update counters explicitly:
import { recordRequest } from './ratelimit';
// Update request counters
recordRequest(platform, modelId, keyId);
Handling Rate Limit Responses and Cooldowns
When receiving a 429 or 402 response, FreeLLMAPI automatically applies escalating cooldowns. The duration depends on how many times the key has been throttled in the last 24 hours:
import { getNextCooldownDuration } from './ratelimit';
const cooldownMs = getNextCooldownDuration(platform, modelId, keyId);
// Automatically applied by the router on error responses
This protects both the local instance and the upstream provider from retry storms while allowing temporary issues to resolve automatically.
Summary
- FreeLLMAPI rate limiting operates across five layers: RPM/RPD requests, TPM/TPD tokens, provider-wide daily caps, escalating cooldowns, and dynamic header observation.
- Core logic resides in
server/src/services/ratelimit.ts, with dynamic quota learning inserver/src/services/provider-quota.ts. - The router validates keys using
isOnCooldown,canUseProvider,canMakeRequest, andcanUseTokensbefore selecting an endpoint. - Failed keys trigger automatic fallback to alternate models; exhausted chains return HTTP 429.
- Provider-wide caps are configurable via
PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>environment variables. - Cooldowns escalate from 2 minutes to 24 hours based on 24-hour throttle history and persist in SQLite.
Frequently Asked Questions
How does FreeLLMAPI store rate limit data between restarts?
FreeLLMAPI persists sliding-window counters and cooldown states in SQLite tables (rate_limit_usage and rate_limit_cooldowns) while maintaining high-performance in-memory maps (windows and cooldowns) for active operations. This hybrid approach ensures durability without sacrificing lookup speed during request routing.
What happens when all API keys for a model hit their rate limits?
When all keys for a specific model fail the validation checks (cooldown, provider caps, or RPM/RPD/TPM/TPD limits), the router immediately proceeds to the next model in the fallback chain. If the entire chain is exhausted, FreeLLMAPI throws a RouteError with HTTP status 429, indicating that no capacity is currently available and the client should implement retry logic with exponential backoff.
How do I configure custom daily caps for OpenRouter or other providers?
Set the environment variable PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> where <PLATFORM> is the uppercase provider name (e.g., OPENROUTER, OPENAI). For example, PROVIDER_DAILY_REQUEST_CAP_OPENROUTER=50 limits that provider to 50 requests per UTC day across all models. Set the value to 0 to disable the provider-wide cap entirely.
Does FreeLLMAPI automatically learn provider rate limits from API responses?
Yes. The server/src/services/provider-quota.ts module parses quota headers (such as x-ratelimit-limit-requests and x-ratelimit-remaining) from provider responses and stores these observations in provider_quota_observations. This empirical data allows the system to adapt to actual provider limits dynamically, supplementing the static configuration with real-world usage patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →