How FreeLLMAPI Manages Rate Limiting: A Deep Dive into Multi-Layer Quota Controls
FreeLLMAPI enforces rate limits through a hierarchical three-layer architecture that combines per-key sliding windows, provider-wide caps, and intelligent cooldown mechanisms, all centralized in server/src/services/ratelimit.ts.
The tashfeenahmed/freellmapi repository implements a sophisticated traffic management system designed to maximize free-tier LLM provider utilization without hitting hard quotas. Unlike simple token bucket algorithms, this implementation tracks requests per minute (rpm), tokens per minute (tpm), requests per day (rpd), and tokens per day (tpd) across multiple dimensions while maintaining ACID-like persistence through SQLite.
Three-Layer Rate Limiting Architecture
Per-Key, Per-Model Windows
At the granular level, FreeLLMAPI tracks consumption through sliding-window counters defined in server/src/services/ratelimit.ts. The functions requestCount() and tokenCount() query the SQLite rate_limit_usage table when available, falling back to in-memory windows (memoryRequestCount, memoryTokenCount) during database unavailability.
The gate functions canMakeRequest() and canUseTokens() execute the core admission logic. These functions combine persisted counts, active in-flight leases, and configured limits to determine whether a new call is permitted. When checking a request, the system validates against both minute-level (rpm/tpm) and day-level (rpd/tpd) constraints simultaneously.
Provider-Wide Daily Caps
FreeLLMAPI aggregates consumption across all models for a given provider key using buildWindowUsageSnapshot(). The canUseProvider() function enforces hard daily limits by comparing the aggregated providerDailyRequestCount against environment-specific caps.
Operators configure these limits through environment variables: PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> overrides the default values defined in DEFAULT_PROVIDER_DAILY_REQUEST_CAPS. This allows the system to adapt to fluctuating free-tier quotas from providers like OpenRouter (approximately 1000 requests/day) without code deployments.
Provider-Wide Minute Caps
Similar to daily enforcement, canUseProviderMinute() applies per-minute request limits using providerMinuteRequestCount. This prevents burst traffic from exhausting minute-level quotas, particularly relevant for providers like NVIDIA NIM that enforce approximately 40 requests per minute.
Cooldown Handling for Rate-Limit Errors
When upstream providers return 429 errors or quota-exceeded signals, the getCooldownDecisionForLimit() function calculates temporary backoff periods. The decision engine considers whether the quota is known, historical hit counts, and any Retry-After header values.
Local endpoints (loopback URLs) receive a fixed short cooldown (LOCAL_ENDPOINT_COOLDOWN_MS), while external providers follow escalating cooldown durations (COOLDOWN_DURATIONS). The system applies extended cooldowns when daily limits exhaust or when heuristic thresholds (two hits within one hour) trigger. These cooldowns persist in the rate_limit_cooldowns table and expire automatically via clearCooldown().
Core Rate Limiting Mechanisms
Lease-Based Concurrency Control
Before dispatching any request, acquireLease() creates an in-memory lease representing the anticipated resource consumption. This lease gets counted against both minute and day windows immediately, preventing race conditions between the check and act phases. The lease ensures that concurrent requests respect quota limits even before the database records the actual usage.
Persistent SQLite Storage
All successful request and token recordings flow through recordUsage(), which writes to the rate_limit_usage table. If the database write fails, the system gracefully degrades to in-memory counters, ensuring availability during SQLite outages while potentially sacrificing accuracy until recovery.
Snapshot Caching for Performance
To minimize database query overhead, FreeLLMAPI maintains a five-second TTL snapshot (windowUsageSnapshot) that aggregates per-key usage statistics. This caching layer allows canMakeRequest() and canUseTokens() to perform fast guard-rail calculations without executing multiple tiny queries for every incoming request.
Configuration and Environment Variables
FreeLLMAPI exposes provider-specific caps through environment variables, enabling runtime configuration of rate limiting behavior:
PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>: Sets the maximum daily requests for a provider accountPROVIDER_MINUTE_REQUEST_CAP_<PLATFORM>: Configures per-minute request thresholds
These variables interface directly with the limit-checking logic in server/src/services/ratelimit.ts (lines 61-68 and 95-103), allowing operators to adjust to evolving free-tier quotas without restarting the service or modifying source code.
Implementation Examples
Checking Request Permissions
import { canMakeRequest, canUseTokens } from './services/ratelimit.js';
// Verify if a key can issue another request for a specific model
const canCall = canMakeRequest('openrouter', 'gpt-4o-mini', 42, {
rpm: 60,
rpd: 1000,
tpm: null,
tpd: null,
});
if (canCall) {
// Safe to dispatch the request
}
Validating Token Budgets
// Check token budget before sending large prompts
const canSpend = canUseTokens('openrouter', 'gpt-4o-mini', 42, 5000, {
tpm: 200_000,
tpd: null,
});
if (canSpend) {
// Proceed with the API call
}
Provider-Wide Quota Verification
import { canUseProvider } from './services/ratelimit.js';
// Check if the entire provider account has remaining budget
const accountOk = canUseProvider('openrouter', 42);
if (!accountOk) {
// Skip all OpenRouter models for this key today
}
Applying Post-Error Cooldowns
import { setCooldown } from './services/ratelimit.js';
// Apply a 2-minute cooldown after receiving a 429 response
setCooldown('openrouter', 'gpt-4o-mini', 42, 2 * 60_000, 'heuristic');
Summary
- Multi-layer enforcement: FreeLLMAPI applies per-key/model windows, provider-wide daily caps, and provider-wide minute caps through
server/src/services/ratelimit.ts. - Race-condition prevention: The
acquireLease()function creates in-memory leases before request dispatch to prevent quota overruns during concurrent access. - Persistent tracking: SQLite tables
rate_limit_usageandrate_limit_cooldownsmaintain state across restarts, with automatic fallback to memory during database failures. - Intelligent backoff: The
getCooldownDecisionForLimit()function implements escalating delays based on error types, headers, and historical patterns. - Environment configuration: Provider caps are tunable via
PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>andPROVIDER_MINUTE_REQUEST_CAP_<PLATFORM>variables.
Frequently Asked Questions
How does FreeLLMAPI prevent concurrent requests from exceeding rate limits?
The system implements lease-based concurrency control through the acquireLease() function in server/src/services/ratelimit.ts. Before dispatching any request, FreeLLMAPI reserves the expected capacity in sliding-window counters, ensuring that simultaneous checks respect the same quota limits without race conditions.
What happens to rate limit data if the SQLite database becomes unavailable?
FreeLLMAPI degrades gracefully to in-memory windows (memoryRequestCount, memoryTokenCount) when the rate_limit_usage table is inaccessible. The recordUsage() function handles write failures by maintaining counters in memory, preserving basic protection at the cost of persistence until database connectivity restores.
How does the cooldown system handle different types of rate limit errors?
The getCooldownDecisionForLimit() function analyzes error context to determine backoff durations. Local endpoints receive fixed short timeouts (LOCAL_ENDPOINT_COOLDOWN_MS), while external providers follow escalating cooldown patterns (COOLDOWN_DURATIONS). When providers include Retry-After headers, the system respects these values; otherwise, it applies heuristics based on daily limit exhaustion or hourly hit thresholds.
Can operators adjust rate limits without modifying code?
Yes. FreeLLMAPI reads provider-wide caps from environment variables (PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> and PROVIDER_MINUTE_REQUEST_CAP_<PLATFORM>) at runtime. These variables override default constants in server/src/services/ratelimit.ts, allowing immediate adaptation to changing free-tier quotas from providers like OpenRouter or NVIDIA NIM.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →