How FreeLLMAPI Handles Rate Limits (RPM, RPD, TPM, TPD): Implementation Deep Dive

FreeLLMAPI enforces rate limits through a layered system that tracks per-model, per-key quotas using sliding-window counters stored in SQLite with in-memory caching, while also respecting provider-wide account caps through environment-configurable thresholds.

FreeLLMAPI protects both individual API keys and entire provider accounts from exceeding usage quotas through a sophisticated rate-limiting architecture. The system implements four distinct quota types—requests per minute (RPM), requests per day (RPD), tokens per minute (TPM), and tokens per day (TPD)—using sliding-window counters that persist to SQLite but fallback to memory when the database is unavailable. All rate-limit logic is centralized in server/src/services/ratelimit.ts and invoked by the routing layer before dispatching requests to upstream providers.

Understanding the Four Rate Limit Types

FreeLLMAPI tracks usage at the granularity of model + API key pairs, allowing fine-grained control over individual integrations.

RPM (Requests Per Minute)

The RPM quota applies a per-model, per-key limit using a sliding-minute window. The system stores counters in the rate_limit_usage SQLite table and maintains an in-memory cache for performance. The requestCount() function (lines 93-97) aggregates recent requests, falling back to memoryRequestCount() (lines 74-78) when the database is unreachable. Before dispatching any request, the canMakeRequest() function (lines 112-133) validates that the current count remains below the configured threshold.

RPD (Requests Per Day)

The RPD quota enforces daily limits based on a sliding-day window calculated from UTC midnight. Using the same rate_limit_usage table structure, the system distinguishes day-wide aggregations through the DAY constant (line 34). The canMakeRequest() function checks both minute and daily windows (lines 124-129), rejecting requests that would exceed the 24-hour allowance.

TPM (Tokens Per Minute)

For TPM tracking, FreeLLMAPI monitors token consumption through rate_limit_usage entries with kind = 'tokens'. The memoryTokenCount() function (lines 80-84) maintains hot counters in memory, while canUseTokens() (lines 134-155) validates requests by adding provisional token estimates (provisionalTokens(), lines 72-78) to the current window total before comparing against limits.tpm.

TPD (Tokens Per Day)

The TPD quota aggregates total tokens consumed per day through tokenCount() (lines 100-108), which queries day-wide sums from the database. The canUseTokens() function validates daily limits at lines 146-152, ensuring that batch operations or large completions do not exhaust the 24-hour token budget.

Sliding-Window Counter Implementation

FreeLLMAPI implements durable counters through SQLite with automatic pruning and memory fallback mechanisms.

Every request or token usage is persisted via recordUsage() (lines 42-48), which writes timestamped entries to rate_limit_usage. To prevent unbounded table growth, the system prunes expired rows every minute using USAGE_PRUNE_INTERVAL_MS (line 103), removing entries outside the current sliding windows.

When SQLite connectivity fails, the rate limiter transparently falls back to in-memory tracking through pushMemoryRequest() and pushMemoryTokens() (lines 80-87). These ephemeral counters ensure that rate limiting continues to function during database maintenance or network partitions, though they reset on process restart.

Provider-Wide Account Caps

Beyond per-key limits, FreeLLMAPI respects quotas that apply to entire provider accounts (such as OpenRouter's free tier limitations).

Configurable Thresholds

The system reads default caps from DEFAULT_PROVIDER_DAILY_REQUEST_CAPS (lines 61-66) and DEFAULT_PROVIDER_DAILY_TOKEN_CAPS (lines 69-74), with provider-specific overrides available through environment variables:

  • PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> – Overrides daily request limits (lines 86-92)
  • PROVIDER_DAILY_TOKEN_CAP_<PLATFORM> – Overrides daily token limits (lines 96-102)
  • PROVIDER_MINUTE_REQUEST_CAP_<PLATFORM> – Overrides per-minute request limits (lines 95-103)

Default values include OpenRouter (1000 requests/day), ModelScope (1800 requests/day), and NVIDIA (40 requests/minute).

Enforcement Functions

The canUseProvider(), canUseProviderMinute(), and canUseProviderTokens() functions (lines 56-61, 80-86, 74-86) query these caps against provider-level usage counters maintained by providerDailyRequestCount(), providerMinuteRequestCount(), and providerDailyTokenCount(). These checks gate routing decisions in server/src/services/provider-quota.ts to prevent exhausting shared provider budgets.

Preventing Race Conditions with In-Flight Leases

To eliminate race conditions where parallel requests simultaneously pass quota checks based on stale counter values, FreeLLMAPI implements a lease-based concurrency control system.

When the router selects a key for dispatch, it immediately calls acquireLease() (line 27-34) to register the pending request. The provisionalRequests() and provisionalTokens() functions (lines 72-78) include these in-flight operations in quota calculations within canMakeRequest() and canUseTokens().

This ensures that concurrent requests against the same model and key correctly account for each other's resource consumption before upstream API calls are initiated.

Rate Limit Enforcement Flow

The complete rate-limiting workflow follows this strict sequence:

  1. Limit Retrieval – Fetch the key's WindowLimits object containing rpm, rpd, tpm, and tpd values from configuration
  2. Lease Acquisition – Call acquireLease() to register the pending request and obtain a lease ID
  3. Hard Limit Validation – Execute canMakeRequest() to verify RPM/RPD compliance, then canUseTokens() to check TPM/TPD against estimated token usage
  4. Provider Cap Check – Validate against canUseProvider(), canUseProviderMinute(), and canUseProviderTokens() if provider-wide limits are configured
  5. Dispatch – If all checks pass, transmit the request to the upstream provider
  6. Usage Recording – On success, recordRequest() and recordTokens() persist actual consumption; on failure, releaseLease() frees the provisional allocation and triggers cooldown logic if applicable
// Example: Validating rate limits before provider dispatch
import { canMakeRequest, canUseTokens, acquireLease, recordRequest, recordTokens } from '@/services/ratelimit';

async function dispatchRequest(platform: string, modelId: string, keyId: number, estimatedTokens: number) {
  const limits = { rpm: 60, rpd: 1000, tpm: 200_000, tpd: 1_000_000 };
  
  // Check per-key limits
  if (!canMakeRequest(platform, modelId, keyId, limits)) {
    throw new Error('Rate limit exceeded: RPM or RPD quota depleted');
  }
  
  if (!canUseTokens(platform, modelId, keyId, estimatedTokens, limits)) {
    throw new Error('Token limit exceeded: TPM or TPD quota depleted');
  }
  
  // Acquire lease to prevent race conditions
  const leaseId = acquireLease(platform, modelId, keyId, estimatedTokens);
  
  try {
    const response = await providerApiCall(platform, modelId, keyId);
    
    // Record actual usage
    recordRequest(platform, modelId, keyId);
    recordTokens(platform, modelId, keyId, response.usage?.total_tokens ?? 0);
    
    return response;
  } finally {
    releaseLease(leaseId);
  }
}

Summary

  • Four quota types (RPM, RPD, TPM, TPD) are tracked per-model and per-key in server/src/services/ratelimit.ts using sliding-window counters
  • Dual-storage architecture persists counters to SQLite (rate_limit_usage table) with automatic pruning every USAGE_PRUNE_INTERVAL_MS, falling back to in-memory windows during database outages
  • Provider-wide caps enforce account-level limits configurable via PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> and related environment variables
  • Lease-based concurrency prevents race conditions by counting in-flight requests (provisionalRequests, provisionalTokens) against quota totals before dispatch
  • Validation functions canMakeRequest() and canUseTokens() gate all outbound traffic, checking both hard limits and provisional allocations

Frequently Asked Questions

How does FreeLLMAPI prevent rate limit errors when multiple requests arrive simultaneously?

FreeLLMAPI uses an in-flight lease system managed by acquireLease() and releaseLease() to track pending requests before they reach the provider. The canMakeRequest() and canUseTokens() functions include provisional counts from active leases in their calculations, ensuring that concurrent requests see each other's resource consumption and fail fast if quotas would be exceeded.

What happens to rate limiting if the SQLite database becomes unavailable?

The system implements a graceful degradation strategy through memory fallback functions (memoryRequestCount(), memoryTokenCount(), pushMemoryRequest(), pushMemoryTokens()). When SQLite writes fail, counters continue operating in ephemeral memory, though usage data is lost on process restart. This ensures rate limiting remains functional during database maintenance.

Where are provider-wide daily caps configured in FreeLLMAPI?

Provider-wide limits are defined in server/src/services/ratelimit.ts through the DEFAULT_PROVIDER_DAILY_REQUEST_CAPS and DEFAULT_PROVIDER_DAILY_TOKEN_CAPS maps (lines 61-74). Operators can override these defaults using environment variables following the pattern PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> (e.g., PROVIDER_DAILY_REQUEST_CAP_OPENROUTER=500).

How does FreeLLMAPI calculate the sliding window for daily limits?

Daily quotas (RPD and TPD) use a sliding-day window based on UTC midnight rather than a rolling 24-hour period. The system distinguishes day boundaries using the DAY constant (line 34) and aggregates usage through providerDailyRequestCount() and tokenCount() functions that filter entries by the current UTC date.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →