# How FreeLLMAPI Handles Rate Limits (RPM, RPD, TPM, TPD): Implementation Deep Dive

> Discover how FreeLLMAPI implements rate limits RPM RPD TPM TPD with a layered system, SQLite storage, in-memory caching, and environment-configurable thresholds. Optimize your API usage.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-09-04

---

**FreeLLMAPI enforces rate limits through a layered system that tracks per-model, per-key quotas using sliding-window counters stored in SQLite with in-memory caching, while also respecting provider-wide account caps through environment-configurable thresholds.**

FreeLLMAPI protects both individual API keys and entire provider accounts from exceeding usage quotas through a sophisticated rate-limiting architecture. The system implements four distinct quota types—requests per minute (RPM), requests per day (RPD), tokens per minute (TPM), and tokens per day (TPD)—using sliding-window counters that persist to SQLite but fallback to memory when the database is unavailable. All rate-limit logic is centralized in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) and invoked by the routing layer before dispatching requests to upstream providers.

## Understanding the Four Rate Limit Types

FreeLLMAPI tracks usage at the granularity of **model + API key pairs**, allowing fine-grained control over individual integrations.

### RPM (Requests Per Minute)

The **RPM** quota applies a per-model, per-key limit using a sliding-minute window. The system stores counters in the `rate_limit_usage` SQLite table and maintains an in-memory cache for performance. The `requestCount()` function (lines 93-97) aggregates recent requests, falling back to `memoryRequestCount()` (lines 74-78) when the database is unreachable. Before dispatching any request, the `canMakeRequest()` function (lines 112-133) validates that the current count remains below the configured threshold.

### RPD (Requests Per Day)

The **RPD** quota enforces daily limits based on a sliding-day window calculated from UTC midnight. Using the same `rate_limit_usage` table structure, the system distinguishes day-wide aggregations through the `DAY` constant (line 34). The `canMakeRequest()` function checks both minute and daily windows (lines 124-129), rejecting requests that would exceed the 24-hour allowance.

### TPM (Tokens Per Minute)

For **TPM** tracking, FreeLLMAPI monitors token consumption through `rate_limit_usage` entries with `kind = 'tokens'`. The `memoryTokenCount()` function (lines 80-84) maintains hot counters in memory, while `canUseTokens()` (lines 134-155) validates requests by adding **provisional token estimates** (`provisionalTokens()`, lines 72-78) to the current window total before comparing against `limits.tpm`.

### TPD (Tokens Per Day)

The **TPD** quota aggregates total tokens consumed per day through `tokenCount()` (lines 100-108), which queries day-wide sums from the database. The `canUseTokens()` function validates daily limits at lines 146-152, ensuring that batch operations or large completions do not exhaust the 24-hour token budget.

## Sliding-Window Counter Implementation

FreeLLMAPI implements durable counters through SQLite with automatic pruning and memory fallback mechanisms.

Every request or token usage is persisted via `recordUsage()` (lines 42-48), which writes timestamped entries to `rate_limit_usage`. To prevent unbounded table growth, the system prunes expired rows every minute using `USAGE_PRUNE_INTERVAL_MS` (line 103), removing entries outside the current sliding windows.

When SQLite connectivity fails, the rate limiter transparently falls back to in-memory tracking through `pushMemoryRequest()` and `pushMemoryTokens()` (lines 80-87). These ephemeral counters ensure that rate limiting continues to function during database maintenance or network partitions, though they reset on process restart.

## Provider-Wide Account Caps

Beyond per-key limits, FreeLLMAPI respects quotas that apply to entire provider accounts (such as OpenRouter's free tier limitations).

### Configurable Thresholds

The system reads default caps from `DEFAULT_PROVIDER_DAILY_REQUEST_CAPS` (lines 61-66) and `DEFAULT_PROVIDER_DAILY_TOKEN_CAPS` (lines 69-74), with provider-specific overrides available through environment variables:

- **`PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>`** – Overrides daily request limits (lines 86-92)
- **`PROVIDER_DAILY_TOKEN_CAP_<PLATFORM>`** – Overrides daily token limits (lines 96-102)  
- **`PROVIDER_MINUTE_REQUEST_CAP_<PLATFORM>`** – Overrides per-minute request limits (lines 95-103)

Default values include OpenRouter (1000 requests/day), ModelScope (1800 requests/day), and NVIDIA (40 requests/minute).

### Enforcement Functions

The `canUseProvider()`, `canUseProviderMinute()`, and `canUseProviderTokens()` functions (lines 56-61, 80-86, 74-86) query these caps against provider-level usage counters maintained by `providerDailyRequestCount()`, `providerMinuteRequestCount()`, and `providerDailyTokenCount()`. These checks gate routing decisions in [`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts) to prevent exhausting shared provider budgets.

## Preventing Race Conditions with In-Flight Leases

To eliminate race conditions where parallel requests simultaneously pass quota checks based on stale counter values, FreeLLMAPI implements a **lease-based concurrency control** system.

When the router selects a key for dispatch, it immediately calls `acquireLease()` (line 27-34) to register the pending request. The `provisionalRequests()` and `provisionalTokens()` functions (lines 72-78) include these in-flight operations in quota calculations within `canMakeRequest()` and `canUseTokens()`.

This ensures that concurrent requests against the same model and key correctly account for each other's resource consumption before upstream API calls are initiated.

## Rate Limit Enforcement Flow

The complete rate-limiting workflow follows this strict sequence:

1. **Limit Retrieval** – Fetch the key's `WindowLimits` object containing `rpm`, `rpd`, `tpm`, and `tpd` values from configuration
2. **Lease Acquisition** – Call `acquireLease()` to register the pending request and obtain a lease ID
3. **Hard Limit Validation** – Execute `canMakeRequest()` to verify RPM/RPD compliance, then `canUseTokens()` to check TPM/TPD against estimated token usage
4. **Provider Cap Check** – Validate against `canUseProvider()`, `canUseProviderMinute()`, and `canUseProviderTokens()` if provider-wide limits are configured
5. **Dispatch** – If all checks pass, transmit the request to the upstream provider
6. **Usage Recording** – On success, `recordRequest()` and `recordTokens()` persist actual consumption; on failure, `releaseLease()` frees the provisional allocation and triggers cooldown logic if applicable

```typescript
// Example: Validating rate limits before provider dispatch
import { canMakeRequest, canUseTokens, acquireLease, recordRequest, recordTokens } from '@/services/ratelimit';

async function dispatchRequest(platform: string, modelId: string, keyId: number, estimatedTokens: number) {
  const limits = { rpm: 60, rpd: 1000, tpm: 200_000, tpd: 1_000_000 };
  
  // Check per-key limits
  if (!canMakeRequest(platform, modelId, keyId, limits)) {
    throw new Error('Rate limit exceeded: RPM or RPD quota depleted');
  }
  
  if (!canUseTokens(platform, modelId, keyId, estimatedTokens, limits)) {
    throw new Error('Token limit exceeded: TPM or TPD quota depleted');
  }
  
  // Acquire lease to prevent race conditions
  const leaseId = acquireLease(platform, modelId, keyId, estimatedTokens);
  
  try {
    const response = await providerApiCall(platform, modelId, keyId);
    
    // Record actual usage
    recordRequest(platform, modelId, keyId);
    recordTokens(platform, modelId, keyId, response.usage?.total_tokens ?? 0);
    
    return response;
  } finally {
    releaseLease(leaseId);
  }
}

```

## Summary

- **Four quota types** (RPM, RPD, TPM, TPD) are tracked per-model and per-key in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) using sliding-window counters
- **Dual-storage architecture** persists counters to SQLite (`rate_limit_usage` table) with automatic pruning every `USAGE_PRUNE_INTERVAL_MS`, falling back to in-memory windows during database outages
- **Provider-wide caps** enforce account-level limits configurable via `PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>` and related environment variables
- **Lease-based concurrency** prevents race conditions by counting in-flight requests (`provisionalRequests`, `provisionalTokens`) against quota totals before dispatch
- **Validation functions** `canMakeRequest()` and `canUseTokens()` gate all outbound traffic, checking both hard limits and provisional allocations

## Frequently Asked Questions

### How does FreeLLMAPI prevent rate limit errors when multiple requests arrive simultaneously?

FreeLLMAPI uses an **in-flight lease system** managed by `acquireLease()` and `releaseLease()` to track pending requests before they reach the provider. The `canMakeRequest()` and `canUseTokens()` functions include provisional counts from active leases in their calculations, ensuring that concurrent requests see each other's resource consumption and fail fast if quotas would be exceeded.

### What happens to rate limiting if the SQLite database becomes unavailable?

The system implements a **graceful degradation** strategy through memory fallback functions (`memoryRequestCount()`, `memoryTokenCount()`, `pushMemoryRequest()`, `pushMemoryTokens()`). When SQLite writes fail, counters continue operating in ephemeral memory, though usage data is lost on process restart. This ensures rate limiting remains functional during database maintenance.

### Where are provider-wide daily caps configured in FreeLLMAPI?

Provider-wide limits are defined in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) through the `DEFAULT_PROVIDER_DAILY_REQUEST_CAPS` and `DEFAULT_PROVIDER_DAILY_TOKEN_CAPS` maps (lines 61-74). Operators can override these defaults using environment variables following the pattern `PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>` (e.g., `PROVIDER_DAILY_REQUEST_CAP_OPENROUTER=500`).

### How does FreeLLMAPI calculate the sliding window for daily limits?

Daily quotas (RPD and TPD) use a **sliding-day window** based on UTC midnight rather than a rolling 24-hour period. The system distinguishes day boundaries using the `DAY` constant (line 34) and aggregates usage through `providerDailyRequestCount()` and `tokenCount()` functions that filter entries by the current UTC date.