# How FreeLLMAPI Manages Rate Limiting: A Deep Dive into Multi-Layer Quota Controls

> Discover how FreeLLMAPI manages rate limiting with a three-layer quota system. Learn about per-key windows, provider caps, and cooldowns to optimize API usage.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-09-02

---

**FreeLLMAPI enforces rate limits through a hierarchical three-layer architecture that combines per-key sliding windows, provider-wide caps, and intelligent cooldown mechanisms, all centralized in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts).**

The `tashfeenahmed/freellmapi` repository implements a sophisticated traffic management system designed to maximize free-tier LLM provider utilization without hitting hard quotas. Unlike simple token bucket algorithms, this implementation tracks **requests per minute (rpm)**, **tokens per minute (tpm)**, **requests per day (rpd)**, and **tokens per day (tpd)** across multiple dimensions while maintaining ACID-like persistence through SQLite.

## Three-Layer Rate Limiting Architecture

### Per-Key, Per-Model Windows

At the granular level, FreeLLMAPI tracks consumption through sliding-window counters defined in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts). The functions `requestCount()` and `tokenCount()` query the SQLite `rate_limit_usage` table when available, falling back to in-memory windows (`memoryRequestCount`, `memoryTokenCount`) during database unavailability.

The gate functions `canMakeRequest()` and `canUseTokens()` execute the core admission logic. These functions combine persisted counts, active in-flight leases, and configured limits to determine whether a new call is permitted. When checking a request, the system validates against both minute-level (rpm/tpm) and day-level (rpd/tpd) constraints simultaneously.

### Provider-Wide Daily Caps

FreeLLMAPI aggregates consumption across all models for a given provider key using `buildWindowUsageSnapshot()`. The `canUseProvider()` function enforces hard daily limits by comparing the aggregated `providerDailyRequestCount` against environment-specific caps.

Operators configure these limits through environment variables: `PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>` overrides the default values defined in `DEFAULT_PROVIDER_DAILY_REQUEST_CAPS`. This allows the system to adapt to fluctuating free-tier quotas from providers like OpenRouter (approximately 1000 requests/day) without code deployments.

### Provider-Wide Minute Caps

Similar to daily enforcement, `canUseProviderMinute()` applies per-minute request limits using `providerMinuteRequestCount`. This prevents burst traffic from exhausting minute-level quotas, particularly relevant for providers like NVIDIA NIM that enforce approximately 40 requests per minute.

### Cooldown Handling for Rate-Limit Errors

When upstream providers return 429 errors or quota-exceeded signals, the `getCooldownDecisionForLimit()` function calculates temporary backoff periods. The decision engine considers whether the quota is known, historical hit counts, and any `Retry-After` header values.

Local endpoints (loopback URLs) receive a fixed short cooldown (`LOCAL_ENDPOINT_COOLDOWN_MS`), while external providers follow escalating cooldown durations (`COOLDOWN_DURATIONS`). The system applies extended cooldowns when daily limits exhaust or when heuristic thresholds (two hits within one hour) trigger. These cooldowns persist in the `rate_limit_cooldowns` table and expire automatically via `clearCooldown()`.

## Core Rate Limiting Mechanisms

### Lease-Based Concurrency Control

Before dispatching any request, `acquireLease()` creates an in-memory lease representing the anticipated resource consumption. This lease gets counted against both minute and day windows immediately, preventing race conditions between the check and act phases. The lease ensures that concurrent requests respect quota limits even before the database records the actual usage.

### Persistent SQLite Storage

All successful request and token recordings flow through `recordUsage()`, which writes to the `rate_limit_usage` table. If the database write fails, the system gracefully degrades to in-memory counters, ensuring availability during SQLite outages while potentially sacrificing accuracy until recovery.

### Snapshot Caching for Performance

To minimize database query overhead, FreeLLMAPI maintains a five-second TTL snapshot (`windowUsageSnapshot`) that aggregates per-key usage statistics. This caching layer allows `canMakeRequest()` and `canUseTokens()` to perform fast guard-rail calculations without executing multiple tiny queries for every incoming request.

## Configuration and Environment Variables

FreeLLMAPI exposes provider-specific caps through environment variables, enabling runtime configuration of rate limiting behavior:

- `PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>`: Sets the maximum daily requests for a provider account
- `PROVIDER_MINUTE_REQUEST_CAP_<PLATFORM>`: Configures per-minute request thresholds

These variables interface directly with the limit-checking logic in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) (lines 61-68 and 95-103), allowing operators to adjust to evolving free-tier quotas without restarting the service or modifying source code.

## Implementation Examples

### Checking Request Permissions

```typescript
import { canMakeRequest, canUseTokens } from './services/ratelimit.js';

// Verify if a key can issue another request for a specific model
const canCall = canMakeRequest('openrouter', 'gpt-4o-mini', 42, {
  rpm: 60,
  rpd: 1000,
  tpm: null,
  tpd: null,
});

if (canCall) {
  // Safe to dispatch the request
}

```

### Validating Token Budgets

```typescript
// Check token budget before sending large prompts
const canSpend = canUseTokens('openrouter', 'gpt-4o-mini', 42, 5000, {
  tpm: 200_000,
  tpd: null,
});

if (canSpend) {
  // Proceed with the API call
}

```

### Provider-Wide Quota Verification

```typescript
import { canUseProvider } from './services/ratelimit.js';

// Check if the entire provider account has remaining budget
const accountOk = canUseProvider('openrouter', 42);

if (!accountOk) {
  // Skip all OpenRouter models for this key today
}

```

### Applying Post-Error Cooldowns

```typescript
import { setCooldown } from './services/ratelimit.js';

// Apply a 2-minute cooldown after receiving a 429 response
setCooldown('openrouter', 'gpt-4o-mini', 42, 2 * 60_000, 'heuristic');

```

## Summary

- **Multi-layer enforcement**: FreeLLMAPI applies per-key/model windows, provider-wide daily caps, and provider-wide minute caps through [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts).
- **Race-condition prevention**: The `acquireLease()` function creates in-memory leases before request dispatch to prevent quota overruns during concurrent access.
- **Persistent tracking**: SQLite tables `rate_limit_usage` and `rate_limit_cooldowns` maintain state across restarts, with automatic fallback to memory during database failures.
- **Intelligent backoff**: The `getCooldownDecisionForLimit()` function implements escalating delays based on error types, headers, and historical patterns.
- **Environment configuration**: Provider caps are tunable via `PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>` and `PROVIDER_MINUTE_REQUEST_CAP_<PLATFORM>` variables.

## Frequently Asked Questions

### How does FreeLLMAPI prevent concurrent requests from exceeding rate limits?

The system implements **lease-based concurrency control** through the `acquireLease()` function in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts). Before dispatching any request, FreeLLMAPI reserves the expected capacity in sliding-window counters, ensuring that simultaneous checks respect the same quota limits without race conditions.

### What happens to rate limit data if the SQLite database becomes unavailable?

FreeLLMAPI degrades gracefully to **in-memory windows** (`memoryRequestCount`, `memoryTokenCount`) when the `rate_limit_usage` table is inaccessible. The `recordUsage()` function handles write failures by maintaining counters in memory, preserving basic protection at the cost of persistence until database connectivity restores.

### How does the cooldown system handle different types of rate limit errors?

The `getCooldownDecisionForLimit()` function analyzes error context to determine backoff durations. Local endpoints receive fixed short timeouts (`LOCAL_ENDPOINT_COOLDOWN_MS`), while external providers follow **escalating cooldown patterns** (`COOLDOWN_DURATIONS`). When providers include `Retry-After` headers, the system respects these values; otherwise, it applies heuristics based on daily limit exhaustion or hourly hit thresholds.

### Can operators adjust rate limits without modifying code?

Yes. FreeLLMAPI reads provider-wide caps from environment variables (`PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>` and `PROVIDER_MINUTE_REQUEST_CAP_<PLATFORM>`) at runtime. These variables override default constants in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts), allowing immediate adaptation to changing free-tier quotas from providers like OpenRouter or NVIDIA NIM.