# How the FreeLLMAPI Rate Limiting System Tracks Per-Key Usage Across Providers

> Learn how FreeLLMAPI's rate limiting system tracks per-key usage across providers using its three-layer architecture. Understand composite keys and ensure efficient API access.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: internals
- Published: 2026-08-31

---

**The rate limiting system in FreeLLMAPI uses a three-layer architecture—in-memory sliding windows, SQLite persistence, and provider-wide caps—to track per-key usage across different AI providers using composite keys formatted as `platform:modelId:keyId:type`.**

The FreeLLMAPI project routes requests to multiple large language model providers while enforcing strict quota limits. Understanding how the system tracks per-key usage across providers is essential for operators managing API keys with varying rate limits. The implementation lives primarily in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) and coordinates durable storage, fast memory caches, and lease-based concurrency control.

## Three-Layer Tracking Architecture

The rate limiting subsystem coordinates three distinct layers to ensure both performance and durability:

- **In-memory sliding windows** – A fast, per-process cache storing recent request timestamps and token counts for each `platform:modelId:keyId` tuple. This layer provides microsecond-level lookups via the `windows` map and `Window` interface, tracked by `memoryRequestCount` and `memoryTokenCount`.

- **SQLite persistence** – Durable counters that survive process restarts and provide the authoritative source for all windows. The `recordUsage` function writes to this layer, while `countPersistedRequests` and `sumPersistedTokens` aggregate historical data.

- **Provider-wide caps** – Optional daily or per-minute limits that apply to entire provider accounts (such as OpenRouter’s free tier). These are managed by `getProviderDailyRequestCap`, `getProviderMinuteRequestCap`, and enforced through `canUseProvider` and `canUseProviderMinute`.

## Per-Key Identification and Counter Logic

Each API key is uniquely identified by a composite string following the pattern `"platform:modelId:keyId:type"`, where `type` is one of four rate limit categories defined in lines 13-14: `rpm` (requests per minute), `rpd` (requests per day), `tpm` (tokens per minute), or `tpd` (tokens per day).

### Request and Token Counting

The system maintains separate counters for requests and tokens using a sliding window algorithm. For **request counting**, the `requestCount` function (lines 93-97) first queries the SQLite database via `countPersistedRequests`, falling back to `memoryRequestCount` if the database is unavailable. Similarly, **token counting** uses `tokenCount` (lines 100-108) with `sumPersistedTokens` and `memoryTokenCount`.

Both implementations use sliding windows—one minute for `rpm`/`tpm` and one day for `rpd`/`tpd`. The helper `pruneTimestamps` (lines 27-30) automatically removes entries older than the window boundary to prevent memory bloat.

```typescript
// Check request limits (rpm/rpd) before routing
if (!canMakeRequest('openrouter', 'gpt-4', keyId, { rpm: 20, rpd: 1000, tpm: null, tpd: null })) {
  // Router skips this key for the current request
}

// Check token budget before sending a large prompt
const tokensNeeded = 350;
if (!canUseTokens('openrouter', 'gpt-4', keyId, tokensNeeded, { tpm: 2000, tpd: null })) {
  // Try alternative key or provider
}

```

## Preventing Race Conditions with Leases

To handle check-then-act race conditions during concurrent requests, the system implements a **lease mechanism**. When the router prepares to dispatch a request, it calls `acquireLease` (lines 26-34), which stores a provisional reservation in the process-local `leases` map.

The lease count is included in limit calculations via `inFlightForKey` and `provisional*` helpers (lines 58-71). This ensures that `canMakeRequest` and `canUseTokens` (lines 21-25 and 42-45) account for in-flight requests before they hit the database, preventing quota overshoot during high-concurrency bursts.

## Enforcing Hard Limits

### Per-Key Gates

The `canMakeRequest` function (lines 23-30) enforces per-key request limits by checking both `rpm` and `rpd` thresholds against the sliding window counters. Similarly, `canUseTokens` (lines 45-53) validates `tpm` and `tpd` limits against current token usage. If any limit is exceeded, these functions return `false`, triggering the router to select an alternative key.

### Provider-Wide Daily and Minute Caps

Some providers enforce account-level quotas rather than per-key limits. The system supports this through `DEFAULT_PROVIDER_DAILY_REQUEST_CAPS` and `DEFAULT_PROVIDER_MINUTE_REQUEST_CAPS` (lines 61-84), which operators can override via environment variables like `PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>`.

The `canUseProvider` (lines 56-60) and `canUseProviderMinute` (lines 80-84) functions aggregate per-key usage across all keys for that provider, blocking new requests when the collective total approaches the provider-wide cap. This prevents a single provider account from exhausting its quota across multiple keys.

## Routing Intelligence with Usage Snapshots

For load-balancing decisions, the router consumes a memoized snapshot of all per-key usage built by `modelWindowUsedFraction` (lines 18-48). This function queries `windowUsageSnapshot` from the database in a single grouped operation, calculating the fraction of each window already consumed.

The router in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) uses this data to weight routing decisions, steering traffic away from keys nearing their limits without affecting the hard enforcement gates. This separation allows for intelligent traffic shaping while maintaining strict quota adherence.

## Summary

- **Composite keys** track usage via the format `platform:modelId:keyId:type` with four rate limit categories (rpm, rpd, tpm, tpd).
- **Dual-layer storage** combines fast in-memory sliding windows with durable SQLite persistence for authoritative counts.
- **Lease-based concurrency control** prevents race conditions by accounting for in-flight requests before database writes occur.
- **Provider-wide caps** protect against exhausting platform-level quotas using environment-variable-configured limits.
- **Usage snapshots** enable the router to make informed load-balancing decisions based on real-time window consumption.

## Frequently Asked Questions

### How does the system handle database downtime when tracking usage?

When SQLite is unavailable, the rate limiting system falls back to in-memory counters via `memoryRequestCount` and `memoryTokenCount`. While this sacrifices durability across restarts, it ensures continuous request processing without hard failures.

### What is the difference between rpm/rpd and tpm/tpd limit types?

The `rpm` (requests per minute) and `rpd` (requests per day) types track discrete API call counts, while `tpm` (tokens per minute) and `tpd` (tokens per day) aggregate the token counts from request payloads. All four use sliding windows but measure different resource dimensions.

### How are provider-wide caps configured for custom deployments?

Operators set default caps in `DEFAULT_PROVIDER_DAILY_REQUEST_CAPS` and `DEFAULT_PROVIDER_MINUTE_REQUEST_CAPS`, or override them per-platform using environment variables following the pattern `PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>` (e.g., `PROVIDER_DAILY_REQUEST_CAP_OPENROUTER=1000`).

### Where does the routing logic consume the usage data?

The router imports `modelWindowUsedFraction` from [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) to build a `windowUsageSnapshot`, which scores candidate keys during the selection phase in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts). This allows the system to prefer keys with higher remaining quota without violating hard limits.