How the FreeLLMAPI Rate Limiting System Tracks Per-Key Usage Across Providers

The rate limiting system in FreeLLMAPI uses a three-layer architecture—in-memory sliding windows, SQLite persistence, and provider-wide caps—to track per-key usage across different AI providers using composite keys formatted as platform:modelId:keyId:type.

The FreeLLMAPI project routes requests to multiple large language model providers while enforcing strict quota limits. Understanding how the system tracks per-key usage across providers is essential for operators managing API keys with varying rate limits. The implementation lives primarily in server/src/services/ratelimit.ts and coordinates durable storage, fast memory caches, and lease-based concurrency control.

Three-Layer Tracking Architecture

The rate limiting subsystem coordinates three distinct layers to ensure both performance and durability:

  • In-memory sliding windows – A fast, per-process cache storing recent request timestamps and token counts for each platform:modelId:keyId tuple. This layer provides microsecond-level lookups via the windows map and Window interface, tracked by memoryRequestCount and memoryTokenCount.

  • SQLite persistence – Durable counters that survive process restarts and provide the authoritative source for all windows. The recordUsage function writes to this layer, while countPersistedRequests and sumPersistedTokens aggregate historical data.

  • Provider-wide caps – Optional daily or per-minute limits that apply to entire provider accounts (such as OpenRouter’s free tier). These are managed by getProviderDailyRequestCap, getProviderMinuteRequestCap, and enforced through canUseProvider and canUseProviderMinute.

Per-Key Identification and Counter Logic

Each API key is uniquely identified by a composite string following the pattern "platform:modelId:keyId:type", where type is one of four rate limit categories defined in lines 13-14: rpm (requests per minute), rpd (requests per day), tpm (tokens per minute), or tpd (tokens per day).

Request and Token Counting

The system maintains separate counters for requests and tokens using a sliding window algorithm. For request counting, the requestCount function (lines 93-97) first queries the SQLite database via countPersistedRequests, falling back to memoryRequestCount if the database is unavailable. Similarly, token counting uses tokenCount (lines 100-108) with sumPersistedTokens and memoryTokenCount.

Both implementations use sliding windows—one minute for rpm/tpm and one day for rpd/tpd. The helper pruneTimestamps (lines 27-30) automatically removes entries older than the window boundary to prevent memory bloat.

// Check request limits (rpm/rpd) before routing
if (!canMakeRequest('openrouter', 'gpt-4', keyId, { rpm: 20, rpd: 1000, tpm: null, tpd: null })) {
  // Router skips this key for the current request
}

// Check token budget before sending a large prompt
const tokensNeeded = 350;
if (!canUseTokens('openrouter', 'gpt-4', keyId, tokensNeeded, { tpm: 2000, tpd: null })) {
  // Try alternative key or provider
}

Preventing Race Conditions with Leases

To handle check-then-act race conditions during concurrent requests, the system implements a lease mechanism. When the router prepares to dispatch a request, it calls acquireLease (lines 26-34), which stores a provisional reservation in the process-local leases map.

The lease count is included in limit calculations via inFlightForKey and provisional* helpers (lines 58-71). This ensures that canMakeRequest and canUseTokens (lines 21-25 and 42-45) account for in-flight requests before they hit the database, preventing quota overshoot during high-concurrency bursts.

Enforcing Hard Limits

Per-Key Gates

The canMakeRequest function (lines 23-30) enforces per-key request limits by checking both rpm and rpd thresholds against the sliding window counters. Similarly, canUseTokens (lines 45-53) validates tpm and tpd limits against current token usage. If any limit is exceeded, these functions return false, triggering the router to select an alternative key.

Provider-Wide Daily and Minute Caps

Some providers enforce account-level quotas rather than per-key limits. The system supports this through DEFAULT_PROVIDER_DAILY_REQUEST_CAPS and DEFAULT_PROVIDER_MINUTE_REQUEST_CAPS (lines 61-84), which operators can override via environment variables like PROVIDER_DAILY_REQUEST_CAP_<PLATFORM>.

The canUseProvider (lines 56-60) and canUseProviderMinute (lines 80-84) functions aggregate per-key usage across all keys for that provider, blocking new requests when the collective total approaches the provider-wide cap. This prevents a single provider account from exhausting its quota across multiple keys.

Routing Intelligence with Usage Snapshots

For load-balancing decisions, the router consumes a memoized snapshot of all per-key usage built by modelWindowUsedFraction (lines 18-48). This function queries windowUsageSnapshot from the database in a single grouped operation, calculating the fraction of each window already consumed.

The router in server/src/services/router.ts uses this data to weight routing decisions, steering traffic away from keys nearing their limits without affecting the hard enforcement gates. This separation allows for intelligent traffic shaping while maintaining strict quota adherence.

Summary

  • Composite keys track usage via the format platform:modelId:keyId:type with four rate limit categories (rpm, rpd, tpm, tpd).
  • Dual-layer storage combines fast in-memory sliding windows with durable SQLite persistence for authoritative counts.
  • Lease-based concurrency control prevents race conditions by accounting for in-flight requests before database writes occur.
  • Provider-wide caps protect against exhausting platform-level quotas using environment-variable-configured limits.
  • Usage snapshots enable the router to make informed load-balancing decisions based on real-time window consumption.

Frequently Asked Questions

How does the system handle database downtime when tracking usage?

When SQLite is unavailable, the rate limiting system falls back to in-memory counters via memoryRequestCount and memoryTokenCount. While this sacrifices durability across restarts, it ensures continuous request processing without hard failures.

What is the difference between rpm/rpd and tpm/tpd limit types?

The rpm (requests per minute) and rpd (requests per day) types track discrete API call counts, while tpm (tokens per minute) and tpd (tokens per day) aggregate the token counts from request payloads. All four use sliding windows but measure different resource dimensions.

How are provider-wide caps configured for custom deployments?

Operators set default caps in DEFAULT_PROVIDER_DAILY_REQUEST_CAPS and DEFAULT_PROVIDER_MINUTE_REQUEST_CAPS, or override them per-platform using environment variables following the pattern PROVIDER_DAILY_REQUEST_CAP_<PLATFORM> (e.g., PROVIDER_DAILY_REQUEST_CAP_OPENROUTER=1000).

Where does the routing logic consume the usage data?

The router imports modelWindowUsedFraction from server/src/services/ratelimit.ts to build a windowUsageSnapshot, which scores candidate keys during the selection phase in server/src/services/router.ts. This allows the system to prefer keys with higher remaining quota without violating hard limits.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →