How FreeLLMAPI Tracks Per-Key Rate Limits: SQLite + Sliding-Window Architecture Explained
FreeLLMAPI enforces per-key rate limits using a hybrid architecture: an embedded SQLite database for durable usage records and an in-memory sliding-window cache for fast real-time checks, with each key uniquely identified by a platform:modelId:keyId triple.
The open-source FreeLLMAPI project implements sophisticated rate limiting at the individual API key level. According to the tashfeenahmed/freellmapi source code, the system combines database persistence with memory caching to track requests, tokens, and concurrency across multiple time windows while maintaining low latency for routing decisions.
Per-Key Identification Strategy
FreeLLMAPI identifies each rate-limited entity using a composite key format.
In server/src/services/ratelimit.ts, keys are structured as platform:modelId:keyId — for example, openai:gpt-4o:key_abc123. This triple ensures complete isolation between different providers, models, and individual key credentials (see the key format comment at lines 13-14).
The usageCacheKey function (lines 404-405) builds cache keys from this triple, guaranteeing that usage statistics for one key never bleed into another.
Dual-Layer Storage: Database + Memory Cache
The rate limiting system operates across two storage layers:
- SQLite database (
getDb()) — persists all usage records for durability and historical analysis - In-memory sliding window — provides sub-millisecond lookups for hot-path routing decisions
Every recorded request is written to both layers simultaneously. The memory cache maintains recent request timestamps in memoryRequestCount and token tallies in memoryTokenCount, while the database stores the authoritative record via countPersisted* queries.
Core Rate Limit Checks
The ratelimit service exposes three primary enforcement functions that the router consults before assigning a key to a request.
Daily and Per-Minute Request Caps
-
canUseProvider()(lines 639-647) — validates againstproviderDailyRequestCount, combining persisted DB counts with in-memory window data plus anyprovisionalProviderRequests(in-flight requests not yet committed) -
canUseProviderMinute()(lines 664-670) — enforcesproviderMinuteRequestCountusing the same dual-source calculation
Both functions return boolean availability status, allowing the router to immediately disqualify exhausted keys.
Concurrency Limits
canUseKeyConcurrency()— prevents a single key from accumulating too many simultaneous leases
The system tracks inFlightForKey (lines 109-120) to limit parallel execution. When a key reaches its concurrency cap, new requests queue or route to alternative keys.
Recording Usage: The recordRequest() and recordTokens() Pipeline
After a request completes, the system commits usage through two dedicated functions:
// server/src/services/ratelimit.ts
// Lines 812-820 and 832-834
import { recordRequest, recordTokens } from './services/ratelimit.js';
// Increment request counter by 1
recordRequest('openai', modelId, keyId);
// Add token consumption to the ledger
recordTokens('openai', modelId, keyId, 150);
Both functions execute pushMemoryRequest to update the sliding window and execute INSERT statements against SQLite. This design ensures the in-memory view stays synchronized with the durable store.
Cooldown Enforcement and Querying
When a key exhausts its limits, FreeLLMAPI applies a cooldown penalty rather than permanent blacklisting.
The setCooldown function registers a temporary ban. The router later queries getNextCooldownDuration (lines 877-878) to determine remaining penalty time:
import { getNextCooldownDuration } from './services/ratelimit.js';
const msRemaining = getNextCooldownDuration('openai', modelId, keyId);
if (msRemaining > 0) {
// Key is rate-limited; router will skip or deprioritize
console.log(`Cooldown active: ${msRemaining / 1000}s remaining`);
}
This mechanism enables automatic key recovery without manual intervention.
Integration with the Routing Engine
The rate limiting service is consumed by server/src/services/router.ts, which orchestrates model selection across available keys.
A typical routing check sequence:
import { canUseProvider, canUseProviderMinute } from './services/ratelimit.js';
if (canUseProvider('openai', keyId) && canUseProviderMinute('openai', keyId)) {
// Key has capacity → assign request
} else {
// Exceeds limits → trigger cooldown, try next key
}
The router penalizes cooldown-state keys in its scoring algorithm, naturally load-balancing across the key pool.
Key Implementation Files
| File | Responsibility |
|---|---|
server/src/services/ratelimit.ts |
Core per-key rate limit logic, SQLite persistence, sliding-window cache, cooldown registry |
server/src/services/router.ts |
Consumes ratelimit API to filter and score keys during request routing |
shared/types.ts |
KeyStatus type definitions used throughout the ratelimit system |
Summary
- Per-key isolation via platform:modelId:keyId triple identifiers ensures usage tracking never crosses key boundaries
- Hybrid storage combines SQLite durability with in-memory speed for rate limit calculations
- Three enforcement dimensions: daily requests, per-minute requests, and concurrent leases
- Automatic cooldown with
getNextCooldownDurationenables graceful key recovery - Router integration allows real-time key selection based on live availability
Frequently Asked Questions
How does FreeLLMAPI prevent rate limit state from being lost on restart?
The system persists all usage records to an embedded SQLite database accessed via getDb(). While the in-memory sliding window provides fast lookups, every recordRequest() and recordTokens() call writes to SQLite. On restart, the countPersisted* queries rebuild the rate limit baseline from durable storage.
What happens when a key hits its per-minute limit but not its daily limit?
canUseProviderMinute() returns false, triggering the cooldown mechanism via setCooldown. The router deprioritizes the key for the cooldown duration. Once getNextCooldownDuration() reaches zero, the key becomes eligible again — potentially within the same day, preserving remaining daily quota.
Why does FreeLLMAPI use a triple identifier instead of just the API key string?
The platform:modelId:keyId format allows the same underlying API key to have independent rate limits across different models and providers. A key registered for both GPT-4o and GPT-3.5-Turbo tracks usage separately per model, preventing one high-traffic model from exhausting quota for another.
Can multiple concurrent requests exceed the limit before the system detects it?
The provisionalProviderRequests counter tracks in-flight requests not yet committed to the database. canUseProvider() includes these provisional requests in its availability calculation, preventing overcommitment even under high concurrency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →