# How FreeLLMAPI Tracks Per-Key Rate Limits: SQLite + Sliding-Window Architecture Explained

> Learn how FreeLLMAPI tracks per-key rate limits with its SQLite and sliding-window architecture. Discover efficient usage monitoring for your API keys.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: internals
- Published: 2026-08-30

---

**FreeLLMAPI enforces per-key rate limits using a hybrid architecture: an embedded SQLite database for durable usage records and an in-memory sliding-window cache for fast real-time checks, with each key uniquely identified by a platform:modelId:keyId triple.**

The open-source FreeLLMAPI project implements sophisticated rate limiting at the individual API key level. According to the [tashfeenahmed/freellmapi](https://github.com/tashfeenahmed/freellmapi) source code, the system combines database persistence with memory caching to track requests, tokens, and concurrency across multiple time windows while maintaining low latency for routing decisions.

## Per-Key Identification Strategy

FreeLLMAPI identifies each rate-limited entity using a composite key format.

In [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts), keys are structured as **platform:modelId:keyId** — for example, `openai:gpt-4o:key_abc123`. This triple ensures complete isolation between different providers, models, and individual key credentials (see the key format comment at lines 13-14).

The `usageCacheKey` function (lines 404-405) builds cache keys from this triple, guaranteeing that usage statistics for one key never bleed into another.

## Dual-Layer Storage: Database + Memory Cache

The rate limiting system operates across two storage layers:

- **SQLite database** (`getDb()`) — persists all usage records for durability and historical analysis
- **In-memory sliding window** — provides sub-millisecond lookups for hot-path routing decisions

Every recorded request is written to both layers simultaneously. The memory cache maintains recent request timestamps in `memoryRequestCount` and token tallies in `memoryTokenCount`, while the database stores the authoritative record via `countPersisted*` queries.

## Core Rate Limit Checks

The ratelimit service exposes three primary enforcement functions that the router consults before assigning a key to a request.

### Daily and Per-Minute Request Caps

- **`canUseProvider()`** (lines 639-647) — validates against `providerDailyRequestCount`, combining persisted DB counts with in-memory window data plus any `provisionalProviderRequests` (in-flight requests not yet committed)

- **`canUseProviderMinute()`** (lines 664-670) — enforces `providerMinuteRequestCount` using the same dual-source calculation

Both functions return boolean availability status, allowing the router to immediately disqualify exhausted keys.

### Concurrency Limits

- **`canUseKeyConcurrency()`** — prevents a single key from accumulating too many simultaneous leases

The system tracks `inFlightForKey` (lines 109-120) to limit parallel execution. When a key reaches its concurrency cap, new requests queue or route to alternative keys.

## Recording Usage: The `recordRequest()` and `recordTokens()` Pipeline

After a request completes, the system commits usage through two dedicated functions:

```ts
// server/src/services/ratelimit.ts
// Lines 812-820 and 832-834

import { recordRequest, recordTokens } from './services/ratelimit.js';

// Increment request counter by 1
recordRequest('openai', modelId, keyId);

// Add token consumption to the ledger
recordTokens('openai', modelId, keyId, 150);

```

Both functions execute `pushMemoryRequest` to update the sliding window and execute INSERT statements against SQLite. This design ensures the in-memory view stays synchronized with the durable store.

## Cooldown Enforcement and Querying

When a key exhausts its limits, FreeLLMAPI applies a cooldown penalty rather than permanent blacklisting.

The `setCooldown` function registers a temporary ban. The router later queries `getNextCooldownDuration` (lines 877-878) to determine remaining penalty time:

```ts
import { getNextCooldownDuration } from './services/ratelimit.js';

const msRemaining = getNextCooldownDuration('openai', modelId, keyId);

if (msRemaining > 0) {
  // Key is rate-limited; router will skip or deprioritize
  console.log(`Cooldown active: ${msRemaining / 1000}s remaining`);
}

```

This mechanism enables automatic key recovery without manual intervention.

## Integration with the Routing Engine

The rate limiting service is consumed by [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts), which orchestrates model selection across available keys.

A typical routing check sequence:

```ts
import { canUseProvider, canUseProviderMinute } from './services/ratelimit.js';

if (canUseProvider('openai', keyId) && canUseProviderMinute('openai', keyId)) {
  // Key has capacity → assign request
} else {
  // Exceeds limits → trigger cooldown, try next key
}

```

The router penalizes cooldown-state keys in its scoring algorithm, naturally load-balancing across the key pool.

## Key Implementation Files

| File | Responsibility |
|------|---------------|
| [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) | Core per-key rate limit logic, SQLite persistence, sliding-window cache, cooldown registry |
| [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) | Consumes ratelimit API to filter and score keys during request routing |
| [`shared/types.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/shared/types.ts) | `KeyStatus` type definitions used throughout the ratelimit system |

## Summary

- **Per-key isolation** via platform:modelId:keyId triple identifiers ensures usage tracking never crosses key boundaries
- **Hybrid storage** combines SQLite durability with in-memory speed for rate limit calculations
- **Three enforcement dimensions**: daily requests, per-minute requests, and concurrent leases
- **Automatic cooldown** with `getNextCooldownDuration` enables graceful key recovery
- **Router integration** allows real-time key selection based on live availability

## Frequently Asked Questions

### How does FreeLLMAPI prevent rate limit state from being lost on restart?

The system persists all usage records to an embedded SQLite database accessed via `getDb()`. While the in-memory sliding window provides fast lookups, every `recordRequest()` and `recordTokens()` call writes to SQLite. On restart, the `countPersisted*` queries rebuild the rate limit baseline from durable storage.

### What happens when a key hits its per-minute limit but not its daily limit?

`canUseProviderMinute()` returns false, triggering the cooldown mechanism via `setCooldown`. The router deprioritizes the key for the cooldown duration. Once `getNextCooldownDuration()` reaches zero, the key becomes eligible again — potentially within the same day, preserving remaining daily quota.

### Why does FreeLLMAPI use a triple identifier instead of just the API key string?

The **platform:modelId:keyId** format allows the same underlying API key to have independent rate limits across different models and providers. A key registered for both GPT-4o and GPT-3.5-Turbo tracks usage separately per model, preventing one high-traffic model from exhausting quota for another.

### Can multiple concurrent requests exceed the limit before the system detects it?

The `provisionalProviderRequests` counter tracks in-flight requests not yet committed to the database. `canUseProvider()` includes these provisional requests in its availability calculation, preventing overcommitment even under high concurrency.