# How FreeLLMAPI Handles Provider Quota Exhaustion: Architecture and Implementation

> Learn how FreeLLMAPI prevents service disruption by extracting quota signals, persisting them, and routing traffic away from exhausted keys using a headroom-based algorithm.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: architecture
- Published: 2026-09-04

---

**FreeLLMAPI prevents service disruption by extracting quota signals from provider responses, persisting them to SQLite tables, and automatically routing traffic away from exhausted keys using a headroom-based weighting algorithm.**

Provider quota management is critical for maintaining uptime in multi-tenant LLM aggregation platforms. According to the tashfeenahmed/freellmapi source code, the system implements a comprehensive quota exhaustion handling mechanism that combines real-time monitoring, persistent state tracking, and intelligent request routing to ensure continuous service availability.

## The Quota Monitoring Pipeline

The platform monitors quota usage through a continuous feedback loop that intercepts every provider response and updates internal state accordingly.

### Extracting Quota Signals from HTTP Responses

Every HTTP response is inspected for quota-related metadata. In [`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts), the `extractContext` and `recordObservation` functions scan for explicit quota headers (e.g., `X-RateLimit-Remaining`), 429/402 error responses, or the absence of quota headers (treated as *unknown* status).

When the system encounters a "real" quota signal—specifically a 429 or 402 status code with `quotaSignal:true`—the [`ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/ratelimit.ts) service immediately increments provisional usage counters and blocks the key for the remainder of the reset window.

### Persisting Observations to SQLite

Quota observations are stored in two complementary SQLite tables to ensure durable state across restarts:

- **`provider_quota_observations`**: An append-only audit trail that records every quota signal detected
- **`provider_quota_state`**: The current best-guess of remaining quota, updated via upsert operations on each observation

This persistence layer prevents "cold-start" scenarios where a server restart might otherwise cause the system to hammer already-exhausted provider keys. The state rows maintain the most recent `limit`, `remaining`, and `reset` values for each key.

## Intelligent Request Routing with Quota Awareness

The routing layer uses quota headroom calculations to prioritize healthy keys and implement graceful degradation when providers reach limits.

### Calculating Quota Headroom

When evaluating candidate keys for a request, the router queries `getKeyQuotaHeadroom` from [`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts). This function returns a fractional value between 0 and 1 indicating the percentage of daily quota remaining.

If the service has no observation history for a specific key, it applies a neutral default headroom of approximately 0.5, allowing new keys to participate in rotation while gathering usage data. As implemented in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts), the `getKeyQuotaHeadroom` call determines whether `quotaWeightingApplies` to the current routing decision.

### Quota-Aware Weighting and Fallback Chains

The router re-orders candidate keys based on quota availability using the `quotaWeighting` logic in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts). Keys with low headroom values are pushed toward the end of the candidate list, effectively deprioritizing them without immediate removal from rotation.

If all keys for a primary provider exhaust their quotas, the system automatically triggers fallback chains (e.g., OpenRouter → Groq → OpenAI). The fallback mechanism respects the same quota weighting logic, ensuring that backup providers are evaluated using identical headroom calculations. This prevents cascading failures when primary providers throttle requests.

## Real-Time Rate Limiting Protection

The ratelimit service acts as a circuit breaker for quota exhaustion. In [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts), the system intercepts every request and monitors for quota-specific error codes. When a 429 or 402 response arrives with `quotaSignal:true`, the service:

1. Increments the provisional usage counter for that specific key
2. Marks the key as blocked when daily limits are reached
3. Prevents additional requests to that key until the reset window expires

This gate operates independently of the routing weights, providing an absolute enforcement mechanism that complements the probabilistic routing adjustments.

## Forecasting and Audit Trails

For operational visibility, [`server/src/services/quota-forecast.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/quota-forecast.ts) provides lightweight prediction capabilities that analyze recent usage patterns to forecast when specific quotas will exhaust. This enables dashboard displays showing "quota-remaining" meters and proactive capacity planning.

The [`server/src/services/request-retention.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/request-retention.ts) service maintains the comprehensive audit log in `provider_quota_observations`, while [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts) applies the "headroom factor" guardrail that trims model selection options when quota availability drops below configured thresholds.

```typescript
// Example: Automatic quota-aware routing through the client
import { fetchLLM } from '@freellmapi/client';

async function generateResponse(prompt: string) {
  const response = await fetchLLM({
    model: 'gpt-4o-mini',
    messages: [{ role: 'user', content: prompt }],
    // The client routes through the server's router, which consults
    // provider-quota and ratelimit services automatically
  });

  if (!response.ok) {
    // If all provider quotas are exhausted, this error indicates
    // complete depletion across the fallback chain
    throw new Error(`Request failed: ${response.status}`);
  }

  return await response.json();
}

```

```typescript
// Example: Manually inspecting quota headroom for monitoring
import { getKeyQuotaHeadroom } from '@freellmapi/server/src/services/provider-quota';

async function checkKeyHealth(keyId: string) {
  const headroom = await getKeyQuotaHeadroom(keyId);
  const percentage = Math.round(headroom * 100);
  
  console.log(`Key ${keyId} has ${percentage}% quota remaining`);
  
  if (headroom < 0.1) {
    console.warn(`Key ${keyId} approaching exhaustion`);
  }
}

```

## Summary

- **Signal Extraction**: The system monitors HTTP responses for quota headers and error codes via [`provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/provider-quota.ts), treating 429/402 errors as explicit quota exhaustion signals.
- **Persistent State**: SQLite tables `provider_quota_observations` and `provider_quota_state` maintain durable quota records that survive server restarts.
- **Headroom Calculation**: The `getKeyQuotaHeadroom` function provides fractional availability metrics used to weight routing decisions.
- **Intelligent Routing**: [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) implements quota-aware weighting and automatic fallback chains to steer traffic away from depleted keys.
- **Circuit Breaking**: [`ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/ratelimit.ts) enforces hard stops on exhausted keys while the reset window remains active.
- **Operational Visibility**: Optional forecasting and retention services provide audit trails and predictive exhaustion alerts.

## Frequently Asked Questions

### How does FreeLLMAPI detect when a provider key has exhausted its quota?

The system detects quota exhaustion by inspecting HTTP response headers and status codes. In [`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts), the `recordObservation` function processes responses containing 429/402 status codes or quota-specific headers like `X-RateLimit-Remaining`. The [`ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/ratelimit.ts) service specifically looks for `quotaSignal:true` flags on error responses to distinguish between general rate limiting and actual quota depletion.

### What happens to requests when all provider keys are exhausted?

When all keys within a provider are blocked, the router triggers configured fallback chains as implemented in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts). The system attempts the next provider in the sequence (e.g., moving from OpenRouter to Groq to OpenAI), applying identical quota headroom calculations to each candidate. If the entire fallback chain is exhausted, the request fails with an error indicating complete quota depletion across all available providers.

### How does the router prioritize keys with available quota?

The router calls `getKeyQuotaHeadroom` for each candidate key and applies `quotaWeighting` logic that sorts keys by their fractional availability. Keys with higher headroom values (closer to 1.0) receive priority placement in the request queue, while keys approaching zero headroom are deprioritized. This weighting applies dynamically on every routing decision without requiring manual key rotation.

### Does quota state persist across server restarts?

Yes. FreeLLMAPI persists quota observations in SQLite tables (`provider_quota_observations` and `provider_quota_state`) rather than relying solely on in-memory storage. This design prevents the system from accidentally reusing exhausted keys immediately after a deployment or crash, as the `provider_quota_state` table maintains the most recent `remaining` and `reset` values for each key across process lifecycles.