How FreeLLMAPI Handles Provider Quota Exhaustion: Architecture and Implementation
FreeLLMAPI prevents service disruption by extracting quota signals from provider responses, persisting them to SQLite tables, and automatically routing traffic away from exhausted keys using a headroom-based weighting algorithm.
Provider quota management is critical for maintaining uptime in multi-tenant LLM aggregation platforms. According to the tashfeenahmed/freellmapi source code, the system implements a comprehensive quota exhaustion handling mechanism that combines real-time monitoring, persistent state tracking, and intelligent request routing to ensure continuous service availability.
The Quota Monitoring Pipeline
The platform monitors quota usage through a continuous feedback loop that intercepts every provider response and updates internal state accordingly.
Extracting Quota Signals from HTTP Responses
Every HTTP response is inspected for quota-related metadata. In server/src/services/provider-quota.ts, the extractContext and recordObservation functions scan for explicit quota headers (e.g., X-RateLimit-Remaining), 429/402 error responses, or the absence of quota headers (treated as unknown status).
When the system encounters a "real" quota signal—specifically a 429 or 402 status code with quotaSignal:true—the ratelimit.ts service immediately increments provisional usage counters and blocks the key for the remainder of the reset window.
Persisting Observations to SQLite
Quota observations are stored in two complementary SQLite tables to ensure durable state across restarts:
provider_quota_observations: An append-only audit trail that records every quota signal detectedprovider_quota_state: The current best-guess of remaining quota, updated via upsert operations on each observation
This persistence layer prevents "cold-start" scenarios where a server restart might otherwise cause the system to hammer already-exhausted provider keys. The state rows maintain the most recent limit, remaining, and reset values for each key.
Intelligent Request Routing with Quota Awareness
The routing layer uses quota headroom calculations to prioritize healthy keys and implement graceful degradation when providers reach limits.
Calculating Quota Headroom
When evaluating candidate keys for a request, the router queries getKeyQuotaHeadroom from server/src/services/provider-quota.ts. This function returns a fractional value between 0 and 1 indicating the percentage of daily quota remaining.
If the service has no observation history for a specific key, it applies a neutral default headroom of approximately 0.5, allowing new keys to participate in rotation while gathering usage data. As implemented in server/src/services/router.ts, the getKeyQuotaHeadroom call determines whether quotaWeightingApplies to the current routing decision.
Quota-Aware Weighting and Fallback Chains
The router re-orders candidate keys based on quota availability using the quotaWeighting logic in server/src/services/router.ts. Keys with low headroom values are pushed toward the end of the candidate list, effectively deprioritizing them without immediate removal from rotation.
If all keys for a primary provider exhaust their quotas, the system automatically triggers fallback chains (e.g., OpenRouter → Groq → OpenAI). The fallback mechanism respects the same quota weighting logic, ensuring that backup providers are evaluated using identical headroom calculations. This prevents cascading failures when primary providers throttle requests.
Real-Time Rate Limiting Protection
The ratelimit service acts as a circuit breaker for quota exhaustion. In server/src/services/ratelimit.ts, the system intercepts every request and monitors for quota-specific error codes. When a 429 or 402 response arrives with quotaSignal:true, the service:
- Increments the provisional usage counter for that specific key
- Marks the key as blocked when daily limits are reached
- Prevents additional requests to that key until the reset window expires
This gate operates independently of the routing weights, providing an absolute enforcement mechanism that complements the probabilistic routing adjustments.
Forecasting and Audit Trails
For operational visibility, server/src/services/quota-forecast.ts provides lightweight prediction capabilities that analyze recent usage patterns to forecast when specific quotas will exhaust. This enables dashboard displays showing "quota-remaining" meters and proactive capacity planning.
The server/src/services/request-retention.ts service maintains the comprehensive audit log in provider_quota_observations, while server/src/services/scoring.ts applies the "headroom factor" guardrail that trims model selection options when quota availability drops below configured thresholds.
// Example: Automatic quota-aware routing through the client
import { fetchLLM } from '@freellmapi/client';
async function generateResponse(prompt: string) {
const response = await fetchLLM({
model: 'gpt-4o-mini',
messages: [{ role: 'user', content: prompt }],
// The client routes through the server's router, which consults
// provider-quota and ratelimit services automatically
});
if (!response.ok) {
// If all provider quotas are exhausted, this error indicates
// complete depletion across the fallback chain
throw new Error(`Request failed: ${response.status}`);
}
return await response.json();
}
// Example: Manually inspecting quota headroom for monitoring
import { getKeyQuotaHeadroom } from '@freellmapi/server/src/services/provider-quota';
async function checkKeyHealth(keyId: string) {
const headroom = await getKeyQuotaHeadroom(keyId);
const percentage = Math.round(headroom * 100);
console.log(`Key ${keyId} has ${percentage}% quota remaining`);
if (headroom < 0.1) {
console.warn(`Key ${keyId} approaching exhaustion`);
}
}
Summary
- Signal Extraction: The system monitors HTTP responses for quota headers and error codes via
provider-quota.ts, treating 429/402 errors as explicit quota exhaustion signals. - Persistent State: SQLite tables
provider_quota_observationsandprovider_quota_statemaintain durable quota records that survive server restarts. - Headroom Calculation: The
getKeyQuotaHeadroomfunction provides fractional availability metrics used to weight routing decisions. - Intelligent Routing:
router.tsimplements quota-aware weighting and automatic fallback chains to steer traffic away from depleted keys. - Circuit Breaking:
ratelimit.tsenforces hard stops on exhausted keys while the reset window remains active. - Operational Visibility: Optional forecasting and retention services provide audit trails and predictive exhaustion alerts.
Frequently Asked Questions
How does FreeLLMAPI detect when a provider key has exhausted its quota?
The system detects quota exhaustion by inspecting HTTP response headers and status codes. In server/src/services/provider-quota.ts, the recordObservation function processes responses containing 429/402 status codes or quota-specific headers like X-RateLimit-Remaining. The ratelimit.ts service specifically looks for quotaSignal:true flags on error responses to distinguish between general rate limiting and actual quota depletion.
What happens to requests when all provider keys are exhausted?
When all keys within a provider are blocked, the router triggers configured fallback chains as implemented in server/src/services/router.ts. The system attempts the next provider in the sequence (e.g., moving from OpenRouter to Groq to OpenAI), applying identical quota headroom calculations to each candidate. If the entire fallback chain is exhausted, the request fails with an error indicating complete quota depletion across all available providers.
How does the router prioritize keys with available quota?
The router calls getKeyQuotaHeadroom for each candidate key and applies quotaWeighting logic that sorts keys by their fractional availability. Keys with higher headroom values (closer to 1.0) receive priority placement in the request queue, while keys approaching zero headroom are deprioritized. This weighting applies dynamically on every routing decision without requiring manual key rotation.
Does quota state persist across server restarts?
Yes. FreeLLMAPI persists quota observations in SQLite tables (provider_quota_observations and provider_quota_state) rather than relying solely on in-memory storage. This design prevents the system from accidentally reusing exhausted keys immediately after a deployment or crash, as the provider_quota_state table maintains the most recent remaining and reset values for each key across process lifecycles.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →