How FreeLLMAPI Handles 429, 5xx, and Timeout Errors During Routing
FreeLLMAPI treats upstream failures as routing signals, applying tiered priority penalties to throttled or failing models and automatically falling back to alternative providers until all candidates are exhausted.
FreeLLMAPI is an open-source routing layer that aggregates multiple LLM providers into a unified API. Understanding how the tashfeenahmed/freellmapi repository handles 429 rate limits, 5xx server errors, and network timeouts during routing is critical for building resilient AI applications that require high availability across distributed model endpoints.
Error Classification and Priority Penalties
The routing engine in router.ts classifies upstream failures into three distinct categories, each triggering specific penalty calculations that influence model selection.
429 Rate-Limit Detection
When a provider returns HTTP 429 (or a 402 quota signal), the router recognizes this as a hard rate-limit. According to router.ts lines 258-292, the code records a 429-penalty defined by the constant PENALTY_PER_429 = 3. This demotes the model’s priority by three positions and places the model-key combination into a cooldown bucket. The penalized model is immediately removed from the current candidate pool, forcing the router to retry the request with the next highest-priority alternative.
5xx Server Error Handling
Any non-limit upstream failure—including HTTP 5xx responses, empty streams, or dead sockets—is classified as a generic failure. As implemented in router.ts line 261, these errors incur PENALTY_PER_FAIL = 1, applying a single priority step demotion. The request retries immediately on the next model in the chain without surfacing the error to the client.
Timeout Identification
Timeouts are identified by matching known timeout markers in the error text, specifically within router.ts lines 699-704. These failures are counted in routing statistics as "timeouts" and incur the same generic failure penalty (PENALTY_PER_FAIL = 1) as 5xx errors. The timeout-specific latency is recorded for future scheduling decisions, allowing the system to prefer faster-responding providers.
The Penalty and Cooldown System
FreeLLMAPI maintains resilient traffic distribution through a dynamic penalty system and temporal cooldown mechanisms.
Penalty Tiers
The router maintains a priority order of available models. Each 429 error adds three priority positions of penalty, while any other failure adds one. This incremental demotion allows the router to quickly steer traffic away from throttled providers while preserving their eligibility for recovery once quotas reset. The penalty calculations are implemented in router.ts lines 258-292 and 1035-1048.
Cooldown Bookkeeping
When a 429 is detected, the system invokes setCooldown (visible in media.ts line 804). The cooldown table records the earliest time the model-key may be reused, as detailed in ratelimit.ts lines 838-898. Cooldowns remain short for transient RPM limits but extend up to 24 hours for daily quota exhaustion. This ensures exhausted providers do not receive traffic until their rate limits have legitimately reset.
Automatic Retry and Fallback Mechanisms
The core routing logic implements transparent failover without client intervention.
The Fallback Loop
The retry loop in fallback-loop.ts catches any upstream error, checks its classification (rate_limited, server_error, or timeout), and automatically selects the next model. This process repeats without exposing intermediate failures to the API consumer. The loop implementation spans fallback-loop.ts lines 185-215 and 592-639.
Provider Exhaustion
Only when all candidates are exhausted does the router raise a RouteError with status 429. At the HTTP endpoint level (routes/proxy.ts lines 656-676 and 1733-1735), the router translates this final state into a rate_limit_error JSON payload for client consumption.
// Example: the router penalises a 429 and retries the next model
try {
const result = await router.routeChat(request);
return result;
} catch (e) {
// All models exhausted – the router throws a RouteError(429)
if (e instanceof RouteError && e.status === 429) {
res.status(429).json({ error: { message: e.message, type: 'rate_limit_error' } });
} else {
// other unexpected errors
res.status(500).json({ error: { message: e.message, type: 'server_error' } });
}
}
Timeout Configuration and Provider Defaults
Each provider module imports provider-timeout.ts to define sensible defaults. Most OpenAI-compatible providers use 60 seconds, while Ollama Cloud configurations allow up to 120 seconds. Calls wrap fetch() with AbortSignal.timeout, converting hung upstream requests into recognized timeout errors that feed the routing penalty system. This implementation appears in providers/openai-compat.ts lines 15-55 and provider-timeout.ts.
// Provider-side timeout definition (openai-compat provider)
import { providerTimeoutMs } from '../lib/provider-timeout.js';
export class OpenAICompatProvider extends BaseProvider {
private readonly timeoutMs = providerTimeoutMs('openai', opts.timeoutMs ?? 60_000);
async chatCompletion(..., options) {
const timeout = options?.timeoutMs ?? this.timeoutMs;
const res = await fetch(url, { signal: AbortSignal.timeout(timeout) });
// …
}
}
Client-Facing Error Translation
While internal routing continues aggressive retries, the public API contract remains stable. Successful retries return the upstream response transparently. When all providers fail, proxy.ts translates the internal RouteError into structured JSON payloads: 429 becomes rate_limit_error, while unexpected terminal failures become server_error. This separation of concerns allows operators to monitor internal routing health without polluting client error handling.
Summary
- 429 errors trigger a severe penalty (3 priority steps) and immediate cooldown, removing the provider from rotation until rate limits reset.
- 5xx errors and timeouts incur a single penalty step and automatic retry on the next available model.
- The fallback loop in
fallback-loop.tshandles error classification and model switching without client exposure. - Cooldown bookkeeping via
ratelimit.tsprevents hammering exhausted providers with retry traffic. - Timeout configurations in
provider-timeout.tsensure hung requests abort cleanly and feed penalty statistics.
Frequently Asked Questions
How does FreeLLMAPI distinguish between a 429 rate limit and other errors?
The router inspects HTTP status codes and response bodies in router.ts and provider-quota.ts. HTTP 429 (and 402 quota signals) are classified as hard rate limits triggering PENALTY_PER_429, while 5xx codes and timeout markers trigger the generic PENALTY_PER_FAIL classification.
What happens when all provider models are exhausted?
When the fallback loop depletes all candidate models, the router throws a RouteError with status 429. The HTTP handler in proxy.ts catches this and returns a JSON error object with type rate_limit_error, signaling to the client that no upstream capacity is currently available.
How long do cooldown periods last for rate-limited models?
Cooldown durations vary based on the rate limit type. Transient RPM limits incur short cooldowns, while daily quota exhaustion triggers cooldowns up to 24 hours. The setCooldown function in ratelimit.ts lines 838-898 calculates these windows based on provider-specific retry-after headers or default heuristic values.
Can timeout durations be customized per provider?
Yes. The provider-timeout.ts module exports providerTimeoutMs(), allowing each provider class to specify custom defaults. Providers like OpenAI-compatible endpoints default to 60 seconds, while Ollama Cloud uses 120 seconds. Individual requests can also override these via the timeoutMs option in the request configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →