# How FreeLLMAPI Handles 429, 5xx, and Timeout Errors During Routing

> Discover how FreeLLMAPI intelligently handles 429, 5xx, and timeout errors. Learn about its tiered priority penalties and automatic fallback to ensure robust API routing and service continuity.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: internals
- Published: 2026-09-01

---

**FreeLLMAPI treats upstream failures as routing signals, applying tiered priority penalties to throttled or failing models and automatically falling back to alternative providers until all candidates are exhausted.**

FreeLLMAPI is an open-source routing layer that aggregates multiple LLM providers into a unified API. Understanding how the tashfeenahmed/freellmapi repository handles 429 rate limits, 5xx server errors, and network timeouts during routing is critical for building resilient AI applications that require high availability across distributed model endpoints.

## Error Classification and Priority Penalties

The routing engine in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) classifies upstream failures into three distinct categories, each triggering specific penalty calculations that influence model selection.

### 429 Rate-Limit Detection

When a provider returns **HTTP 429** (or a 402 quota signal), the router recognizes this as a *hard* rate-limit. According to [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) lines 258-292, the code records a **429-penalty** defined by the constant `PENALTY_PER_429 = 3`. This demotes the model’s priority by three positions and places the model-key combination into a cooldown bucket. The penalized model is immediately removed from the current candidate pool, forcing the router to retry the request with the next highest-priority alternative.

### 5xx Server Error Handling

Any non-limit upstream failure—including **HTTP 5xx** responses, empty streams, or dead sockets—is classified as a generic failure. As implemented in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) line 261, these errors incur `PENALTY_PER_FAIL = 1`, applying a single priority step demotion. The request retries immediately on the next model in the chain without surfacing the error to the client.

### Timeout Identification

Timeouts are identified by matching known timeout markers in the error text, specifically within [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) lines 699-704. These failures are counted in routing statistics as "timeouts" and incur the same generic failure penalty (`PENALTY_PER_FAIL = 1`) as 5xx errors. The timeout-specific latency is recorded for future scheduling decisions, allowing the system to prefer faster-responding providers.

## The Penalty and Cooldown System

FreeLLMAPI maintains resilient traffic distribution through a dynamic penalty system and temporal cooldown mechanisms.

### Penalty Tiers

The router maintains a **priority order** of available models. Each 429 error adds three priority positions of penalty, while any other failure adds one. This incremental demotion allows the router to quickly steer traffic away from throttled providers while preserving their eligibility for recovery once quotas reset. The penalty calculations are implemented in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) lines 258-292 and 1035-1048.

### Cooldown Bookkeeping

When a 429 is detected, the system invokes `setCooldown` (visible in [`media.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/media.ts) line 804). The cooldown table records the earliest time the model-key may be reused, as detailed in [`ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/ratelimit.ts) lines 838-898. Cooldowns remain short for transient RPM limits but extend up to 24 hours for daily quota exhaustion. This ensures exhausted providers do not receive traffic until their rate limits have legitimately reset.

## Automatic Retry and Fallback Mechanisms

The core routing logic implements transparent failover without client intervention.

### The Fallback Loop

The **retry loop** in [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts) catches any upstream error, checks its classification (`rate_limited`, `server_error`, or `timeout`), and automatically selects the next model. This process repeats without exposing intermediate failures to the API consumer. The loop implementation spans [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts) lines 185-215 and 592-639.

### Provider Exhaustion

Only when **all** candidates are exhausted does the router raise a `RouteError` with status 429. At the HTTP endpoint level ([`routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/routes/proxy.ts) lines 656-676 and 1733-1735), the router translates this final state into a `rate_limit_error` JSON payload for client consumption.

```typescript
// Example: the router penalises a 429 and retries the next model
try {
  const result = await router.routeChat(request);
  return result;
} catch (e) {
  // All models exhausted – the router throws a RouteError(429)
  if (e instanceof RouteError && e.status === 429) {
    res.status(429).json({ error: { message: e.message, type: 'rate_limit_error' } });
  } else {
    // other unexpected errors
    res.status(500).json({ error: { message: e.message, type: 'server_error' } });
  }
}

```

## Timeout Configuration and Provider Defaults

Each provider module imports [`provider-timeout.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/provider-timeout.ts) to define sensible defaults. Most OpenAI-compatible providers use 60 seconds, while Ollama Cloud configurations allow up to 120 seconds. Calls wrap `fetch()` with `AbortSignal.timeout`, converting hung upstream requests into recognized timeout errors that feed the routing penalty system. This implementation appears in [`providers/openai-compat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/providers/openai-compat.ts) lines 15-55 and [`provider-timeout.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/provider-timeout.ts).

```typescript
// Provider-side timeout definition (openai-compat provider)
import { providerTimeoutMs } from '../lib/provider-timeout.js';

export class OpenAICompatProvider extends BaseProvider {
  private readonly timeoutMs = providerTimeoutMs('openai', opts.timeoutMs ?? 60_000);
  
  async chatCompletion(..., options) {
    const timeout = options?.timeoutMs ?? this.timeoutMs;
    const res = await fetch(url, { signal: AbortSignal.timeout(timeout) });
    // …
  }
}

```

## Client-Facing Error Translation

While internal routing continues aggressive retries, the public API contract remains stable. Successful retries return the upstream response transparently. When all providers fail, [`proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/proxy.ts) translates the internal `RouteError` into structured JSON payloads: **429** becomes `rate_limit_error`, while unexpected terminal failures become `server_error`. This separation of concerns allows operators to monitor internal routing health without polluting client error handling.

## Summary

- **429 errors** trigger a severe penalty (3 priority steps) and immediate cooldown, removing the provider from rotation until rate limits reset.
- **5xx errors** and **timeouts** incur a single penalty step and automatic retry on the next available model.
- The **fallback loop** in [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts) handles error classification and model switching without client exposure.
- **Cooldown bookkeeping** via [`ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/ratelimit.ts) prevents hammering exhausted providers with retry traffic.
- **Timeout configurations** in [`provider-timeout.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/provider-timeout.ts) ensure hung requests abort cleanly and feed penalty statistics.

## Frequently Asked Questions

### How does FreeLLMAPI distinguish between a 429 rate limit and other errors?

The router inspects HTTP status codes and response bodies in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) and [`provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/provider-quota.ts). HTTP 429 (and 402 quota signals) are classified as hard rate limits triggering `PENALTY_PER_429`, while 5xx codes and timeout markers trigger the generic `PENALTY_PER_FAIL` classification.

### What happens when all provider models are exhausted?

When the fallback loop depletes all candidate models, the router throws a `RouteError` with status 429. The HTTP handler in [`proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/proxy.ts) catches this and returns a JSON error object with type `rate_limit_error`, signaling to the client that no upstream capacity is currently available.

### How long do cooldown periods last for rate-limited models?

Cooldown durations vary based on the rate limit type. Transient RPM limits incur short cooldowns, while daily quota exhaustion triggers cooldowns up to 24 hours. The `setCooldown` function in [`ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/ratelimit.ts) lines 838-898 calculates these windows based on provider-specific retry-after headers or default heuristic values.

### Can timeout durations be customized per provider?

Yes. The [`provider-timeout.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/provider-timeout.ts) module exports `providerTimeoutMs()`, allowing each provider class to specify custom defaults. Providers like OpenAI-compatible endpoints default to 60 seconds, while Ollama Cloud uses 120 seconds. Individual requests can also override these via the `timeoutMs` option in the request configuration.