# How FreeLLMAPI Handles 429 Rate Limit Errors with Automatic Failover

> Learn how FreeLLMAPI automatically handles 429 rate limit errors. Discover its automatic failover strategy that reroutes requests to the next available provider to ensure uninterrupted service.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-08-31

---

**FreeLLMAPI treats HTTP 429 responses as hard health signals that trigger a -3 priority penalty, temporarily demoting the affected model while automatically rerouting requests to the next available provider.**

The open-source FreeLLMAPI gateway ensures continuous LLM availability by converting rate limit errors into intelligent routing decisions. When a provider returns a 429 status, the system doesn't fail the request; instead, it penalizes the exhausted model and seamlessly fails over to alternative keys. This article examines the three-layer architecture in `tashfeenahmed/freellmapi` that enables automatic failover without client-side retry logic.

## Three-Layer Error Handling Architecture

FreeLLMAPI implements a tightly-coupled error handling system across three distinct layers to transform raw 429 responses into routing intelligence.

### Error Classification Layer

When a provider returns an HTTP 429 response (or any error text containing "rate limit"), the **`classifyAttemptError`** function in [`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts) (line 52) tags the error as `rate_limited`. This classification determines the penalty path for the fallback loop, distinguishing rate limits from transient network errors or content moderation flags.

### Model-Level Penalty Recording

The fallback loop invokes **`recordRateLimitHit(modelDbId)`**, defined in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) (line 91), to register a severe penalty for the offending model. Unlike ordinary failures that incur a -1 priority penalty, rate limit errors trigger **`PENALTY_PER_429 = 3`**, pushing the model down the candidate queue by three priority slots. The penalty state is stored in an in-memory map called `rateLimitPenalties`, enabling microsecond-fast lookups during routing decisions.

### Dynamic Priority Routing

During each routing round, the **`getPenalty`** function retrieves the current penalty weight for every model and adds it to the base priority score. Higher penalties demote models to the bottom of the candidate list. To ensure recovery, FreeLLMAPI implements automatic penalty decay: every **`DECAY_INTERVAL_MS = 120000`** (2 minutes), the system subtracts **`DECAY_AMOUNT = 1`** from each model's penalty count (lines 58-65 in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)), allowing demoted models to climb back up the priority queue as their quotas refresh.

## Automatic Failover Execution Flow

The complete failover sequence operates as follows:

1. **Request dispatch** – A client request is sent to the highest-priority model/key pair.
2. **429 detection** – The provider returns HTTP 429, triggering **`isRateLimitSignal(err)`** inside the fallback loop.
3. **Classification** – The error is classified as `rate_limited` (lines 32-34 in [`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts)).
4. **Penalty recording** – `recordRateLimitHit` forwards to **`recordModelFailure`** with the triple-weight penalty (lines 73-81 in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)).
5. **Priority adjustment** – The updated penalty is factored into the model's composite priority score during the next routing calculation.
6. **Automatic failover** – The router skips the penalized model (and any API keys sharing it), instantly retrying the request with the next viable model in the candidate list.
7. **Recovery tracking** – If no further 429s occur, the 2-minute decay interval gradually reduces the penalty until the model returns to its original priority tier.

## Key Implementation Files

The rate limit handling logic spans three critical source files:

- **[`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)** – Maintains penalty state via `recordModelFailure`, `recordRateLimitHit`, and the decay scheduler.
- **[`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts)** – Orchestrates request attempts and classifies errors using `classifyAttemptError`.
- **[`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts)** – Provides **`isRateLimitSignal`** helpers to detect rate limit patterns in provider responses.

## Client Integration Example

FreeLLMAPI handles all failover logic server-side, so clients receive a successful response even when the primary model is rate-limited:

```typescript
// Example: Fetching from the FreeLLMAPI /v1/chat/completions endpoint
import fetch from 'node-fetch';

async function generateResponse(prompt: string) {
  const resp = await fetch('https://api.freellmapi.com/v1/chat/completions', {
    method: 'POST',
    headers: { 
      'Content-Type': 'application/json', 
      Authorization: 'Bearer <your-token>' 
    },
    body: JSON.stringify({ 
      model: 'gpt-4o-mini', 
      messages: [{ role: 'user', content: prompt }] 
    })
  });

  // 429 errors are handled internally via automatic failover
  const data = await resp.json();
  return data.choices?.[0]?.message?.content;
}

generateResponse('Explain quantum tunneling');

```

If `gpt-4o-mini` hits its per-minute quota, the service automatically routes the request to the next available model (e.g., a free-tier alternative) and returns the answer transparently.

For debugging purposes, you can inspect the current penalty map:

```typescript
import { getAllPenalties } from '@freellmapi/server/src/services/router';

console.log(getAllPenalties());
// => [{ modelDbId: 42, count: 3, penalty: 6 }, …]

```

## Summary

- **Hard health signals** – FreeLLMAPI treats 429 errors as severe infrastructure signals rather than transient faults.
- **Triple-weight penalties** – Rate limits incur a -3 priority penalty compared to -1 for standard failures, stored in the `rateLimitPenalties` map.
- **Automatic decay** – Penalties decrease by 1 every 2 minutes, allowing models to recover gradually without manual intervention.
- **Server-side failover** – Clients receive seamless responses while the gateway handles model switching internally via [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts).

## Frequently Asked Questions

### How long does a model remain penalized after a 429 error?

The penalty decays automatically every 2 minutes (`DECAY_INTERVAL_MS = 120000`), reducing the penalty counter by 1 (`DECAY_AMOUNT = 1`) until it reaches zero. A model that receives a single 429 will return to full priority after approximately 6 minutes.

### What is the difference between a 429 penalty and a regular failure?

Regular connection or timeout errors trigger `recordModelFailure` with a weight of -1, while rate limit errors invoke `recordRateLimitHit`, which applies a `PENALTY_PER_429 = 3` weight. This heavier penalty ensures rate-limited models are deprioritized faster than those experiencing temporary network issues.

### Which source files contain the rate limit handling logic?

The primary logic resides in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) (penalty storage and decay), [`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts) (error classification and orchestration), and [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) (detection helpers).

### Does the client need to implement retry logic for 429 errors?

No. FreeLLMAPI handles all 429 detection, penalty application, and model failover internally. The client receives a standard 200 response from whichever healthy model ultimately fulfills the request, eliminating the need for client-side exponential backoff or retry loops.