How FreeLLMAPI Handles 429 Rate Limit Errors with Automatic Failover

FreeLLMAPI treats HTTP 429 responses as hard health signals that trigger a -3 priority penalty, temporarily demoting the affected model while automatically rerouting requests to the next available provider.

The open-source FreeLLMAPI gateway ensures continuous LLM availability by converting rate limit errors into intelligent routing decisions. When a provider returns a 429 status, the system doesn't fail the request; instead, it penalizes the exhausted model and seamlessly fails over to alternative keys. This article examines the three-layer architecture in tashfeenahmed/freellmapi that enables automatic failover without client-side retry logic.

Three-Layer Error Handling Architecture

FreeLLMAPI implements a tightly-coupled error handling system across three distinct layers to transform raw 429 responses into routing intelligence.

Error Classification Layer

When a provider returns an HTTP 429 response (or any error text containing "rate limit"), the classifyAttemptError function in server/src/lib/fallback-loop.ts (line 52) tags the error as rate_limited. This classification determines the penalty path for the fallback loop, distinguishing rate limits from transient network errors or content moderation flags.

Model-Level Penalty Recording

The fallback loop invokes recordRateLimitHit(modelDbId), defined in server/src/services/router.ts (line 91), to register a severe penalty for the offending model. Unlike ordinary failures that incur a -1 priority penalty, rate limit errors trigger PENALTY_PER_429 = 3, pushing the model down the candidate queue by three priority slots. The penalty state is stored in an in-memory map called rateLimitPenalties, enabling microsecond-fast lookups during routing decisions.

Dynamic Priority Routing

During each routing round, the getPenalty function retrieves the current penalty weight for every model and adds it to the base priority score. Higher penalties demote models to the bottom of the candidate list. To ensure recovery, FreeLLMAPI implements automatic penalty decay: every DECAY_INTERVAL_MS = 120000 (2 minutes), the system subtracts DECAY_AMOUNT = 1 from each model's penalty count (lines 58-65 in server/src/services/router.ts), allowing demoted models to climb back up the priority queue as their quotas refresh.

Automatic Failover Execution Flow

The complete failover sequence operates as follows:

  1. Request dispatch – A client request is sent to the highest-priority model/key pair.
  2. 429 detection – The provider returns HTTP 429, triggering isRateLimitSignal(err) inside the fallback loop.
  3. Classification – The error is classified as rate_limited (lines 32-34 in server/src/lib/fallback-loop.ts).
  4. Penalty recording – recordRateLimitHit forwards to recordModelFailure with the triple-weight penalty (lines 73-81 in server/src/services/router.ts).
  5. Priority adjustment – The updated penalty is factored into the model's composite priority score during the next routing calculation.
  6. Automatic failover – The router skips the penalized model (and any API keys sharing it), instantly retrying the request with the next viable model in the candidate list.
  7. Recovery tracking – If no further 429s occur, the 2-minute decay interval gradually reduces the penalty until the model returns to its original priority tier.

Key Implementation Files

The rate limit handling logic spans three critical source files:

Client Integration Example

FreeLLMAPI handles all failover logic server-side, so clients receive a successful response even when the primary model is rate-limited:

// Example: Fetching from the FreeLLMAPI /v1/chat/completions endpoint
import fetch from 'node-fetch';

async function generateResponse(prompt: string) {
  const resp = await fetch('https://api.freellmapi.com/v1/chat/completions', {
    method: 'POST',
    headers: { 
      'Content-Type': 'application/json', 
      Authorization: 'Bearer <your-token>' 
    },
    body: JSON.stringify({ 
      model: 'gpt-4o-mini', 
      messages: [{ role: 'user', content: prompt }] 
    })
  });

  // 429 errors are handled internally via automatic failover
  const data = await resp.json();
  return data.choices?.[0]?.message?.content;
}

generateResponse('Explain quantum tunneling');

If gpt-4o-mini hits its per-minute quota, the service automatically routes the request to the next available model (e.g., a free-tier alternative) and returns the answer transparently.

For debugging purposes, you can inspect the current penalty map:

import { getAllPenalties } from '@freellmapi/server/src/services/router';

console.log(getAllPenalties());
// => [{ modelDbId: 42, count: 3, penalty: 6 }, …]

Summary

  • Hard health signals – FreeLLMAPI treats 429 errors as severe infrastructure signals rather than transient faults.
  • Triple-weight penalties – Rate limits incur a -3 priority penalty compared to -1 for standard failures, stored in the rateLimitPenalties map.
  • Automatic decay – Penalties decrease by 1 every 2 minutes, allowing models to recover gradually without manual intervention.
  • Server-side failover – Clients receive seamless responses while the gateway handles model switching internally via fallback-loop.ts.

Frequently Asked Questions

How long does a model remain penalized after a 429 error?

The penalty decays automatically every 2 minutes (DECAY_INTERVAL_MS = 120000), reducing the penalty counter by 1 (DECAY_AMOUNT = 1) until it reaches zero. A model that receives a single 429 will return to full priority after approximately 6 minutes.

What is the difference between a 429 penalty and a regular failure?

Regular connection or timeout errors trigger recordModelFailure with a weight of -1, while rate limit errors invoke recordRateLimitHit, which applies a PENALTY_PER_429 = 3 weight. This heavier penalty ensures rate-limited models are deprioritized faster than those experiencing temporary network issues.

Which source files contain the rate limit handling logic?

The primary logic resides in server/src/services/router.ts (penalty storage and decay), server/src/lib/fallback-loop.ts (error classification and orchestration), and server/src/services/ratelimit.ts (detection helpers).

Does the client need to implement retry logic for 429 errors?

No. FreeLLMAPI handles all 429 detection, penalty application, and model failover internally. The client receives a standard 200 response from whichever healthy model ultimately fulfills the request, eliminating the need for client-side exponential backoff or retry loops.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →