# How to Perform Automatic Failover with FreeLLMAPI: Complete Implementation Guide

> Learn how to perform automatic failover with FreeLLMAPI. Discover how its intelligent router maintains a fallback chain, retrying requests on alternative providers for seamless operation.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-08-29

---

**FreeLLMAPI implements automatic failover through an intelligent request router that maintains a prioritized fallback chain of models, automatically retrying requests on alternative providers when encountering rate limits (429) or server errors (5xx) while applying temporary cooldown penalties to unhealthy keys.**

FreeLLMAPI is an open-source aggregation layer that unifies access to multiple free-tier LLM providers through a single OpenAI-compatible API. Understanding how to leverage its **automatic failover** capabilities ensures your applications remain resilient when individual providers hit rate limits or experience outages.

## Architecture of the Failover System

The failover logic resides in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts), where the **Router** orchestrates request distribution across a dynamically ordered fallback chain. When you send a request to `/v1/chat/completions`, the router evaluates multiple providers sequentially until successful or exhausted.

### The Fallback Chain Strategy

The router constructs the fallback chain using selectable routing strategies including **priority**, **balanced**, and **smartest**. These strategies leverage the scoring engine in [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts) to rank models by reliability, speed, and intelligence metrics, ensuring optimal provider selection.

### Health and Capability Validation

Before selecting a provider, the router performs three critical validations:

- **Key health**: Periodic probes in [`server/src/services/health.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/health.ts) classify keys as `healthy`, `rate_limited`, `invalid`, or `error`
- **Quota enforcement**: SQLite-backed in-memory counters tracked in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) via functions like `canMakeRequest()` and `isOnCooldown()`
- **Capability matching**: Model-level features (vision, tool-calling, context window) validated through `modelsWithOverriddenField()` and `scopeAllows()`

## How the Failover Mechanism Works

When a request encounters failure, FreeLLMAPI's transparent failover mechanism activates without client intervention.

### Trigger Conditions and Penalty Application

If a provider returns a **429** (rate limit) or **5xx** error, the router immediately:

1. Applies a cooldown to the offending key using `acquireLease()` and `releaseLease()` in [`ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/ratelimit.ts)
2. Records the failure via `recordRateLimitHit()` or `recordModelFailure()` (lines 48-53 in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts))
3. Increments temporary penalties that demote the model in the fallback chain

These penalties automatically decay over time via `DECAY_INTERVAL_MS`, allowing providers to recover once limits reset.

### Retry Logic and Exhaustion Handling

The router attempts **up to 20 retries** (or until a wall-clock budget expires) against subsequent models in the chain. The `summarizeExhaustion()` function logs the complete attempt trail to the request table, enabling dashboard visualization of the exact failover path taken for each request.

## Implementing Automatic Failover in Practice

No additional configuration is required to enable failover—it's transparent for all requests. However, you can observe and interact with the mechanism using the following patterns.

### Standard API Request

The simplest implementation uses bearer token authentication. If the primary provider fails, the router automatically promotes the next available model:

```bash
curl https://api.freellmapi.com/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-<your-token>" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{ "role": "user", "content": "Explain automatic failover." }],
        "max_tokens": 500
      }'

```

### Node.js Client Implementation

Use the official client package (`@freellmapi/client`) to handle failover transparently in JavaScript applications:

```javascript
import { FreeLLMAPI } from '@freellmapi/client';

const api = new FreeLLMAPI({ token: 'freellmapi-<your-token>' });

async function generateContent() {
  const resp = await api.chatCompletions.create({
    model: 'gpt-4o-mini',
    messages: [{ role: 'user', content: 'What is automatic failover?' }],
    max_tokens: 500,
  });
  console.log(resp.choices[0].message.content);
}

generateContent().catch(console.error);

```

### Debugging Failover Activity

To observe the failover chain in action, use verbose mode to inspect the `x-free-llm-retry-attempts` response header:

```bash
curl https://api.freellmapi.com/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-<token>" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{ "role": "user", "content": "Trigger failover test." }],
        "max_tokens": 200
      }' -v

```

The verbose output reveals the request-ID and the number of providers attempted before success. The dashboard UI (`client/` directory) also displays the full attempt trail using data from the `summarizeExhaustion` function.

## Core Components Reference

| Component | Source File | Responsibility |
|-----------|-------------|----------------|
| **Router** | [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) | Picks highest-priority healthy model, handles 429/5xx retries |
| **Rate Limiter** | [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) | SQLite-backed RPM/RPD/TPM tracking, cooldown management |
| **Health Monitor** | [`server/src/services/health.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/health.ts) | Periodic key status probes (healthy, rate_limited, invalid) |
| **Scoring Engine** | [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts) | Bandit algorithm calculating reliability/intelligence scores |
| **Provider Adapters** | `server/src/providers/*.ts` | Individual upstream LLM implementations (Google, Groq, Cerebras) |

## Summary

- **Automatic failover** in FreeLLMAPI occurs in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) through a prioritized fallback chain of models.
- The router checks **health status**, **rate limits**, and **capabilities** before each attempt using [`health.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/health.ts) and [`ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/ratelimit.ts).
- **429 and 5xx errors** trigger cooldowns via `acquireLease()` and penalize providers through `recordModelFailure()`.
- **Up to 20 retry attempts** are made against alternative providers; the client receives only the final successful response.
- Request history and failover trails are accessible via the `summarizeExhaustion` function and dashboard UI.

## Frequently Asked Questions

### How many retry attempts does FreeLLMAPI perform during failover?

FreeLLMAPI attempts **up to approximately 20 retries** or continues until a wall-clock retry budget is exhausted. This behavior is hardcoded in the router logic within [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts), ensuring comprehensive coverage of the fallback chain without indefinite loops.

### Can I configure which providers participate in the automatic failover chain?

Yes, the fallback chain is constructed based on selectable **routing strategies** (priority, balanced, smartest) defined in the router configuration. The [`scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/scoring.ts) engine ranks available models, and you can influence selection by configuring model priorities, though the automatic demotion of unhealthy providers happens dynamically through the penalty system.

### How does FreeLLMAPI handle rate limit cooldowns?

When a provider returns a 429 error, the router invokes `recordRateLimitHit()` and applies a cooldown using `acquireLease()` and `releaseLease()` functions in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts). This temporarily removes the provider from consideration; penalties decay automatically over time according to `DECAY_INTERVAL_MS`, allowing the provider to return to the rotation once limits reset.

### Where can I view the failover attempt history for a specific request?

The `summarizeExhaustion()` function in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) logs all retry attempts to the SQLite request table. You can view the complete attempt trail—including which providers were tried and why they failed—through the dashboard UI in the `client/` directory or by inspecting the `x-free-llm-retry-attempts` header in API responses when using verbose mode.