How to Perform Automatic Failover with FreeLLMAPI: Complete Implementation Guide

FreeLLMAPI implements automatic failover through an intelligent request router that maintains a prioritized fallback chain of models, automatically retrying requests on alternative providers when encountering rate limits (429) or server errors (5xx) while applying temporary cooldown penalties to unhealthy keys.

FreeLLMAPI is an open-source aggregation layer that unifies access to multiple free-tier LLM providers through a single OpenAI-compatible API. Understanding how to leverage its automatic failover capabilities ensures your applications remain resilient when individual providers hit rate limits or experience outages.

Architecture of the Failover System

The failover logic resides in server/src/services/router.ts, where the Router orchestrates request distribution across a dynamically ordered fallback chain. When you send a request to /v1/chat/completions, the router evaluates multiple providers sequentially until successful or exhausted.

The Fallback Chain Strategy

The router constructs the fallback chain using selectable routing strategies including priority, balanced, and smartest. These strategies leverage the scoring engine in server/src/services/scoring.ts to rank models by reliability, speed, and intelligence metrics, ensuring optimal provider selection.

Health and Capability Validation

Before selecting a provider, the router performs three critical validations:

  • Key health: Periodic probes in server/src/services/health.ts classify keys as healthy, rate_limited, invalid, or error
  • Quota enforcement: SQLite-backed in-memory counters tracked in server/src/services/ratelimit.ts via functions like canMakeRequest() and isOnCooldown()
  • Capability matching: Model-level features (vision, tool-calling, context window) validated through modelsWithOverriddenField() and scopeAllows()

How the Failover Mechanism Works

When a request encounters failure, FreeLLMAPI's transparent failover mechanism activates without client intervention.

Trigger Conditions and Penalty Application

If a provider returns a 429 (rate limit) or 5xx error, the router immediately:

  1. Applies a cooldown to the offending key using acquireLease() and releaseLease() in ratelimit.ts
  2. Records the failure via recordRateLimitHit() or recordModelFailure() (lines 48-53 in router.ts)
  3. Increments temporary penalties that demote the model in the fallback chain

These penalties automatically decay over time via DECAY_INTERVAL_MS, allowing providers to recover once limits reset.

Retry Logic and Exhaustion Handling

The router attempts up to 20 retries (or until a wall-clock budget expires) against subsequent models in the chain. The summarizeExhaustion() function logs the complete attempt trail to the request table, enabling dashboard visualization of the exact failover path taken for each request.

Implementing Automatic Failover in Practice

No additional configuration is required to enable failover—it's transparent for all requests. However, you can observe and interact with the mechanism using the following patterns.

Standard API Request

The simplest implementation uses bearer token authentication. If the primary provider fails, the router automatically promotes the next available model:

curl https://api.freellmapi.com/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-<your-token>" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{ "role": "user", "content": "Explain automatic failover." }],
        "max_tokens": 500
      }'

Node.js Client Implementation

Use the official client package (@freellmapi/client) to handle failover transparently in JavaScript applications:

import { FreeLLMAPI } from '@freellmapi/client';

const api = new FreeLLMAPI({ token: 'freellmapi-<your-token>' });

async function generateContent() {
  const resp = await api.chatCompletions.create({
    model: 'gpt-4o-mini',
    messages: [{ role: 'user', content: 'What is automatic failover?' }],
    max_tokens: 500,
  });
  console.log(resp.choices[0].message.content);
}

generateContent().catch(console.error);

Debugging Failover Activity

To observe the failover chain in action, use verbose mode to inspect the x-free-llm-retry-attempts response header:

curl https://api.freellmapi.com/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-<token>" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{ "role": "user", "content": "Trigger failover test." }],
        "max_tokens": 200
      }' -v

The verbose output reveals the request-ID and the number of providers attempted before success. The dashboard UI (client/ directory) also displays the full attempt trail using data from the summarizeExhaustion function.

Core Components Reference

Component Source File Responsibility
Router server/src/services/router.ts Picks highest-priority healthy model, handles 429/5xx retries
Rate Limiter server/src/services/ratelimit.ts SQLite-backed RPM/RPD/TPM tracking, cooldown management
Health Monitor server/src/services/health.ts Periodic key status probes (healthy, rate_limited, invalid)
Scoring Engine server/src/services/scoring.ts Bandit algorithm calculating reliability/intelligence scores
Provider Adapters server/src/providers/*.ts Individual upstream LLM implementations (Google, Groq, Cerebras)

Summary

  • Automatic failover in FreeLLMAPI occurs in server/src/services/router.ts through a prioritized fallback chain of models.
  • The router checks health status, rate limits, and capabilities before each attempt using health.ts and ratelimit.ts.
  • 429 and 5xx errors trigger cooldowns via acquireLease() and penalize providers through recordModelFailure().
  • Up to 20 retry attempts are made against alternative providers; the client receives only the final successful response.
  • Request history and failover trails are accessible via the summarizeExhaustion function and dashboard UI.

Frequently Asked Questions

How many retry attempts does FreeLLMAPI perform during failover?

FreeLLMAPI attempts up to approximately 20 retries or continues until a wall-clock retry budget is exhausted. This behavior is hardcoded in the router logic within server/src/services/router.ts, ensuring comprehensive coverage of the fallback chain without indefinite loops.

Can I configure which providers participate in the automatic failover chain?

Yes, the fallback chain is constructed based on selectable routing strategies (priority, balanced, smartest) defined in the router configuration. The scoring.ts engine ranks available models, and you can influence selection by configuring model priorities, though the automatic demotion of unhealthy providers happens dynamically through the penalty system.

How does FreeLLMAPI handle rate limit cooldowns?

When a provider returns a 429 error, the router invokes recordRateLimitHit() and applies a cooldown using acquireLease() and releaseLease() functions in server/src/services/ratelimit.ts. This temporarily removes the provider from consideration; penalties decay automatically over time according to DECAY_INTERVAL_MS, allowing the provider to return to the rotation once limits reset.

Where can I view the failover attempt history for a specific request?

The summarizeExhaustion() function in router.ts logs all retry attempts to the SQLite request table. You can view the complete attempt trail—including which providers were tried and why they failed—through the dashboard UI in the client/ directory or by inspecting the x-free-llm-retry-attempts header in API responses when using verbose mode.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →