How FreeLLMAPI Handles Automatic Failover for LLM Providers: A Deep Dive into the Retry Engine

FreeLLMAPI guarantees every request reaches a working LLM through a single shared retry/failover loop that automatically routes around failed providers, implements intelligent cooldowns, and returns structured error responses when the provider pool is exhausted.

FreeLLMAPI's automatic failover system ensures high availability by transparently switching between LLM providers when keys fail, rate limits hit, or models become unavailable. This article examines the core mechanisms, code paths, and configuration options that power this resilience layer according to the tashfeenahmed/freellmapi source code.

The Core Failover Loop Architecture

All OpenAI-compatible surfaces—including /chat/completions, /v1/completions, and Anthropic endpoints—delegate to a unified retry engine. In server/src/lib/fallback-loop.ts, the runFallbackLoop function orchestrates every failover attempt through a consistent six-step pipeline.

Three Core Mechanisms Protecting Request Reliability

Mechanism Purpose Location
Retry budget & hop limit Caps failover hops at FALLBACK_MAX_RETRIES = 20 and enforces a wall-clock budget (fallback_time_budget_ms, default 45s) fallback-loop.ts 58-66
Cooldown & model-level benching Records per-key failures and applies timed cooldowns; model-wide failure windows (15min, 3 failures) bench models across all keys fallback-loop.ts 81-105, 42-48
Routing & exhaustion handling Builds structured ExhaustionBody responses preserving exact failure reasons when the provider pool exhausts fallback-loop.ts 676-785, 272-292

Step-by-Step Automatic Failover Flow

The automatic failover for LLM providers in FreeLLMAPI follows a deterministic sequence from candidate selection through final response or exhaustion.

1. Candidate Selection via Dynamic Scoring

The loop calls hooks.route(attempt) to query services/router.ts. The router scores each enabled model/key pair using a bandit scorer with dynamic penalties, respecting the fallback_config chain or a profile-specific override.

// Conceptual flow from fallback-loop.ts
const candidate = await hooks.route({
  attempt,
  skipKeys,      // Keys currently on cooldown
  skipModels,    // Models benched due to failures
  preferredProvider // Optional affinity hint
});

2. Request Dispatch and Success Recording

The surface dispatches to the selected provider. On success, recordUpstreamSuccess immediately clears any active cooldowns for that key/model pair.

3. Error Classification

When providers return errors, classifyAttemptError in server/src/lib/error-classify.ts maps raw responses to AttemptErrorClass enums: auth, rate_limited, model_not_found, timeout, empty_completion, etc.

4. Failure Bookkeeping and Cooldown Application

Auth failures (401) trigger recordAuthFailure, which:

  • Adds the key to skipKeys
  • Benches the key for 5 minutes
  • Launches an immediate re-validation job

Retryable failures (429, 5xx, timeouts, empty completions) invoke recordRetryableFailure, which calls:

  • cooldownForError to apply error-specific cooldown durations
  • noteModelFailure to update model-level failure windows
// From fallback-loop.ts - model-level benching logic
if (modelFailureWindow.exceedsThreshold(3, '15m')) {
  skipModels.add(modelId);  // Bench across ALL keys
  cooldownUntil.set(modelId, Date.now() + MODEL_BENCH_DURATION);
}

5. Budget Verification Before Retry

Before each new attempt, the loop enforces guardrails:

  • Wall-clock budget: getFallbackTimeBudgetMs() checks elapsed time
  • Circuit breaker: breakerLimit prevents runaway cascades
if (elapsed > timeBudget || attempt >= FALLBACK_MAX_RETRIES) {
  throw exhaustedRetryError(buildExhaustionBody(attemptTrail));
}

6. Loop Termination or Exhaustion Response

Successful responses return immediately. When limits exhaust, exhaustedRetryError constructs a standards-compliant JSON payload:

{
  "status": 429,
  "type": "rate_limit_error",
  "message": "All models rate-limited. Trail: openai/gpt-4o (key1): 429; anthropic/claude-3 (key2): 429"
}

Practical Usage: Automatic Failover in Action

Basic Request (Failover is Invisible)

No configuration changes are needed for automatic failover for LLM providers to activate:

curl -X POST https://api.freellmapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $YOUR_FREELLMAPI_KEY" \
  -d '{
        "model": "gpt-4o",
        "messages": [{"role":"user","content":"What is the capital of France?"}]
      }'

If the primary OpenAI key fails, FreeLLMAPI automatically retries Anthropic, Azure OpenAI, or other configured providers until success or exhaustion.

Inspecting Failover Diagnostics

Enable detailed headers by setting EXPOSE_FALLBACK_DETAIL_HEADER=1:

curl -i \
  -H "Authorization: Bearer $YOUR_FREELLMAPI_KEY" \
  -H "Accept: application/json" \
  -X POST https://api.freellmapi.com/v1/chat/completions \
  -d '{"model":"gpt-4o","messages":[{"role":"user","content":"Hello"}]}'

Response headers from setFallbackHeaders (560-575):

Header Example Value Meaning
X-Fallback-Attempts 2 Total hops before success
X-Fallback-Trail openai/modelX key1=rate_limited; anthropic/modelY key2=success Provider/model/key and outcome per hop
X-Fallback-Detail openai/modelX key1=rate_limited t=0+120ms msg=Quota exceeded; ... Timing and redacted error summaries

Bypassing Automatic Failover for Testing

To force a specific provider and disable automatic routing:


# Via settings API

curl -X PUT https://api.freellmapi.com/fallback \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{"fallback_time_budget_ms": 0}'

With zero budget, runFallbackLoop completes exactly one attempt, surfacing provider-specific errors directly.

Key Files in the Automatic Failover System

File Responsibility
server/src/lib/fallback-loop.ts Core engine: runFallbackLoop, cooldown logic, exhaustion handling
server/src/routes/fallback.ts HTTP configuration endpoints, routing score exposure, token diagnostics
server/src/services/router.ts Candidate scoring, dynamic penalties, profile-aware selection
server/src/services/ratelimit.ts Cooldown state: setCooldown, getCooldownDecisionForLimit
server/src/services/model-retirement.ts Upstream model retirement detection and chain removal
server/src/lib/error-classify.ts Provider error → internal AttemptErrorClass mapping
server/src/lib/error-redaction.ts PII/security redaction before logging/response

Cooldown Strategy and Failure Windows

FreeLLMAPI implements multi-layer cooldowns to prevent thundering herds:

  • Key-level cooldowns: 5 minutes for auth failures; variable for rate limits based on Retry-After headers
  • Model-level benching: 15-minute window, 3 failures trigger removal from all keys' candidate pools
  • Dynamic backoff: Exponential increase for repeated 429 responses from the same provider

The noteModelFailure function in fallback-loop.ts 81-105 maintains sliding windows per model across all provider keys, ensuring systemic model issues (e.g., deprecated endpoints) trigger rapid chain-wide exclusion.

Summary

  • FreeLLMAPI's automatic failover for LLM providers operates through a single shared loop in fallback-loop.ts, ensuring identical behavior across all API surfaces
  • Dynamic router scoring in services/router.ts selects optimal candidates using bandit algorithms with penalty adjustments
  • Multi-layer cooldowns protect against retry storms: per-key auth timeouts, error-type-specific delays, and model-wide failure windows
  • Configurable guardrails—20 hop limit and 45-second default budget—prevent unbounded retries
  • Structured exhaustion responses preserve failure context for client debugging when all providers fail
  • Diagnostic headers expose hop trails without exposing raw credentials or full error messages

Frequently Asked Questions

What triggers automatic failover in FreeLLMAPI?

Automatic failover activates on any classified error: authentication failures (401), rate limits (429), server errors (5xx), timeouts, or empty completions. The classifyAttemptError function in server/src/lib/error-classify.ts determines retry eligibility. Each error type maps to specific cooldown behavior—auth failures immediately bench keys for 5 minutes, while rate limits respect provider Retry-After headers.

How can I monitor which providers handled my requests?

Set EXPOSE_FALLBACK_DETAIL_HEADER=1 in server environment variables. Responses then include X-Fallback-Attempts, X-Fallback-Trail, and X-Fallback-Detail headers generated by setFallbackHeaders in fallback-loop.ts. These reveal the hop count, provider sequence, and error classifications without exposing actual API keys or sensitive error details.

Does FreeLLMAPI charge for failed provider attempts?

The source code tracks token consumption and attempt counts separately. Failed attempts increment counters but typically do not bill for non-existent completions. The fallback_time_budget_ms limit prevents runaway costs from slow-failing providers. Exact billing behavior depends on your deployment's metering configuration in server/src/services/.

Can I customize the automatic failover chain per API key?

Yes. The fallback_config field in profiles allows per-key override of the default provider chain. The hooks.route call accepts a preferredProvider hint, and the router in services/router.ts applies profile-specific penalties. Configure via PUT /fallback endpoints defined in server/src/routes/fallback.ts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →