How FreeLLMAPI's Automatic Failover Mechanism Works: Multi-Provider Resilience Explained

FreeLLMAPI's automatic failover mechanism is a budget-aware, hedged retry loop that transparently routes requests across multiple LLM providers when upstream failures occur, utilizing circuit breakers, degraded mode logic, and real-time abort controllers to maximize availability.

FreeLLMAPI implements a sophisticated automatic failover system that keeps requests alive even when individual upstream providers become unresponsive or return errors. According to the tashfeenahmed/freellmapi source code, this mechanism centers on a shared fallback loop in server/src/lib/fallback-loop.ts that coordinates retries across providers like OpenAI, Anthropic, and Groq while enforcing strict time budgets and health checks.

Core Components of the Automatic Failover System

The automatic failover mechanism comprises five integrated subsystems that work together to ensure high availability.

Retry Budget and Request Hedging

At the heart of the system lies a wall-clock budget that limits how long the failover loop may attempt new providers. In server/src/lib/fallback-loop.ts, the getFallbackTimeBudgetMs() function enforces a default 45 second budget per request. If this budget expires mid-attempt, the in-flight request is aborted via AbortController through the abortInFlight() function, which produces a HedgeAbortError. This error is treated as a non-provider-health failure, allowing the loop to proceed to the next viable provider without penalizing the circuit breaker.

Circuit Breaker Protection

To prevent cascading failures, the failover mechanism implements a circuit breaker governed by the max_consecutive_upstream_fails setting. This counter tracks consecutive retryable upstream failures—specifically HTTP 429 (rate limit), 5xx server errors, and network timeouts. When the count exceeds the configured limit, the loop stops immediately and returns a 503 service_unavailable response with the upstream_unhealthy error code, protecting downstream clients from prolonged delays.

Degraded Mode State Management

When a significant portion of enabled providers become unhealthy, the system transitions to degraded mode via the logic in server/src/services/degradation.ts. The isDegraded() and updateDegradationState() functions monitor health ratios using thresholds like DEGRADED_HEALTHY_RATIO and DEGRADED_ENTRY_GRACE_MS. In degraded mode, the bandit-exploration floor is disabled, restricting the failover loop to only those providers with existing healthy scores rather than attempting experimental or uncertain routes.

Content Classification Failover

The system handles content filtering through classification-based failover implemented in server/src/routes/proxy.ts. When a provider returns a bare "safe" or "unsafe" token instead of a valid completion, the isUpstreamClassificationOutput() function treats this as an empty completion. The loop then throws an error that triggers an immediate retry with the next available provider, ensuring that content filtering on one provider does not block the entire request.

Observability and Tracing Headers

Every failover hop is tracked via custom headers injected by setFallbackHeaders() and formatAttemptDetail() in fallback-loop.ts. The X-Fallback-Attempts header counts total tries, while X-Fallback-Trail records the provider path taken, including failure reasons like openrouter/gpt-4 key3=rate_limited. When the EXPOSE_FALLBACK_DETAIL_HEADER flag is enabled, X-Fallback-Detail provides per-hop timing telemetry, giving developers complete visibility into the failover chain.

Execution Flow of the Failover Loop

The runFallbackLoop function orchestrates the automatic failover mechanism through the following deterministic sequence:

  1. Health Check Consultation: Upon entry, isDegraded() evaluates whether the system is in degraded mode; if true, only pre-scored healthy providers are considered for selection.

  2. Initial Attempt: The loop begins with attempt 0 against the selected provider.

  3. Budget Validation: Before each retry (attempt ≥ 1), the system verifies the remaining retry budget via getFallbackTimeBudgetMs(). Exhaustion triggers an immediate exhaustion response without starting new attempts.

  4. Mid-Flight Abort: During active attempts, if the budget expires, abortInFlight() cancels the request via AbortController, yielding a HedgeAbortError that triggers continuation to the next provider.

  5. Circuit Breaker Check: Consecutive upstream failures increment the max_consecutive_upstream_fails counter; exceeding the limit aborts the loop with a 503 error.

  6. Classification Handling: Returns of "safe" or "unsafe" tokens trigger immediate failover to the next provider without decrementing health scores.

  7. Header Injection: Upon success or final failure, the response includes X-Fallback-* headers documenting the full attempt history.

Implementing Automatic Failover in Practice

Developers interact with the automatic failover mechanism transparently—no client-side code is required to enable it. However, you can inspect and influence the behavior using standard HTTP headers.

Basic Request with Failover Inspection

The following Node.js example demonstrates how to observe the failover trail:

import fetch from 'node-fetch';

// The request will automatically hop between providers if any fail.
// No extra client-side code is needed – the proxy does the work.
const resp = await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${process.env.FREELLMAPI_KEY}`,
    'Content-Type': 'application/json',
    // Opt-in detailed hop timing (optional)
    'X-Expose-Fallback-Detail': '1',
  },
  body: JSON.stringify({
    model: 'gpt-4o',
    messages: [{role: 'user', content: 'Explain quantum tunnelling'}],
  }),
});

const data = await resp.json();

console.log('Response:', data.choices?.[0]?.message?.content);

// Inspect the fail-over trace supplied by the proxy
console.log('X-Fallback-Attempts:', resp.headers.get('X-Fallback-Attempts'));
console.log('X-Fallback-Trail:',   resp.headers.get('X-Fallback-Trail'));
console.log('X-Fallback-Detail:',  resp.headers.get('X-Fallback-Detail'));

When a provider times out or hits a quota, X-Fallback-Trail contains entries such as openrouter/gpt-4 key3=rate_limited, and the request automatically continues with the next viable provider.

Runtime Budget Overrides

You can adjust the default 45-second budget for individual requests using the X-Fallback-Settings header:

// Override the default 45s budget for this request only
const resp = await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${process.env.FREELLMAPI_KEY}`,
    'Content-Type': 'application/json',
    // Setting the runtime key (see docs/env/01-variables.md)
    'X-Fallback-Settings': JSON.stringify({fallback_time_budget_ms: 20000}),
  },
  body: JSON.stringify({model: 'gpt-4o', messages: [{role: 'user', content: 'Hi'}]}),
});

The fallback_time_budget_ms override is respected by getFallbackTimeBudgetMs() inside server/src/lib/fallback-loop.ts.

Summary

  • Budget-Aware Hedging: The automatic failover mechanism enforces a strict 45-second wall-clock budget via getFallbackTimeBudgetMs(), aborting in-flight requests with AbortController when timeouts occur mid-attempt.

  • Circuit Breaker Safety: Consecutive upstream failures tracked by max_consecutive_upstream_fails trigger a 503 upstream_unhealthy response to prevent cascade failures.

  • Degraded Mode Isolation: The isDegraded() state machine in degradation.ts restricts routing to proven healthy providers when system health drops below DEGRADED_HEALTHY_RATIO.

  • Classification Handling: Content filter responses ("safe"/"unsafe") from isUpstreamClassificationOutput() trigger immediate provider switching without health penalties.

  • Full Observability: Headers like X-Fallback-Trail and X-Fallback-Detail provide complete visibility into provider selection and timing across the failover chain.

Frequently Asked Questions

What triggers the automatic failover mechanism in FreeLLMAPI?

The mechanism activates when an upstream provider returns specific error conditions, including HTTP 429 rate limits, 5xx server errors, network timeouts, or content classification tokens ("safe" or "unsafe"). Additionally, if a request exceeds internal timing thresholds or the AbortController triggers a HedgeAbortError due to budget exhaustion mid-flight, the system automatically selects the next available provider from the healthy pool.

How does the retry budget prevent infinite failover loops?

The retry budget enforces a hard 45-second limit (configurable via FALLBACK_TIME_BUDGET_MS) managed by getFallbackTimeBudgetMs() in fallback-loop.ts. Before initiating any retry attempt, the loop checks remaining budget; if exhausted, it returns an exhaustion response rather than starting a new attempt. This guarantees that requests cannot spawn infinite retries regardless of provider failure rates.

What is the difference between a HedgeAbortError and a circuit breaker failure?

A HedgeAbortError occurs when the time budget expires during an in-flight request, triggering abortInFlight() to cancel the connection; this is treated as a timeout failure, not a provider health issue. Conversely, a circuit breaker failure occurs when max_consecutive_upstream_fails is exceeded due to retryable errors (429, 5xx), causing the loop to return a 503 upstream_unhealthy error and halt further attempts immediately.

How can developers monitor which providers were attempted during failover?

Developers can inspect the X-Fallback-Attempts, X-Fallback-Trail, and X-Fallback-Detail response headers injected by setFallbackHeaders() in fallback-loop.ts. The X-Fallback-Trail header lists each attempted provider with failure reasons (e.g., rate_limited), while X-Fallback-Detail provides granular timing data when the EXPOSE_FALLBACK_DETAIL_HEADER flag is enabled via the X-Expose-Fallback-Detail: 1 request header.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →