# How FreeLLMAPI's Automatic Failover Mechanism Works: Multi-Provider Resilience Explained

> Discover FreeLLMAPI's automatic failover mechanism. Learn how this budget-aware system routes requests across multiple LLM providers using circuit breakers and real-time abort controllers to ensure maximum availability.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-09-02

---

**FreeLLMAPI's automatic failover mechanism is a budget-aware, hedged retry loop that transparently routes requests across multiple LLM providers when upstream failures occur, utilizing circuit breakers, degraded mode logic, and real-time abort controllers to maximize availability.**

FreeLLMAPI implements a sophisticated automatic failover system that keeps requests alive even when individual upstream providers become unresponsive or return errors. According to the tashfeenahmed/freellmapi source code, this mechanism centers on a shared fallback loop in [`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts) that coordinates retries across providers like OpenAI, Anthropic, and Groq while enforcing strict time budgets and health checks.

## Core Components of the Automatic Failover System

The automatic failover mechanism comprises five integrated subsystems that work together to ensure high availability.

### Retry Budget and Request Hedging

At the heart of the system lies a **wall-clock budget** that limits how long the failover loop may attempt new providers. In [`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts), the `getFallbackTimeBudgetMs()` function enforces a default **45 second** budget per request. If this budget expires **mid-attempt**, the in-flight request is aborted via `AbortController` through the `abortInFlight()` function, which produces a `HedgeAbortError`. This error is treated as a non-provider-health failure, allowing the loop to proceed to the next viable provider without penalizing the circuit breaker.

### Circuit Breaker Protection

To prevent cascading failures, the failover mechanism implements a circuit breaker governed by the `max_consecutive_upstream_fails` setting. This counter tracks consecutive retryable upstream failures—specifically HTTP 429 (rate limit), 5xx server errors, and network timeouts. When the count exceeds the configured limit, the loop stops immediately and returns a **503 service_unavailable** response with the `upstream_unhealthy` error code, protecting downstream clients from prolonged delays.

### Degraded Mode State Management

When a significant portion of enabled providers become unhealthy, the system transitions to **degraded mode** via the logic in [`server/src/services/degradation.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/degradation.ts). The `isDegraded()` and `updateDegradationState()` functions monitor health ratios using thresholds like `DEGRADED_HEALTHY_RATIO` and `DEGRADED_ENTRY_GRACE_MS`. In degraded mode, the bandit-exploration floor is disabled, restricting the failover loop to only those providers with existing healthy scores rather than attempting experimental or uncertain routes.

### Content Classification Failover

The system handles content filtering through **classification-based failover** implemented in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts). When a provider returns a bare `"safe"` or `"unsafe"` token instead of a valid completion, the `isUpstreamClassificationOutput()` function treats this as an empty completion. The loop then throws an error that triggers an immediate retry with the next available provider, ensuring that content filtering on one provider does not block the entire request.

### Observability and Tracing Headers

Every failover hop is tracked via custom headers injected by `setFallbackHeaders()` and `formatAttemptDetail()` in [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts). The `X-Fallback-Attempts` header counts total tries, while `X-Fallback-Trail` records the provider path taken, including failure reasons like `openrouter/gpt-4 key3=rate_limited`. When the `EXPOSE_FALLBACK_DETAIL_HEADER` flag is enabled, `X-Fallback-Detail` provides per-hop timing telemetry, giving developers complete visibility into the failover chain.

## Execution Flow of the Failover Loop

The `runFallbackLoop` function orchestrates the automatic failover mechanism through the following deterministic sequence:

1. **Health Check Consultation**: Upon entry, `isDegraded()` evaluates whether the system is in degraded mode; if true, only pre-scored healthy providers are considered for selection.

2. **Initial Attempt**: The loop begins with attempt 0 against the selected provider.

3. **Budget Validation**: Before each retry (attempt ≥ 1), the system verifies the remaining retry budget via `getFallbackTimeBudgetMs()`. Exhaustion triggers an immediate exhaustion response without starting new attempts.

4. **Mid-Flight Abort**: During active attempts, if the budget expires, `abortInFlight()` cancels the request via `AbortController`, yielding a `HedgeAbortError` that triggers continuation to the next provider.

5. **Circuit Breaker Check**: Consecutive upstream failures increment the `max_consecutive_upstream_fails` counter; exceeding the limit aborts the loop with a 503 error.

6. **Classification Handling**: Returns of `"safe"` or `"unsafe"` tokens trigger immediate failover to the next provider without decrementing health scores.

7. **Header Injection**: Upon success or final failure, the response includes `X-Fallback-*` headers documenting the full attempt history.

## Implementing Automatic Failover in Practice

Developers interact with the automatic failover mechanism transparently—no client-side code is required to enable it. However, you can inspect and influence the behavior using standard HTTP headers.

### Basic Request with Failover Inspection

The following Node.js example demonstrates how to observe the failover trail:

```typescript
import fetch from 'node-fetch';

// The request will automatically hop between providers if any fail.
// No extra client-side code is needed – the proxy does the work.
const resp = await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${process.env.FREELLMAPI_KEY}`,
    'Content-Type': 'application/json',
    // Opt-in detailed hop timing (optional)
    'X-Expose-Fallback-Detail': '1',
  },
  body: JSON.stringify({
    model: 'gpt-4o',
    messages: [{role: 'user', content: 'Explain quantum tunnelling'}],
  }),
});

const data = await resp.json();

console.log('Response:', data.choices?.[0]?.message?.content);

// Inspect the fail-over trace supplied by the proxy
console.log('X-Fallback-Attempts:', resp.headers.get('X-Fallback-Attempts'));
console.log('X-Fallback-Trail:',   resp.headers.get('X-Fallback-Trail'));
console.log('X-Fallback-Detail:',  resp.headers.get('X-Fallback-Detail'));

```

When a provider times out or hits a quota, `X-Fallback-Trail` contains entries such as `openrouter/gpt-4 key3=rate_limited`, and the request automatically continues with the next viable provider.

### Runtime Budget Overrides

You can adjust the default 45-second budget for individual requests using the `X-Fallback-Settings` header:

```typescript
// Override the default 45s budget for this request only
const resp = await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${process.env.FREELLMAPI_KEY}`,
    'Content-Type': 'application/json',
    // Setting the runtime key (see docs/env/01-variables.md)
    'X-Fallback-Settings': JSON.stringify({fallback_time_budget_ms: 20000}),
  },
  body: JSON.stringify({model: 'gpt-4o', messages: [{role: 'user', content: 'Hi'}]}),
});

```

The `fallback_time_budget_ms` override is respected by `getFallbackTimeBudgetMs()` inside [`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts).

## Summary

- **Budget-Aware Hedging**: The automatic failover mechanism enforces a strict 45-second wall-clock budget via `getFallbackTimeBudgetMs()`, aborting in-flight requests with `AbortController` when timeouts occur mid-attempt.

- **Circuit Breaker Safety**: Consecutive upstream failures tracked by `max_consecutive_upstream_fails` trigger a 503 `upstream_unhealthy` response to prevent cascade failures.

- **Degraded Mode Isolation**: The `isDegraded()` state machine in [`degradation.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/degradation.ts) restricts routing to proven healthy providers when system health drops below `DEGRADED_HEALTHY_RATIO`.

- **Classification Handling**: Content filter responses (`"safe"`/`"unsafe"`) from `isUpstreamClassificationOutput()` trigger immediate provider switching without health penalties.

- **Full Observability**: Headers like `X-Fallback-Trail` and `X-Fallback-Detail` provide complete visibility into provider selection and timing across the failover chain.

## Frequently Asked Questions

### What triggers the automatic failover mechanism in FreeLLMAPI?

The mechanism activates when an upstream provider returns specific error conditions, including HTTP 429 rate limits, 5xx server errors, network timeouts, or content classification tokens (`"safe"` or `"unsafe"`). Additionally, if a request exceeds internal timing thresholds or the `AbortController` triggers a `HedgeAbortError` due to budget exhaustion mid-flight, the system automatically selects the next available provider from the healthy pool.

### How does the retry budget prevent infinite failover loops?

The retry budget enforces a hard **45-second limit** (configurable via `FALLBACK_TIME_BUDGET_MS`) managed by `getFallbackTimeBudgetMs()` in [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts). Before initiating any retry attempt, the loop checks remaining budget; if exhausted, it returns an exhaustion response rather than starting a new attempt. This guarantees that requests cannot spawn infinite retries regardless of provider failure rates.

### What is the difference between a HedgeAbortError and a circuit breaker failure?

A **HedgeAbortError** occurs when the time budget expires during an in-flight request, triggering `abortInFlight()` to cancel the connection; this is treated as a timeout failure, not a provider health issue. Conversely, a **circuit breaker failure** occurs when `max_consecutive_upstream_fails` is exceeded due to retryable errors (429, 5xx), causing the loop to return a 503 `upstream_unhealthy` error and halt further attempts immediately.

### How can developers monitor which providers were attempted during failover?

Developers can inspect the `X-Fallback-Attempts`, `X-Fallback-Trail`, and `X-Fallback-Detail` response headers injected by `setFallbackHeaders()` in [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts). The `X-Fallback-Trail` header lists each attempted provider with failure reasons (e.g., `rate_limited`), while `X-Fallback-Detail` provides granular timing data when the `EXPOSE_FALLBACK_DETAIL_HEADER` flag is enabled via the `X-Expose-Fallback-Detail: 1` request header.