# How FreeLLMAPI's Smart Routing and Failover Mechanism Works: A Technical Deep Dive

> Discover how FreeLLMAPI's smart routing and failover mechanism ensures reliable LLM access. Learn about bandit-style routing, Thompson sampling, and automatic retries for optimal performance.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-08-31

---

**FreeLLMAPI routes every request through a bandit-style router that dynamically scores every enabled model/key pair across reliability, speed, and intelligence axes, then selects the highest-scoring candidate using Thompson sampling while a surrounding fallback loop automatically retries with alternative providers when individual keys fail or hit quota limits.**

FreeLLMAPI is an open-source LLM gateway developed by tashfeenahmed/freellmapi designed to distribute traffic across multiple providers without vendor lock-in. Understanding how FreeLLMAPI's smart routing and failover mechanism works reveals a sophisticated, self-optimizing system that continuously learns provider performance and gracefully handles outages through automatic demotion and retry logic.

## The Eight-Step Routing Pipeline

The routing engine in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) executes a deterministic pipeline for every request. Each step refines the candidate pool until a single optimal provider is selected.

### 1. Building the Candidate Chain

The process begins in `resolveRoutingChain` and `orderChain` (lines ≈ 1881‑1890). The system queries the SQLite `models` table and joins entries with their associated API-key rows. Tier-fallback information stored in `match_tier` is attached to each candidate, creating an ordered list of potentially usable upstreams.

### 2. Gathering Per-Model Statistics

Recent request outcomes are cached in a decay-weighted `statsCache` via `refreshStatsCache` (lines ≈ 1850‑1870). This cache tracks successes, failures, token-throughput, first-byte latency, and monthly token usage for every model/key combination. The decay weighting ensures recent behavior influences routing decisions more heavily than historical data.

### 3. Computing Multi-Axis Quality Scores

For each candidate, `scoreChainEntry` (lines ≈ 5540‑5570) calculates three independent quality axes:

- **Reliability** – A Beta posterior distribution derived from successes versus failures, optionally seeded with community priors.
- **Speed** – Tokens-per-second throughput combined with first-byte latency into a composite performance score.
- **Intelligence** – A static rank derived from the catalog-provided tier and intelligence metadata.

### 4. Enforcing Budget Headroom Guardrails

Before final selection, the router applies financial and rate-limit protections. The functions `headroomFactor` and `rateWindowHeadroomFactor` (lines ≈ 5790‑5805) calculate multiplicative penalties based on remaining monthly token quotas and per-minute rate-limit windows. When a model approaches its budget ceiling, its score is proportionally reduced to prevent overages.

### 5. Applying Rate Limit and Failure Penalties

Transient failures trigger immediate score adjustments. The system maintains a global `rateLimitPenalties` map using `recordRateLimitHit` and `recordModelFailure` (lines ≈ 5660‑5665). Penalties decay over time and are applied as a multiplicative `rateLimitFactor` during scoring, effectively quarantining unstable providers until they recover.

### 6. Combining Scores with Strategy Weights

The active **routing strategy**—`priority`, `balanced`, `smartest`, or custom—determines a weight vector (`RoutingWeights`). The `combineScore` function merges the three quality axes using these weights. Operators can override weights for specific models via `applyModelWeightOverride` (lines ≈ 6060‑6070), enabling per-provider fine-tuning without code changes.

### 7. Selecting the Optimal Route

The candidate chain undergoes final sorting in `orderChain`: first by `tier`, then by computed `score`, then by manual `priority`. The top entry is returned as a `RouteResult` object containing the provider name, model ID, and decrypted API key. This occurs within `routeRequest` (lines ≈ 1881‑1900).

### 8. Executing the Failover Loop

If the selected key cannot be used—due to cooldown, quota exhaustion, or decryption errors—the router throws a `RouteError`. The [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts) wrapper (lines ≈ 800‑825) catches this exception, records the failure via `recordRateLimitHit` or `recordModelFailure`, releases any acquired lease, and re-invokes `routeRequest` with the failing key and model added to exclusion lists (`skipKeys`, `skipModels`). This repeats until a usable upstream is found or the chain is exhausted.

## Advanced Bandit-Style Optimizations

Beyond the core pipeline, FreeLLMAPI implements several bandit-algorithm enhancements to balance exploration against exploitation.

### Thompson Sampling for Exploration

When the routing strategy is not set to `priority`, the router samples from the Beta posterior using `sampleBeta` rather than using the expected reliability value. This injects intrinsic exploration, allowing potentially better-performing providers to prove themselves without a separate exploration phase.

### Anti-Starvation Protection

An optional `EXPLORE_CHANCE` flag ensures completely unmeasured models (cold starts) receive a guaranteed minimum selection probability. This prevents permanent starvation of new providers while the system gathers initial performance statistics.

### Time-of-Day Performance Tuning

Operators can enable `getPeakHoursConfig` to apply time-of-day re-weighting. During configured windows, the system temporarily tilts scores toward cheaper or higher-throughput providers, optimizing for cost or latency based on predictable traffic patterns.

### Community-Driven Reliability Priors

When enabled, the system merges aggregated reliability data from other FreeLLMAPI installations into the Beta posterior. This gives newly added models a sensible starting reliability score based on community experience, accelerating convergence to optimal routing.

## Code Implementation Examples

### Basic Request Routing

The simplest integration calls `routeRequest` directly from server route handlers:

```typescript
import { routeRequest } from '../services/router.js';

// Estimate of output tokens the client will request (default 1000)
const estimatedTokens = 1200;

// Call the router; it returns the selected provider, model, and decrypted API key.
const result = routeRequest(estimatedTokens);

// Result fields you can forward to the upstream provider:
console.log(`Using ${result.provider.name} – model ${result.modelId}`);
console.log(`API key (redacted): ${result.apiKey.slice(0, 4)}…`);

```

### Automatic Failover Implementation

For production resilience, wrap calls in the built-in fallback loop:

```typescript
import { routeRequest } from '../services/router.js';
import { invokeProvider } from '../providers/base.js';
import { fallbackLoop } from '../lib/fallback-loop.js';

async function handleChat(request) {
  return await fallbackLoop(async (skipKeys, skipModels) => {
    // `routeRequest` will throw a RouteError if the whole chain is exhausted.
    const route = routeRequest(/*estimatedTokens=*/1000, skipKeys, undefined, false, false, skipModels);
    // Perform the actual upstream call.
    const response = await invokeProvider(route.provider, {
      model: route.modelId,
      apiKey: route.apiKey,
      // …other request fields…
    });
    // Record success so the model penalty decays.
    recordSuccess(route.modelDbId);
    return response;
  });
}

```

### Dynamic Strategy Adjustment

Change routing behavior at runtime without restarting the server:

```typescript
import { setRoutingStrategy, getRoutingStrategy } from '../services/router.js';

// Switch to the “smartest” strategy (reliability‑driven)
setRoutingStrategy('smartest');
console.log('Current strategy:', getRoutingStrategy());

```

## Key Configuration and Source Files

| File | Role |
|------|------|
| [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) | Main routing engine implementing chain building, scoring, and strategy logic. |
| [`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts) | Wraps `routeRequest` with retry logic that excludes failed keys and records penalties. |
| [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts) | Defines reliability, speed, and intelligence scoring functions and the `combineScore` combinator. |
| [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) | Implements per-key and per-model rate-limit checks, lease acquisition, and cooldown handling. |
| [`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts) | Calculates headroom factors based on monthly token budgets and usable key counts. |

All routing parameters, including strategy weights and peak-hours configuration, are persisted in the SQLite `settings` table, enabling runtime tuning without redeployment.

## Summary

- **FreeLLMAPI** uses a **bandit-style router** that scores every model/key pair across reliability, speed, and intelligence axes.
- **Thompson sampling** provides intrinsic exploration, while configurable **routing strategies** allow operators to prioritize cost, speed, or reliability.
- **Headroom guardrails** and **rate-limit penalties** automatically protect budgets and quarantine failing providers.
- A **fallback loop** in [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts) catches routing errors and retries with excluded failed keys until a healthy upstream is found.
- All configuration is stored in SQLite and adjustable at runtime via `setRoutingStrategy` and related APIs.

## Frequently Asked Questions

### What routing strategies does FreeLLMAPI support?

FreeLLMAPI supports multiple strategies including `priority` (manual ordering), `balanced` (even distribution), and `smartest` (reliability-driven). Each strategy applies a distinct weight vector in `combineScore` to emphasize different axes—such as cost versus latency—when calculating final route scores.

### How does FreeLLMAPI handle rate limit errors from upstream providers?

When a provider returns a 429 or connection failure, the system invokes `recordRateLimitHit` or `recordModelFailure` to inject a multiplicative penalty into the `rateLimitPenalties` map. This penalty decays over time and reduces the provider’s selection probability. Simultaneously, the [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts) wrapper catches the error and retries with the failing key excluded.

### Can FreeLLMAPI route to unmeasured or newly added models?

Yes. The optional `EXPLORE_CHANCE` flag guarantees that unmeasured models receive a minimum selection probability, preventing starvation. Additionally, enabling community priors seeds the Beta posterior with aggregated reliability data from other installations, giving new models a realistic starting score before local statistics are gathered.

### Where is the routing configuration stored and how can it be modified?

All settings—including strategy weights, peak-hours configuration, and model overrides—are stored in the SQLite `settings` table. Operators can modify behavior at runtime using functions like `setRoutingStrategy` and `applyModelWeightOverride` without requiring a server restart or code redeployment.