How FreeLLMAPI's Smart Routing and Failover Mechanism Works: A Technical Deep Dive
FreeLLMAPI routes every request through a bandit-style router that dynamically scores every enabled model/key pair across reliability, speed, and intelligence axes, then selects the highest-scoring candidate using Thompson sampling while a surrounding fallback loop automatically retries with alternative providers when individual keys fail or hit quota limits.
FreeLLMAPI is an open-source LLM gateway developed by tashfeenahmed/freellmapi designed to distribute traffic across multiple providers without vendor lock-in. Understanding how FreeLLMAPI's smart routing and failover mechanism works reveals a sophisticated, self-optimizing system that continuously learns provider performance and gracefully handles outages through automatic demotion and retry logic.
The Eight-Step Routing Pipeline
The routing engine in server/src/services/router.ts executes a deterministic pipeline for every request. Each step refines the candidate pool until a single optimal provider is selected.
1. Building the Candidate Chain
The process begins in resolveRoutingChain and orderChain (lines ≈ 1881‑1890). The system queries the SQLite models table and joins entries with their associated API-key rows. Tier-fallback information stored in match_tier is attached to each candidate, creating an ordered list of potentially usable upstreams.
2. Gathering Per-Model Statistics
Recent request outcomes are cached in a decay-weighted statsCache via refreshStatsCache (lines ≈ 1850‑1870). This cache tracks successes, failures, token-throughput, first-byte latency, and monthly token usage for every model/key combination. The decay weighting ensures recent behavior influences routing decisions more heavily than historical data.
3. Computing Multi-Axis Quality Scores
For each candidate, scoreChainEntry (lines ≈ 5540‑5570) calculates three independent quality axes:
- Reliability – A Beta posterior distribution derived from successes versus failures, optionally seeded with community priors.
- Speed – Tokens-per-second throughput combined with first-byte latency into a composite performance score.
- Intelligence – A static rank derived from the catalog-provided tier and intelligence metadata.
4. Enforcing Budget Headroom Guardrails
Before final selection, the router applies financial and rate-limit protections. The functions headroomFactor and rateWindowHeadroomFactor (lines ≈ 5790‑5805) calculate multiplicative penalties based on remaining monthly token quotas and per-minute rate-limit windows. When a model approaches its budget ceiling, its score is proportionally reduced to prevent overages.
5. Applying Rate Limit and Failure Penalties
Transient failures trigger immediate score adjustments. The system maintains a global rateLimitPenalties map using recordRateLimitHit and recordModelFailure (lines ≈ 5660‑5665). Penalties decay over time and are applied as a multiplicative rateLimitFactor during scoring, effectively quarantining unstable providers until they recover.
6. Combining Scores with Strategy Weights
The active routing strategy—priority, balanced, smartest, or custom—determines a weight vector (RoutingWeights). The combineScore function merges the three quality axes using these weights. Operators can override weights for specific models via applyModelWeightOverride (lines ≈ 6060‑6070), enabling per-provider fine-tuning without code changes.
7. Selecting the Optimal Route
The candidate chain undergoes final sorting in orderChain: first by tier, then by computed score, then by manual priority. The top entry is returned as a RouteResult object containing the provider name, model ID, and decrypted API key. This occurs within routeRequest (lines ≈ 1881‑1900).
8. Executing the Failover Loop
If the selected key cannot be used—due to cooldown, quota exhaustion, or decryption errors—the router throws a RouteError. The fallback-loop.ts wrapper (lines ≈ 800‑825) catches this exception, records the failure via recordRateLimitHit or recordModelFailure, releases any acquired lease, and re-invokes routeRequest with the failing key and model added to exclusion lists (skipKeys, skipModels). This repeats until a usable upstream is found or the chain is exhausted.
Advanced Bandit-Style Optimizations
Beyond the core pipeline, FreeLLMAPI implements several bandit-algorithm enhancements to balance exploration against exploitation.
Thompson Sampling for Exploration
When the routing strategy is not set to priority, the router samples from the Beta posterior using sampleBeta rather than using the expected reliability value. This injects intrinsic exploration, allowing potentially better-performing providers to prove themselves without a separate exploration phase.
Anti-Starvation Protection
An optional EXPLORE_CHANCE flag ensures completely unmeasured models (cold starts) receive a guaranteed minimum selection probability. This prevents permanent starvation of new providers while the system gathers initial performance statistics.
Time-of-Day Performance Tuning
Operators can enable getPeakHoursConfig to apply time-of-day re-weighting. During configured windows, the system temporarily tilts scores toward cheaper or higher-throughput providers, optimizing for cost or latency based on predictable traffic patterns.
Community-Driven Reliability Priors
When enabled, the system merges aggregated reliability data from other FreeLLMAPI installations into the Beta posterior. This gives newly added models a sensible starting reliability score based on community experience, accelerating convergence to optimal routing.
Code Implementation Examples
Basic Request Routing
The simplest integration calls routeRequest directly from server route handlers:
import { routeRequest } from '../services/router.js';
// Estimate of output tokens the client will request (default 1000)
const estimatedTokens = 1200;
// Call the router; it returns the selected provider, model, and decrypted API key.
const result = routeRequest(estimatedTokens);
// Result fields you can forward to the upstream provider:
console.log(`Using ${result.provider.name} – model ${result.modelId}`);
console.log(`API key (redacted): ${result.apiKey.slice(0, 4)}…`);
Automatic Failover Implementation
For production resilience, wrap calls in the built-in fallback loop:
import { routeRequest } from '../services/router.js';
import { invokeProvider } from '../providers/base.js';
import { fallbackLoop } from '../lib/fallback-loop.js';
async function handleChat(request) {
return await fallbackLoop(async (skipKeys, skipModels) => {
// `routeRequest` will throw a RouteError if the whole chain is exhausted.
const route = routeRequest(/*estimatedTokens=*/1000, skipKeys, undefined, false, false, skipModels);
// Perform the actual upstream call.
const response = await invokeProvider(route.provider, {
model: route.modelId,
apiKey: route.apiKey,
// …other request fields…
});
// Record success so the model penalty decays.
recordSuccess(route.modelDbId);
return response;
});
}
Dynamic Strategy Adjustment
Change routing behavior at runtime without restarting the server:
import { setRoutingStrategy, getRoutingStrategy } from '../services/router.js';
// Switch to the “smartest” strategy (reliability‑driven)
setRoutingStrategy('smartest');
console.log('Current strategy:', getRoutingStrategy());
Key Configuration and Source Files
| File | Role |
|---|---|
server/src/services/router.ts |
Main routing engine implementing chain building, scoring, and strategy logic. |
server/src/lib/fallback-loop.ts |
Wraps routeRequest with retry logic that excludes failed keys and records penalties. |
server/src/services/scoring.ts |
Defines reliability, speed, and intelligence scoring functions and the combineScore combinator. |
server/src/services/ratelimit.ts |
Implements per-key and per-model rate-limit checks, lease acquisition, and cooldown handling. |
server/src/services/provider-quota.ts |
Calculates headroom factors based on monthly token budgets and usable key counts. |
All routing parameters, including strategy weights and peak-hours configuration, are persisted in the SQLite settings table, enabling runtime tuning without redeployment.
Summary
- FreeLLMAPI uses a bandit-style router that scores every model/key pair across reliability, speed, and intelligence axes.
- Thompson sampling provides intrinsic exploration, while configurable routing strategies allow operators to prioritize cost, speed, or reliability.
- Headroom guardrails and rate-limit penalties automatically protect budgets and quarantine failing providers.
- A fallback loop in
fallback-loop.tscatches routing errors and retries with excluded failed keys until a healthy upstream is found. - All configuration is stored in SQLite and adjustable at runtime via
setRoutingStrategyand related APIs.
Frequently Asked Questions
What routing strategies does FreeLLMAPI support?
FreeLLMAPI supports multiple strategies including priority (manual ordering), balanced (even distribution), and smartest (reliability-driven). Each strategy applies a distinct weight vector in combineScore to emphasize different axes—such as cost versus latency—when calculating final route scores.
How does FreeLLMAPI handle rate limit errors from upstream providers?
When a provider returns a 429 or connection failure, the system invokes recordRateLimitHit or recordModelFailure to inject a multiplicative penalty into the rateLimitPenalties map. This penalty decays over time and reduces the provider’s selection probability. Simultaneously, the fallback-loop.ts wrapper catches the error and retries with the failing key excluded.
Can FreeLLMAPI route to unmeasured or newly added models?
Yes. The optional EXPLORE_CHANCE flag guarantees that unmeasured models receive a minimum selection probability, preventing starvation. Additionally, enabling community priors seeds the Beta posterior with aggregated reliability data from other installations, giving new models a realistic starting score before local statistics are gathered.
Where is the routing configuration stored and how can it be modified?
All settings—including strategy weights, peak-hours configuration, and model overrides—are stored in the SQLite settings table. Operators can modify behavior at runtime using functions like setRoutingStrategy and applyModelWeightOverride without requiring a server restart or code redeployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →