How to Perform Automatic Failover with FreeLLMAPI: Complete Implementation Guide
FreeLLMAPI implements automatic failover through an intelligent request router that maintains a prioritized fallback chain of models, automatically retrying requests on alternative providers when encountering rate limits (429) or server errors (5xx) while applying temporary cooldown penalties to unhealthy keys.
FreeLLMAPI is an open-source aggregation layer that unifies access to multiple free-tier LLM providers through a single OpenAI-compatible API. Understanding how to leverage its automatic failover capabilities ensures your applications remain resilient when individual providers hit rate limits or experience outages.
Architecture of the Failover System
The failover logic resides in server/src/services/router.ts, where the Router orchestrates request distribution across a dynamically ordered fallback chain. When you send a request to /v1/chat/completions, the router evaluates multiple providers sequentially until successful or exhausted.
The Fallback Chain Strategy
The router constructs the fallback chain using selectable routing strategies including priority, balanced, and smartest. These strategies leverage the scoring engine in server/src/services/scoring.ts to rank models by reliability, speed, and intelligence metrics, ensuring optimal provider selection.
Health and Capability Validation
Before selecting a provider, the router performs three critical validations:
- Key health: Periodic probes in
server/src/services/health.tsclassify keys ashealthy,rate_limited,invalid, orerror - Quota enforcement: SQLite-backed in-memory counters tracked in
server/src/services/ratelimit.tsvia functions likecanMakeRequest()andisOnCooldown() - Capability matching: Model-level features (vision, tool-calling, context window) validated through
modelsWithOverriddenField()andscopeAllows()
How the Failover Mechanism Works
When a request encounters failure, FreeLLMAPI's transparent failover mechanism activates without client intervention.
Trigger Conditions and Penalty Application
If a provider returns a 429 (rate limit) or 5xx error, the router immediately:
- Applies a cooldown to the offending key using
acquireLease()andreleaseLease()inratelimit.ts - Records the failure via
recordRateLimitHit()orrecordModelFailure()(lines 48-53 inrouter.ts) - Increments temporary penalties that demote the model in the fallback chain
These penalties automatically decay over time via DECAY_INTERVAL_MS, allowing providers to recover once limits reset.
Retry Logic and Exhaustion Handling
The router attempts up to 20 retries (or until a wall-clock budget expires) against subsequent models in the chain. The summarizeExhaustion() function logs the complete attempt trail to the request table, enabling dashboard visualization of the exact failover path taken for each request.
Implementing Automatic Failover in Practice
No additional configuration is required to enable failover—it's transparent for all requests. However, you can observe and interact with the mechanism using the following patterns.
Standard API Request
The simplest implementation uses bearer token authentication. If the primary provider fails, the router automatically promotes the next available model:
curl https://api.freellmapi.com/v1/chat/completions \
-H "Authorization: Bearer freellmapi-<your-token>" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{ "role": "user", "content": "Explain automatic failover." }],
"max_tokens": 500
}'
Node.js Client Implementation
Use the official client package (@freellmapi/client) to handle failover transparently in JavaScript applications:
import { FreeLLMAPI } from '@freellmapi/client';
const api = new FreeLLMAPI({ token: 'freellmapi-<your-token>' });
async function generateContent() {
const resp = await api.chatCompletions.create({
model: 'gpt-4o-mini',
messages: [{ role: 'user', content: 'What is automatic failover?' }],
max_tokens: 500,
});
console.log(resp.choices[0].message.content);
}
generateContent().catch(console.error);
Debugging Failover Activity
To observe the failover chain in action, use verbose mode to inspect the x-free-llm-retry-attempts response header:
curl https://api.freellmapi.com/v1/chat/completions \
-H "Authorization: Bearer freellmapi-<token>" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{ "role": "user", "content": "Trigger failover test." }],
"max_tokens": 200
}' -v
The verbose output reveals the request-ID and the number of providers attempted before success. The dashboard UI (client/ directory) also displays the full attempt trail using data from the summarizeExhaustion function.
Core Components Reference
| Component | Source File | Responsibility |
|---|---|---|
| Router | server/src/services/router.ts |
Picks highest-priority healthy model, handles 429/5xx retries |
| Rate Limiter | server/src/services/ratelimit.ts |
SQLite-backed RPM/RPD/TPM tracking, cooldown management |
| Health Monitor | server/src/services/health.ts |
Periodic key status probes (healthy, rate_limited, invalid) |
| Scoring Engine | server/src/services/scoring.ts |
Bandit algorithm calculating reliability/intelligence scores |
| Provider Adapters | server/src/providers/*.ts |
Individual upstream LLM implementations (Google, Groq, Cerebras) |
Summary
- Automatic failover in FreeLLMAPI occurs in
server/src/services/router.tsthrough a prioritized fallback chain of models. - The router checks health status, rate limits, and capabilities before each attempt using
health.tsandratelimit.ts. - 429 and 5xx errors trigger cooldowns via
acquireLease()and penalize providers throughrecordModelFailure(). - Up to 20 retry attempts are made against alternative providers; the client receives only the final successful response.
- Request history and failover trails are accessible via the
summarizeExhaustionfunction and dashboard UI.
Frequently Asked Questions
How many retry attempts does FreeLLMAPI perform during failover?
FreeLLMAPI attempts up to approximately 20 retries or continues until a wall-clock retry budget is exhausted. This behavior is hardcoded in the router logic within server/src/services/router.ts, ensuring comprehensive coverage of the fallback chain without indefinite loops.
Can I configure which providers participate in the automatic failover chain?
Yes, the fallback chain is constructed based on selectable routing strategies (priority, balanced, smartest) defined in the router configuration. The scoring.ts engine ranks available models, and you can influence selection by configuring model priorities, though the automatic demotion of unhealthy providers happens dynamically through the penalty system.
How does FreeLLMAPI handle rate limit cooldowns?
When a provider returns a 429 error, the router invokes recordRateLimitHit() and applies a cooldown using acquireLease() and releaseLease() functions in server/src/services/ratelimit.ts. This temporarily removes the provider from consideration; penalties decay automatically over time according to DECAY_INTERVAL_MS, allowing the provider to return to the rotation once limits reset.
Where can I view the failover attempt history for a specific request?
The summarizeExhaustion() function in router.ts logs all retry attempts to the SQLite request table. You can view the complete attempt trail—including which providers were tried and why they failed—through the dashboard UI in the client/ directory or by inspecting the x-free-llm-retry-attempts header in API responses when using verbose mode.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →