How FreeLLMAPI Handles API Errors and Retries: A Deep Dive Into the Fallback Loop Architecture
FreeLLMAPI routes every request through a single shared fallback loop that centralizes error classification, automatic retries, intelligent cooldowns, and exhaustion handling across all OpenAI-compatible providers.
FreeLLMAPI is an open-source LLM gateway that aggregates free tiers from multiple providers (Groq, Together AI, Anthropic, etc.) behind a unified OpenAI-compatible API. Understanding how FreeLLMAPI handles API errors and retries is essential for operators deploying this proxy in production, as its resilience mechanisms directly impact uptime and latency. This article examines the core architecture implemented in tashfeenahmed/freellmapi, tracing the path from error classification to final client response.
The Centralized Fallback Loop Pattern
Unlike per-provider retry logic scattered across adapters, FreeLLMAPI implements a single shared fallback loop in server/src/lib/fallback-loop.ts that governs all retry behavior. This design ensures consistent handling whether you're hitting chat completions, embeddings, or Anthropic-native endpoints.
The loop operates through a hooks-based contract. Callers provide:
route(attempt)— selects the provider/key/model for each attemptdispatch(route, attempt, context)— executes the actual HTTP requestonFatal,onExhausted,onRoutingExhausted— lifecycle callbacks
The loop manages cross-cutting concerns: retry budgeting, cooldown calculation, circuit breaking, and observability headers. This keeps provider adapters thin and focused on serialization rather than resilience logic.
Error Classification: Four Fatal Categories
Before any retry decision, errors pass through server/src/lib/error-classify.ts, a pure-function classifier with no side effects. This module categorizes failures into four mutually exclusive buckets that determine routing behavior.
Retryable Errors
Retryable errors trigger automatic backoff and key rotation. The classifier flags:
- Network timeouts and TCP/TLS failures
- HTTP 5xx responses (with nuanced handling per provider)
- Explicit rate-limit signals (429 with
Retry-After) - Daily quota exhaustion markers
- Transient provider degradation
The isRetryableError function (lines 6-30) uses pattern matching on error.status, error.code, and provider-specific error message substrings to identify these cases.
// Error classification determines retry vs. fatal behavior
import { isRetryableError, isKeyAuthError, isModelNotFoundError } from './server/src/lib/error-classify';
function handleProviderError(err: unknown) {
if (isRetryableError(err)) {
// Falls back to next key/model with cooldown
return 'retry';
}
if (isKeyAuthError(err)) {
// Permanently blacklists this key
return 'auth-fatal';
}
if (isModelNotFoundError(err)) {
// Permanently blacklists this model for this provider
return 'model-fatal';
}
// Provider-level fatal: circuit-break this provider entirely
return 'provider-fatal';
}
Key-Authentication Fatal Errors
401 Unauthorized and invalid API key responses immediately blacklist the offending key. The isKeyAuthError function (lines 31-47) detects these to prevent burning retries on permanently invalid credentials. These errors skip cooldown entirely—keys enter skipKeys permanently until manual intervention.
Model-Level Fatal Errors
Model not found (404/410), model forbidden (403), and context-too-large errors indicate the request itself is invalid, not the provider infrastructure. The classifier uses isModelNotFoundError (lines 99-106) and isModelAccessForbiddenError (lines 111-119) to flag these. Affected models enter skipModels for the current request chain only, allowing other requests to retry them later.
Provider-Level Fatal Errors
Persistent 5xx storms, transport-layer failures, or explicit degraded-deployment markers trigger isProviderLevelError (lines 86-100). These circuit-break the entire provider for the request duration, falling back to alternate platforms entirely.
Retry Decision Logic and Budgeting
When isRetryableError returns true, the fallback loop executes recordRetryableFailure() (lines 93-104), which:
- Adds the key/model to ephemeral skip sets
- Computes cooldown duration via
cooldownDecisionForError()(lines 102-120)
Dynamic Cooldown Calculation
Cooldowns are context-sensitive rather than fixed:
| Error Type | Default Cooldown | Override Mechanism |
|---|---|---|
| Generic 5xx | 5 minutes | None |
Rate limit with Retry-After |
Header value | Respects provider guidance |
| Daily quota exhausted | 24 hours | Configurable per provider |
| Connection timeout | 30 seconds | FALLBACK_TIMEOUT_BACKOFF_MS |
The cooldownDecisionForError function analyzes error codes, response headers, and provider heuristics to select appropriate backoffs. Implementations respect the explicit over implicit principle: a provider's Retry-After header always wins over defaults.
Wall-Clock Retry Budget
FreeLLMAPI enforces a global time budget rather than simple attempt counting. FALLBACK_TIME_BUDGET_MS defaults to 45 seconds, configurable via getFallbackTimeBudgetMs() (lines 32-46). The loop aborts when Date.now() - startTime > budget, returning a controlled exhaustion response regardless of remaining candidates.
This prevents retry storms that degrade perceived latency. A request that fails through three quick timeouts returns faster than one attempting a fourth 30-second connection.
Circuit Breaker Integration
After configurable consecutive upstream failures, onExhausted (lines 81-85) can short-circuit with HTTP 503. This protects downstream providers from cascading overload and gives clients immediate feedback for retry decisions.
Cooldown Persistence and Key Routing
Retryable failures propagate to server/src/services/ratelimit.ts via setCooldown() (lines 24-30). This service maintains:
- In-memory LRU for hot-path cooldown checks
- Optional Redis backend for multi-instance consistency
Subsequent requests automatically exclude cooling keys through the skip-set mechanism. The routing layer queries getCooldownDecisionForLimit() before assignment, ensuring no request is routed to a key known to be in penalty.
// Cooldown integration in practice
import { setCooldown } from './server/src/services/ratelimit';
// Called automatically by fallback loop on retryable failure
await setCooldown({
key: 'sk-groq-xxx',
model: 'llama-3.3-70b',
platform: 'groq',
until: Date.now() + cooldownMs,
reason: error.code,
});
Exhaustion Response Construction
When all candidates are exhausted—whether through cooldowns, blacklists, or budget depletion—exhaustedRetryError() (lines 76-99) constructs a truthful, OpenAI-compatible error body:
- Status codes: 429 (rate limited), 502 (bad gateway), 503 (service unavailable), 404 (model not found), 413 (context too large)
retryAtMstimestamp for client-side backoff- Human-readable
messagewith attempt trail - Optional
fallbackDetailwith per-hop timing when debugging is enabled
This ensures clients receive actionable errors without exposing raw provider internals.
Observability: Tracing Every Hop
FreeLLMAPI surfaces retry behavior through standardized response headers set by setFallbackHeaders() (lines 53-66):
| Header | Purpose |
|---|---|
X-Fallback-Attempts |
Count of failed hops before success |
X-Fallback-Trail |
Compact log: groq/llama-3.3-70b key1=timeout; together/mistral key2=rate_limited |
X-Routed-Via |
Final successful provider/model |
With expose_fallback_detail_header: true, X-Fallback-Detail (lines 83-95) adds:
- Per-hop millisecond timings
- Redacted provider error messages (PII scrubbed)
- Retry-after values received from upstream
These headers enable client-side analytics and debugging without server log access.
Production Usage Example
// Production proxy wrapper with custom retry tuning
import { runFallbackLoop, FALLBACK_MAX_RETRIES } from './server/src/lib/fallback-loop';
async function proxyToFreeLLM(requestBody: unknown, preferredModel: string) {
const startTime = Date.now();
const result = await runFallbackLoop({
maxRetries: FALLBACK_MAX_RETRIES, // default 5
timeBudgetMs: 30000, // stricter than default 45s
state: {
skipKeys: new Set(),
skipModels: new Set(),
skipPlatforms: new Set(),
},
route: (attempt) => ({
url: getProviderUrl(attempt),
key: selectLeastLoadedKey(attempt),
model: attempt === 0 ? preferredModel : 'auto',
}),
dispatch: async (route, attempt, ctx) => {
const res = await fetch(route.url, {
method: 'POST',
headers: {
'Authorization': `Bearer ${route.key}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({ ...requestBody, model: route.model }),
signal: AbortSignal.timeout(8000), // per-attempt timeout
});
if (!res.ok) {
const err = new Error(`HTTP ${res.status}`);
(err as any).status = res.status;
(err as any).headers = Object.fromEntries(res.headers);
throw err; // Loop will classify and decide retry vs. fatal
}
return res.json();
},
logFailure: (route, err, attempt) => {
metrics.increment('freellmapi.provider_error', {
provider: route.platform,
error_class: err.constructor.name,
attempt: String(attempt),
});
},
onExhausted: (exhaustion) => {
alerts.trigger('freellmapi.routing_exhausted', {
duration: Date.now() - startTime,
trail: exhaustion.trail,
});
},
});
return result;
}
Summary
- Single fallback loop (
server/src/lib/fallback-loop.ts) centralizes all retry logic across FreeLLMAPI's provider surface - Four-tier error classification (
server/src/lib/error-classify.ts) distinguishes retryable, auth-fatal, model-fatal, and provider-fatal failures - Dynamic cooldowns respect provider
Retry-Afterheaders and implement sensible defaults for quota exhaustion (24h) vs. transient errors (minutes) - Wall-clock budgeting (default 45s) prevents retry storms and ensures predictable latency
- Observability headers (
X-Fallback-Attempts,X-Fallback-Trail, optionalX-Fallback-Detail) expose routing decisions to clients - Cooldown persistence (
server/src/services/ratelimit.ts) prevents repeated routing to penalized keys across concurrent requests
Frequently Asked Questions
How does FreeLLMAPI prevent infinite retry loops?
FreeLLMAPI enforces a wall-clock time budget (FALLBACK_TIME_BUDGET_MS, default 45 seconds) rather than counting attempts. The loop aborts when this budget elapses, returning a structured exhaustion response. Additionally, retryable failures trigger cooldown periods that remove keys from consideration, naturally limiting attempts as the candidate pool shrinks.
Can I customize retry behavior for specific providers?
Yes. The route() hook accepts the attempt number, allowing conditional logic per-provider. You can also adjust timeBudgetMs per-request or override FALLBACK_MAX_RETRIES. However, error classification (isRetryableError, etc.) is shared across providers to ensure consistent handling of standard HTTP semantics.
What happens when all providers are rate-limited simultaneously?
If all keys enter cooldown before success, exhaustedRetryError() returns HTTP 429 with retryAtMs set to the earliest cooldown expiration. The response includes the full X-Fallback-Trail showing which providers were tried and why each failed. Clients can use retryAtMs for efficient backoff rather than polling.
Are failed attempts visible in billing or usage logs?
While FreeLLMAPI does not meter upstream failures as "usage," every failed hop is recorded in the request trace via logFailure(). Operators can forward these to observability platforms. The X-Fallback-Attempts header in successful responses indicates overhead that may affect latency SLOs even when no client-visible error occurs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →