How FreeLLMAPI Handles API Errors and Retries: A Deep Dive Into the Fallback Loop Architecture

FreeLLMAPI routes every request through a single shared fallback loop that centralizes error classification, automatic retries, intelligent cooldowns, and exhaustion handling across all OpenAI-compatible providers.

FreeLLMAPI is an open-source LLM gateway that aggregates free tiers from multiple providers (Groq, Together AI, Anthropic, etc.) behind a unified OpenAI-compatible API. Understanding how FreeLLMAPI handles API errors and retries is essential for operators deploying this proxy in production, as its resilience mechanisms directly impact uptime and latency. This article examines the core architecture implemented in tashfeenahmed/freellmapi, tracing the path from error classification to final client response.

The Centralized Fallback Loop Pattern

Unlike per-provider retry logic scattered across adapters, FreeLLMAPI implements a single shared fallback loop in server/src/lib/fallback-loop.ts that governs all retry behavior. This design ensures consistent handling whether you're hitting chat completions, embeddings, or Anthropic-native endpoints.

The loop operates through a hooks-based contract. Callers provide:

  • route(attempt) — selects the provider/key/model for each attempt
  • dispatch(route, attempt, context) — executes the actual HTTP request
  • onFatal, onExhausted, onRoutingExhausted — lifecycle callbacks

The loop manages cross-cutting concerns: retry budgeting, cooldown calculation, circuit breaking, and observability headers. This keeps provider adapters thin and focused on serialization rather than resilience logic.

Error Classification: Four Fatal Categories

Before any retry decision, errors pass through server/src/lib/error-classify.ts, a pure-function classifier with no side effects. This module categorizes failures into four mutually exclusive buckets that determine routing behavior.

Retryable Errors

Retryable errors trigger automatic backoff and key rotation. The classifier flags:

  • Network timeouts and TCP/TLS failures
  • HTTP 5xx responses (with nuanced handling per provider)
  • Explicit rate-limit signals (429 with Retry-After)
  • Daily quota exhaustion markers
  • Transient provider degradation

The isRetryableError function (lines 6-30) uses pattern matching on error.status, error.code, and provider-specific error message substrings to identify these cases.

// Error classification determines retry vs. fatal behavior
import { isRetryableError, isKeyAuthError, isModelNotFoundError } from './server/src/lib/error-classify';

function handleProviderError(err: unknown) {
  if (isRetryableError(err)) {
    // Falls back to next key/model with cooldown
    return 'retry';
  }
  if (isKeyAuthError(err)) {
    // Permanently blacklists this key
    return 'auth-fatal';
  }
  if (isModelNotFoundError(err)) {
    // Permanently blacklists this model for this provider
    return 'model-fatal';
  }
  // Provider-level fatal: circuit-break this provider entirely
  return 'provider-fatal';
}

Key-Authentication Fatal Errors

401 Unauthorized and invalid API key responses immediately blacklist the offending key. The isKeyAuthError function (lines 31-47) detects these to prevent burning retries on permanently invalid credentials. These errors skip cooldown entirely—keys enter skipKeys permanently until manual intervention.

Model-Level Fatal Errors

Model not found (404/410), model forbidden (403), and context-too-large errors indicate the request itself is invalid, not the provider infrastructure. The classifier uses isModelNotFoundError (lines 99-106) and isModelAccessForbiddenError (lines 111-119) to flag these. Affected models enter skipModels for the current request chain only, allowing other requests to retry them later.

Provider-Level Fatal Errors

Persistent 5xx storms, transport-layer failures, or explicit degraded-deployment markers trigger isProviderLevelError (lines 86-100). These circuit-break the entire provider for the request duration, falling back to alternate platforms entirely.

Retry Decision Logic and Budgeting

When isRetryableError returns true, the fallback loop executes recordRetryableFailure() (lines 93-104), which:

  1. Adds the key/model to ephemeral skip sets
  2. Computes cooldown duration via cooldownDecisionForError() (lines 102-120)

Dynamic Cooldown Calculation

Cooldowns are context-sensitive rather than fixed:

Error Type Default Cooldown Override Mechanism
Generic 5xx 5 minutes None
Rate limit with Retry-After Header value Respects provider guidance
Daily quota exhausted 24 hours Configurable per provider
Connection timeout 30 seconds FALLBACK_TIMEOUT_BACKOFF_MS

The cooldownDecisionForError function analyzes error codes, response headers, and provider heuristics to select appropriate backoffs. Implementations respect the explicit over implicit principle: a provider's Retry-After header always wins over defaults.

Wall-Clock Retry Budget

FreeLLMAPI enforces a global time budget rather than simple attempt counting. FALLBACK_TIME_BUDGET_MS defaults to 45 seconds, configurable via getFallbackTimeBudgetMs() (lines 32-46). The loop aborts when Date.now() - startTime > budget, returning a controlled exhaustion response regardless of remaining candidates.

This prevents retry storms that degrade perceived latency. A request that fails through three quick timeouts returns faster than one attempting a fourth 30-second connection.

Circuit Breaker Integration

After configurable consecutive upstream failures, onExhausted (lines 81-85) can short-circuit with HTTP 503. This protects downstream providers from cascading overload and gives clients immediate feedback for retry decisions.

Cooldown Persistence and Key Routing

Retryable failures propagate to server/src/services/ratelimit.ts via setCooldown() (lines 24-30). This service maintains:

  • In-memory LRU for hot-path cooldown checks
  • Optional Redis backend for multi-instance consistency

Subsequent requests automatically exclude cooling keys through the skip-set mechanism. The routing layer queries getCooldownDecisionForLimit() before assignment, ensuring no request is routed to a key known to be in penalty.

// Cooldown integration in practice
import { setCooldown } from './server/src/services/ratelimit';

// Called automatically by fallback loop on retryable failure
await setCooldown({
  key: 'sk-groq-xxx',
  model: 'llama-3.3-70b',
  platform: 'groq',
  until: Date.now() + cooldownMs,
  reason: error.code,
});

Exhaustion Response Construction

When all candidates are exhausted—whether through cooldowns, blacklists, or budget depletion—exhaustedRetryError() (lines 76-99) constructs a truthful, OpenAI-compatible error body:

  • Status codes: 429 (rate limited), 502 (bad gateway), 503 (service unavailable), 404 (model not found), 413 (context too large)
  • retryAtMs timestamp for client-side backoff
  • Human-readable message with attempt trail
  • Optional fallbackDetail with per-hop timing when debugging is enabled

This ensures clients receive actionable errors without exposing raw provider internals.

Observability: Tracing Every Hop

FreeLLMAPI surfaces retry behavior through standardized response headers set by setFallbackHeaders() (lines 53-66):

Header Purpose
X-Fallback-Attempts Count of failed hops before success
X-Fallback-Trail Compact log: groq/llama-3.3-70b key1=timeout; together/mistral key2=rate_limited
X-Routed-Via Final successful provider/model

With expose_fallback_detail_header: true, X-Fallback-Detail (lines 83-95) adds:

  • Per-hop millisecond timings
  • Redacted provider error messages (PII scrubbed)
  • Retry-after values received from upstream

These headers enable client-side analytics and debugging without server log access.

Production Usage Example

// Production proxy wrapper with custom retry tuning
import { runFallbackLoop, FALLBACK_MAX_RETRIES } from './server/src/lib/fallback-loop';

async function proxyToFreeLLM(requestBody: unknown, preferredModel: string) {
  const startTime = Date.now();
  
  const result = await runFallbackLoop({
    maxRetries: FALLBACK_MAX_RETRIES, // default 5
    timeBudgetMs: 30000,              // stricter than default 45s
    state: {
      skipKeys: new Set(),
      skipModels: new Set(),
      skipPlatforms: new Set(),
    },
    
    route: (attempt) => ({
      url: getProviderUrl(attempt),
      key: selectLeastLoadedKey(attempt),
      model: attempt === 0 ? preferredModel : 'auto',
    }),
    
    dispatch: async (route, attempt, ctx) => {
      const res = await fetch(route.url, {
        method: 'POST',
        headers: {
          'Authorization': `Bearer ${route.key}`,
          'Content-Type': 'application/json',
        },
        body: JSON.stringify({ ...requestBody, model: route.model }),
        signal: AbortSignal.timeout(8000), // per-attempt timeout
      });
      
      if (!res.ok) {
        const err = new Error(`HTTP ${res.status}`);
        (err as any).status = res.status;
        (err as any).headers = Object.fromEntries(res.headers);
        throw err; // Loop will classify and decide retry vs. fatal
      }
      
      return res.json();
    },
    
    logFailure: (route, err, attempt) => {
      metrics.increment('freellmapi.provider_error', {
        provider: route.platform,
        error_class: err.constructor.name,
        attempt: String(attempt),
      });
    },
    
    onExhausted: (exhaustion) => {
      alerts.trigger('freellmapi.routing_exhausted', {
        duration: Date.now() - startTime,
        trail: exhaustion.trail,
      });
    },
  });
  
  return result;
}

Summary

  • Single fallback loop (server/src/lib/fallback-loop.ts) centralizes all retry logic across FreeLLMAPI's provider surface
  • Four-tier error classification (server/src/lib/error-classify.ts) distinguishes retryable, auth-fatal, model-fatal, and provider-fatal failures
  • Dynamic cooldowns respect provider Retry-After headers and implement sensible defaults for quota exhaustion (24h) vs. transient errors (minutes)
  • Wall-clock budgeting (default 45s) prevents retry storms and ensures predictable latency
  • Observability headers (X-Fallback-Attempts, X-Fallback-Trail, optional X-Fallback-Detail) expose routing decisions to clients
  • Cooldown persistence (server/src/services/ratelimit.ts) prevents repeated routing to penalized keys across concurrent requests

Frequently Asked Questions

How does FreeLLMAPI prevent infinite retry loops?

FreeLLMAPI enforces a wall-clock time budget (FALLBACK_TIME_BUDGET_MS, default 45 seconds) rather than counting attempts. The loop aborts when this budget elapses, returning a structured exhaustion response. Additionally, retryable failures trigger cooldown periods that remove keys from consideration, naturally limiting attempts as the candidate pool shrinks.

Can I customize retry behavior for specific providers?

Yes. The route() hook accepts the attempt number, allowing conditional logic per-provider. You can also adjust timeBudgetMs per-request or override FALLBACK_MAX_RETRIES. However, error classification (isRetryableError, etc.) is shared across providers to ensure consistent handling of standard HTTP semantics.

What happens when all providers are rate-limited simultaneously?

If all keys enter cooldown before success, exhaustedRetryError() returns HTTP 429 with retryAtMs set to the earliest cooldown expiration. The response includes the full X-Fallback-Trail showing which providers were tried and why each failed. Clients can use retryAtMs for efficient backoff rather than polling.

Are failed attempts visible in billing or usage logs?

While FreeLLMAPI does not meter upstream failures as "usage," every failed hop is recorded in the request trace via logFailure(). Operators can forward these to observability platforms. The X-Fallback-Attempts header in successful responses indicates overhead that may affect latency SLOs even when no client-visible error occurs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →