# How FreeLLMAPI Handles API Errors and Retries: A Deep Dive Into the Fallback Loop Architecture

> Discover how FreeLLMAPI handles API errors and retries with its fallback loop architecture. Learn about centralized error classification, automatic retries, and cooldowns for OpenAI-compatible providers.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-08-28

---

**FreeLLMAPI routes every request through a single shared fallback loop that centralizes error classification, automatic retries, intelligent cooldowns, and exhaustion handling across all OpenAI-compatible providers.**

FreeLLMAPI is an open-source LLM gateway that aggregates free tiers from multiple providers (Groq, Together AI, Anthropic, etc.) behind a unified OpenAI-compatible API. Understanding how FreeLLMAPI handles API errors and retries is essential for operators deploying this proxy in production, as its resilience mechanisms directly impact uptime and latency. This article examines the core architecture implemented in `tashfeenahmed/freellmapi`, tracing the path from error classification to final client response.

## The Centralized Fallback Loop Pattern

Unlike per-provider retry logic scattered across adapters, FreeLLMAPI implements a **single shared fallback loop** in [`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts) that governs all retry behavior. This design ensures consistent handling whether you're hitting chat completions, embeddings, or Anthropic-native endpoints.

The loop operates through a **hooks-based contract**. Callers provide:
- `route(attempt)` — selects the provider/key/model for each attempt
- `dispatch(route, attempt, context)` — executes the actual HTTP request
- `onFatal`, `onExhausted`, `onRoutingExhausted` — lifecycle callbacks

The loop manages cross-cutting concerns: **retry budgeting**, **cooldown calculation**, **circuit breaking**, and **observability headers**. This keeps provider adapters thin and focused on serialization rather than resilience logic.

## Error Classification: Four Fatal Categories

Before any retry decision, errors pass through [`server/src/lib/error-classify.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/error-classify.ts), a pure-function classifier with no side effects. This module categorizes failures into four mutually exclusive buckets that determine routing behavior.

### Retryable Errors

**Retryable** errors trigger automatic backoff and key rotation. The classifier flags:

- Network timeouts and TCP/TLS failures
- HTTP 5xx responses (with nuanced handling per provider)
- Explicit rate-limit signals (429 with `Retry-After`)
- Daily quota exhaustion markers
- Transient provider degradation

The `isRetryableError` function (lines 6-30) uses pattern matching on `error.status`, `error.code`, and provider-specific error message substrings to identify these cases.

```typescript
// Error classification determines retry vs. fatal behavior
import { isRetryableError, isKeyAuthError, isModelNotFoundError } from './server/src/lib/error-classify';

function handleProviderError(err: unknown) {
  if (isRetryableError(err)) {
    // Falls back to next key/model with cooldown
    return 'retry';
  }
  if (isKeyAuthError(err)) {
    // Permanently blacklists this key
    return 'auth-fatal';
  }
  if (isModelNotFoundError(err)) {
    // Permanently blacklists this model for this provider
    return 'model-fatal';
  }
  // Provider-level fatal: circuit-break this provider entirely
  return 'provider-fatal';
}

```

### Key-Authentication Fatal Errors

**401 Unauthorized** and invalid API key responses immediately blacklist the offending key. The `isKeyAuthError` function (lines 31-47) detects these to prevent burning retries on permanently invalid credentials. These errors skip cooldown entirely—keys enter `skipKeys` permanently until manual intervention.

### Model-Level Fatal Errors

**Model not found** (404/410), **model forbidden** (403), and **context-too-large** errors indicate the request itself is invalid, not the provider infrastructure. The classifier uses `isModelNotFoundError` (lines 99-106) and `isModelAccessForbiddenError` (lines 111-119) to flag these. Affected models enter `skipModels` for the current request chain only, allowing other requests to retry them later.

### Provider-Level Fatal Errors

Persistent 5xx storms, transport-layer failures, or explicit degraded-deployment markers trigger `isProviderLevelError` (lines 86-100). These **circuit-break** the entire provider for the request duration, falling back to alternate platforms entirely.

## Retry Decision Logic and Budgeting

When `isRetryableError` returns true, the fallback loop executes `recordRetryableFailure()` (lines 93-104), which:

1. Adds the key/model to ephemeral skip sets
2. Computes cooldown duration via `cooldownDecisionForError()` (lines 102-120)

### Dynamic Cooldown Calculation

Cooldowns are context-sensitive rather than fixed:

| Error Type | Default Cooldown | Override Mechanism |
|------------|-----------------|-------------------|
| Generic 5xx | 5 minutes | None |
| Rate limit with `Retry-After` | Header value | Respects provider guidance |
| Daily quota exhausted | 24 hours | Configurable per provider |
| Connection timeout | 30 seconds | `FALLBACK_TIMEOUT_BACKOFF_MS` |

The `cooldownDecisionForError` function analyzes error codes, response headers, and provider heuristics to select appropriate backoffs. Implementations respect the **explicit over implicit** principle: a provider's `Retry-After` header always wins over defaults.

### Wall-Clock Retry Budget

FreeLLMAPI enforces a **global time budget** rather than simple attempt counting. `FALLBACK_TIME_BUDGET_MS` defaults to **45 seconds**, configurable via `getFallbackTimeBudgetMs()` (lines 32-46). The loop aborts when `Date.now() - startTime > budget`, returning a controlled exhaustion response regardless of remaining candidates.

This prevents **retry storms** that degrade perceived latency. A request that fails through three quick timeouts returns faster than one attempting a fourth 30-second connection.

### Circuit Breaker Integration

After configurable consecutive upstream failures, `onExhausted` (lines 81-85) can short-circuit with HTTP 503. This protects downstream providers from cascading overload and gives clients immediate feedback for retry decisions.

## Cooldown Persistence and Key Routing

Retryable failures propagate to [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) via `setCooldown()` (lines 24-30). This service maintains:

- **In-memory LRU** for hot-path cooldown checks
- **Optional Redis backend** for multi-instance consistency

Subsequent requests automatically exclude cooling keys through the skip-set mechanism. The routing layer queries `getCooldownDecisionForLimit()` before assignment, ensuring no request is routed to a key known to be in penalty.

```typescript
// Cooldown integration in practice
import { setCooldown } from './server/src/services/ratelimit';

// Called automatically by fallback loop on retryable failure
await setCooldown({
  key: 'sk-groq-xxx',
  model: 'llama-3.3-70b',
  platform: 'groq',
  until: Date.now() + cooldownMs,
  reason: error.code,
});

```

## Exhaustion Response Construction

When all candidates are exhausted—whether through cooldowns, blacklists, or budget depletion—`exhaustedRetryError()` (lines 76-99) constructs a **truthful, OpenAI-compatible error body**:

- Status codes: 429 (rate limited), 502 (bad gateway), 503 (service unavailable), 404 (model not found), 413 (context too large)
- `retryAtMs` timestamp for client-side backoff
- Human-readable `message` with attempt trail
- Optional `fallbackDetail` with per-hop timing when debugging is enabled

This ensures clients receive actionable errors without exposing raw provider internals.

## Observability: Tracing Every Hop

FreeLLMAPI surfaces retry behavior through **standardized response headers** set by `setFallbackHeaders()` (lines 53-66):

| Header | Purpose |
|--------|---------|
| `X-Fallback-Attempts` | Count of failed hops before success |
| `X-Fallback-Trail` | Compact log: `groq/llama-3.3-70b key1=timeout; together/mistral key2=rate_limited` |
| `X-Routed-Via` | Final successful provider/model |

With `expose_fallback_detail_header: true`, `X-Fallback-Detail` (lines 83-95) adds:

- Per-hop millisecond timings
- Redacted provider error messages (PII scrubbed)
- Retry-after values received from upstream

These headers enable client-side analytics and debugging without server log access.

## Production Usage Example

```typescript
// Production proxy wrapper with custom retry tuning
import { runFallbackLoop, FALLBACK_MAX_RETRIES } from './server/src/lib/fallback-loop';

async function proxyToFreeLLM(requestBody: unknown, preferredModel: string) {
  const startTime = Date.now();
  
  const result = await runFallbackLoop({
    maxRetries: FALLBACK_MAX_RETRIES, // default 5
    timeBudgetMs: 30000,              // stricter than default 45s
    state: {
      skipKeys: new Set(),
      skipModels: new Set(),
      skipPlatforms: new Set(),
    },
    
    route: (attempt) => ({
      url: getProviderUrl(attempt),
      key: selectLeastLoadedKey(attempt),
      model: attempt === 0 ? preferredModel : 'auto',
    }),
    
    dispatch: async (route, attempt, ctx) => {
      const res = await fetch(route.url, {
        method: 'POST',
        headers: {
          'Authorization': `Bearer ${route.key}`,
          'Content-Type': 'application/json',
        },
        body: JSON.stringify({ ...requestBody, model: route.model }),
        signal: AbortSignal.timeout(8000), // per-attempt timeout
      });
      
      if (!res.ok) {
        const err = new Error(`HTTP ${res.status}`);
        (err as any).status = res.status;
        (err as any).headers = Object.fromEntries(res.headers);
        throw err; // Loop will classify and decide retry vs. fatal
      }
      
      return res.json();
    },
    
    logFailure: (route, err, attempt) => {
      metrics.increment('freellmapi.provider_error', {
        provider: route.platform,
        error_class: err.constructor.name,
        attempt: String(attempt),
      });
    },
    
    onExhausted: (exhaustion) => {
      alerts.trigger('freellmapi.routing_exhausted', {
        duration: Date.now() - startTime,
        trail: exhaustion.trail,
      });
    },
  });
  
  return result;
}

```

## Summary

- **Single fallback loop** ([`server/src/lib/fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/fallback-loop.ts)) centralizes all retry logic across FreeLLMAPI's provider surface
- **Four-tier error classification** ([`server/src/lib/error-classify.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/error-classify.ts)) distinguishes retryable, auth-fatal, model-fatal, and provider-fatal failures
- **Dynamic cooldowns** respect provider `Retry-After` headers and implement sensible defaults for quota exhaustion (24h) vs. transient errors (minutes)
- **Wall-clock budgeting** (default 45s) prevents retry storms and ensures predictable latency
- **Observability headers** (`X-Fallback-Attempts`, `X-Fallback-Trail`, optional `X-Fallback-Detail`) expose routing decisions to clients
- **Cooldown persistence** ([`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts)) prevents repeated routing to penalized keys across concurrent requests

## Frequently Asked Questions

### How does FreeLLMAPI prevent infinite retry loops?

FreeLLMAPI enforces a **wall-clock time budget** (`FALLBACK_TIME_BUDGET_MS`, default 45 seconds) rather than counting attempts. The loop aborts when this budget elapses, returning a structured exhaustion response. Additionally, retryable failures trigger cooldown periods that remove keys from consideration, naturally limiting attempts as the candidate pool shrinks.

### Can I customize retry behavior for specific providers?

Yes. The `route()` hook accepts the attempt number, allowing conditional logic per-provider. You can also adjust `timeBudgetMs` per-request or override `FALLBACK_MAX_RETRIES`. However, error classification (`isRetryableError`, etc.) is shared across providers to ensure consistent handling of standard HTTP semantics.

### What happens when all providers are rate-limited simultaneously?

If all keys enter cooldown before success, `exhaustedRetryError()` returns HTTP 429 with `retryAtMs` set to the earliest cooldown expiration. The response includes the full `X-Fallback-Trail` showing which providers were tried and why each failed. Clients can use `retryAtMs` for efficient backoff rather than polling.

### Are failed attempts visible in billing or usage logs?

While FreeLLMAPI does not meter upstream failures as "usage," every failed hop is recorded in the **request trace** via `logFailure()`. Operators can forward these to observability platforms. The `X-Fallback-Attempts` header in successful responses indicates overhead that may affect latency SLOs even when no client-visible error occurs.