# How FreeLLMAPI Handles Embeddings Failover: Family-Based Provider Fallback

> Discover how FreeLLMAPI ensures embedding request success with its family-based failover strategy. Learn how it routes requests and handles provider fallback for reliable vector generation.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-08-29

---

**FreeLLMAPI implements a family-based failover strategy that routes embedding requests through prioritized providers within the same model family, returning vectors from the first successful call or a 503 error if all providers fail.**

FreeLLMAPI, an open-source unified API for large language models, ensures high availability for embedding generation through a sophisticated failover mechanism. Unlike chat completions that can fallback across different model families, embeddings failover operates strictly within a resolved family boundary. This design guarantees consistent vector dimensions and embedding semantics while maximizing uptime.

## The Family-Based Failover Architecture

At the heart of FreeLLMAPI's embeddings failover is the **embedding family** abstraction. When a request arrives, the system resolves the supplied model name (e.g., `gemini-embedding-001`, `openrouter/auto`) to a specific family such as `gemini`, `openrouter`, or `huggingface`.

All providers belonging to the same family are stored in the `embedding_models` database table, each assigned a **priority value**. This prioritization determines the order in which providers are attempted during failover.

### Provider Resolution and Prioritization

The resolution process begins in [`server/src/services/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/embeddings.ts) via the `resolveFamily(model)` function. This helper returns the requested family or, if the model is omitted or set to `auto`, the default family defined in application settings.

Once resolved, the system queries the database for every *enabled* provider within that specific family, ordered by ascending priority. This ordered list becomes the failover sequence for the request.

## Inside the `runEmbeddings` Failover Logic

The core failover implementation resides in the **`runEmbeddings`** function exported from [`server/src/services/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/embeddings.ts). This async function implements a sequential retry pattern with the following flow:

1. **Family Resolution** – Determine the target embedding family using `resolveFamily(model)`.
2. **Provider Selection** – Fetch enabled providers from the `embedding_models` table sorted by priority.
3. **Sequential Execution** – Iterate through providers, calling the low-level `openAiStyleEmbed` helper for each.
4. **Success or Continue** – Return vectors immediately upon success; capture `EmbeddingsError` and continue to the next provider on failure.

The `openAiStyleEmbed` helper handles provider-specific API formats, managing different URLs and authentication schemes for Gemini, OpenRouter, NVIDIA, and HuggingFace endpoints.

```typescript
// Direct service usage within the codebase
import { runEmbeddings } from '../services/embeddings.js';

const vectors = await runEmbeddings('gemini-embedding-001', ['hello world']);
// Returns: { vectors: number[][], inputTokens: number | null }

```

## Error Handling and HTTP Status Codes

When a provider fails, FreeLLMAPI captures specific error conditions including HTTP 429 rate-limit responses, 502 upstream gateway failures, and malformed JSON responses. These are wrapped in an **`EmbeddingsError`** class and stored as `lastError` while the loop proceeds to the next provider.

If the function exhausts all providers in the family without success, it throws a final `EmbeddingsError` with:

- **Status Code:** `503` (Service Unavailable)
- **Message Format:** `All providers for embedding family '<family>' failed (last: <lastError.message>)`

This 503 status code signals to clients that the entire embedding family is unavailable, distinguishing between transient provider issues and systemic family outages.

```bash

# Public API request targeting default family

curl -X POST https://api.freellmapi.com/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model":"auto","input":["How many moons does Mars have?"],"dimensions":768}'

```

If the default provider (e.g., Gemini) returns a 429 throttling error, the server automatically attempts the next provider in the priority queue (e.g., OpenRouter) without client intervention.

## API Endpoints for Embeddings

The failover logic is exposed through two public API endpoints defined in the route handlers:

**`POST /embeddings`** – Implemented in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts), this endpoint forwards requests directly to `runEmbeddings` with the standard OpenAI-compatible request format.

**`POST /api/embeddings`** – Located in [`server/src/routes/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/embeddings.ts), this endpoint provides Ollama compatibility and similarly delegates to `runEmbeddings`.

Both endpoints return the unified response format containing the embedding vectors and token counts, or propagate the 503 error when failover is exhausted.

## Summary

- **Family-scoped failover:** Embeddings only failover between providers within the same family, never across families.
- **Priority-driven selection:** Providers are ordered by database priority values in the `embedding_models` table.
- **Sequential retry pattern:** The `runEmbeddings` function attempts providers one-by-one using `openAiStyleEmbed`.
- **503 termination:** Complete family failure results in a 503 status code with cumulative error details.
- **Dual endpoint support:** Both `/embeddings` and `/api/embeddings` routes leverage the same core failover logic.

## Frequently Asked Questions

### Can FreeLLMAPI failover to a different embedding family if all providers in the current family fail?

No. According to the `freellmapi` source code, embeddings failover is strictly constrained to the resolved family. If all providers in the `gemini` family fail, the system returns a 503 error rather than attempting providers from the `huggingface` or `openrouter` families. To use a different family, you must explicitly request a different model name or change the default family setting.

### What HTTP status code does FreeLLMAPI return when all embedding providers fail?

When every provider in an embedding family fails, FreeLLMAPI returns HTTP status **503** (Service Unavailable). The JSON error body includes the message: `All providers for embedding family '<family>' failed (last: <specific_error>)`. This indicates systemic unavailability of the entire family rather than a single provider error.

### How are provider priorities configured for embedding failover?

Provider priorities are stored in the `embedding_models` database table as numeric values. The `runEmbeddings` function queries this table for enabled providers within the target family, ordering results by priority. Lower priority values typically indicate higher precedence in the failover sequence, though the exact ordering logic depends on the SQL query implementation in [`server/src/services/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/embeddings.ts).

### Does the embeddings failover mechanism differ from chat completion routing?

Yes. While chat completions in FreeLLMAPI can fallback across different model families (e.g., from GPT to Claude), **embeddings never failover across families**. This distinction ensures vector consistency—different families may produce embeddings with varying dimensions or semantic representations. The embedding family is fixed at request time via `resolveFamily(model)` and remains constant throughout the failover loop.