How FreeLLMAPI Handles Embeddings Failover: Family-Based Provider Fallback

FreeLLMAPI implements a family-based failover strategy that routes embedding requests through prioritized providers within the same model family, returning vectors from the first successful call or a 503 error if all providers fail.

FreeLLMAPI, an open-source unified API for large language models, ensures high availability for embedding generation through a sophisticated failover mechanism. Unlike chat completions that can fallback across different model families, embeddings failover operates strictly within a resolved family boundary. This design guarantees consistent vector dimensions and embedding semantics while maximizing uptime.

The Family-Based Failover Architecture

At the heart of FreeLLMAPI's embeddings failover is the embedding family abstraction. When a request arrives, the system resolves the supplied model name (e.g., gemini-embedding-001, openrouter/auto) to a specific family such as gemini, openrouter, or huggingface.

All providers belonging to the same family are stored in the embedding_models database table, each assigned a priority value. This prioritization determines the order in which providers are attempted during failover.

Provider Resolution and Prioritization

The resolution process begins in server/src/services/embeddings.ts via the resolveFamily(model) function. This helper returns the requested family or, if the model is omitted or set to auto, the default family defined in application settings.

Once resolved, the system queries the database for every enabled provider within that specific family, ordered by ascending priority. This ordered list becomes the failover sequence for the request.

Inside the runEmbeddings Failover Logic

The core failover implementation resides in the runEmbeddings function exported from server/src/services/embeddings.ts. This async function implements a sequential retry pattern with the following flow:

  1. Family Resolution – Determine the target embedding family using resolveFamily(model).
  2. Provider Selection – Fetch enabled providers from the embedding_models table sorted by priority.
  3. Sequential Execution – Iterate through providers, calling the low-level openAiStyleEmbed helper for each.
  4. Success or Continue – Return vectors immediately upon success; capture EmbeddingsError and continue to the next provider on failure.

The openAiStyleEmbed helper handles provider-specific API formats, managing different URLs and authentication schemes for Gemini, OpenRouter, NVIDIA, and HuggingFace endpoints.

// Direct service usage within the codebase
import { runEmbeddings } from '../services/embeddings.js';

const vectors = await runEmbeddings('gemini-embedding-001', ['hello world']);
// Returns: { vectors: number[][], inputTokens: number | null }

Error Handling and HTTP Status Codes

When a provider fails, FreeLLMAPI captures specific error conditions including HTTP 429 rate-limit responses, 502 upstream gateway failures, and malformed JSON responses. These are wrapped in an EmbeddingsError class and stored as lastError while the loop proceeds to the next provider.

If the function exhausts all providers in the family without success, it throws a final EmbeddingsError with:

  • Status Code: 503 (Service Unavailable)
  • Message Format: All providers for embedding family '<family>' failed (last: <lastError.message>)

This 503 status code signals to clients that the entire embedding family is unavailable, distinguishing between transient provider issues and systemic family outages.


# Public API request targeting default family

curl -X POST https://api.freellmapi.com/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model":"auto","input":["How many moons does Mars have?"],"dimensions":768}'

If the default provider (e.g., Gemini) returns a 429 throttling error, the server automatically attempts the next provider in the priority queue (e.g., OpenRouter) without client intervention.

API Endpoints for Embeddings

The failover logic is exposed through two public API endpoints defined in the route handlers:

POST /embeddings – Implemented in server/src/routes/proxy.ts, this endpoint forwards requests directly to runEmbeddings with the standard OpenAI-compatible request format.

POST /api/embeddings – Located in server/src/routes/embeddings.ts, this endpoint provides Ollama compatibility and similarly delegates to runEmbeddings.

Both endpoints return the unified response format containing the embedding vectors and token counts, or propagate the 503 error when failover is exhausted.

Summary

  • Family-scoped failover: Embeddings only failover between providers within the same family, never across families.
  • Priority-driven selection: Providers are ordered by database priority values in the embedding_models table.
  • Sequential retry pattern: The runEmbeddings function attempts providers one-by-one using openAiStyleEmbed.
  • 503 termination: Complete family failure results in a 503 status code with cumulative error details.
  • Dual endpoint support: Both /embeddings and /api/embeddings routes leverage the same core failover logic.

Frequently Asked Questions

Can FreeLLMAPI failover to a different embedding family if all providers in the current family fail?

No. According to the freellmapi source code, embeddings failover is strictly constrained to the resolved family. If all providers in the gemini family fail, the system returns a 503 error rather than attempting providers from the huggingface or openrouter families. To use a different family, you must explicitly request a different model name or change the default family setting.

What HTTP status code does FreeLLMAPI return when all embedding providers fail?

When every provider in an embedding family fails, FreeLLMAPI returns HTTP status 503 (Service Unavailable). The JSON error body includes the message: All providers for embedding family '<family>' failed (last: <specific_error>). This indicates systemic unavailability of the entire family rather than a single provider error.

How are provider priorities configured for embedding failover?

Provider priorities are stored in the embedding_models database table as numeric values. The runEmbeddings function queries this table for enabled providers within the target family, ordering results by priority. Lower priority values typically indicate higher precedence in the failover sequence, though the exact ordering logic depends on the SQL query implementation in server/src/services/embeddings.ts.

Does the embeddings failover mechanism differ from chat completion routing?

Yes. While chat completions in FreeLLMAPI can fallback across different model families (e.g., from GPT to Claude), embeddings never failover across families. This distinction ensures vector consistency—different families may produce embeddings with varying dimensions or semantic representations. The embedding family is fixed at request time via resolveFamily(model) and remains constant throughout the failover loop.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →