How to Use Embeddings with FreeLLMAPI: A Complete Guide to the OpenAI-Compatible API

FreeLLMAPI exposes a unified OpenAI-compatible embeddings endpoint that routes requests to free providers while guaranteeing vector-space consistency through a family-based architecture.

FreeLLMAPI is an open-source gateway that aggregates free LLM and embedding tiers from multiple providers into a single standardized interface. When you use embeddings with FreeLLMAPI, you interact with standard REST endpoints or SDKs while the service handles provider selection, credential encryption, and dimensional compatibility behind the scenes.

Understanding the Embedding Architecture

The core embedding logic resides in server/src/services/embeddings.ts. This service implements a family-based routing system where each model belongs to a specific family defined by its model ID and vector dimensions.

Family-Based Routing and Vector Consistency

FreeLLMAPI organizes models into families to ensure mathematical compatibility. The router strictly avoids failover across families because vectors from different dimensional spaces cannot be meaningfully compared. When you send a request, the service looks up encrypted credentials in the api_keys table, then forwards the request via proxyFetch (implemented in server/src/lib/proxy.ts) to the provider's /embeddings endpoint.

Supported Provider Platforms

The system supports major platforms defined in EMBEDDING_PLATFORMS, including Google, NVIDIA, OpenRouter, Cloudflare, HuggingFace, and Cohere. You can view the current catalog by calling GET /api/embeddings, which returns all available embedding families configured in your instance.

Core API Endpoints

The HTTP interface defined in server/src/routes/embeddings.ts provides four primary endpoints for embedding operations.

Retrieve Available Models

Send a GET request to /api/embeddings to list all configured embedding families. This returns metadata including model IDs, dimensions, and provider information.

Generate Embeddings

The POST /api/embeddings endpoint accepts OpenAI-standard JSON payloads with model and input fields. Set model to "auto" to use the configured default provider, or specify a concrete family name like "text-embedding-3-small" for deterministic routing.

Example structure:

curl -X POST http://localhost:3000/api/embeddings \
  -H "Content-Type: application/json" \
  -d '{
        "model":"auto",
        "input":["Free LLMAPI aggregates free tiers from many providers."]
      }'

Register Custom Providers

For private or local embedding endpoints (such as Ollama), use POST /api/embeddings/custom to register a new family. This endpoint invokes registerCustomEmbeddingModel and stores your configuration with parameters including keyId, modelId, family, dimensions, and maxInputTokens.

Implementation Examples

cURL Commands

Default provider selection:

curl -X POST http://localhost:3000/api/embeddings \
  -H "Content-Type: application/json" \
  -d '{
        "model":"auto",
        "input":["Free LLMAPI aggregates free tiers from many providers."]
      }'

Specific family selection:

curl -X POST http://localhost:3000/api/embeddings \
  -H "Content-Type: application/json" \
  -d '{
        "model":"text-embedding-3-small",
        "input":["Embedding vectors are useful for similarity search."]
      }'

Custom provider registration:

curl -X POST http://localhost:3000/api/embeddings/custom \
  -H "Content-Type: application/json" \
  -d '{
        "keyId": 42,
        "modelId": "my-ollama-embed",
        "displayName": "Ollama Embedding",
        "family": "ollama-embed",
        "dimensions": 768,
        "maxInputTokens": 4096,
        "quotaLabel": "default"
      }'

Python Client Integration

Since FreeLLMAPI is OpenAI-compatible, use the official openai Python SDK by setting the base_url to your local instance:

import openai

client = openai.OpenAI(base_url="http://localhost:3000/v1")
resp = client.embeddings.create(
    model="auto",                     # or a concrete family name

    input=["Free LLMAPI makes embeddings free!"]
)

vectors = resp.data[0].embedding   # list of floats

JavaScript and TypeScript Usage

The JavaScript implementation follows identical patterns using the OpenAI SDK:

import { OpenAI } from "openai";

const client = new OpenAI({ baseURL: "http://localhost:3000/v1" });
const result = await client.embeddings.create({
  model: "auto",
  input: ["Embedding vectors power RAG pipelines."]
});
console.log(result.data[0].embedding);

Key Files and Implementation Details

Understanding the source structure helps debug and extend embedding functionality:

Summary

  • FreeLLMAPI provides an OpenAI-compatible /api/embeddings endpoint that abstracts multiple free providers.
  • The family-based routing system ensures vector consistency by preventing cross-dimensional failover.
  • Use POST /api/embeddings with model: "auto" for default routing, or specify exact families for deterministic behavior.
  • Register custom embedding endpoints via POST /api/embeddings/custom to integrate local models like Ollama.
  • All responses follow the standard OpenAI embeddings format, ensuring compatibility with existing SDKs and tools.

Frequently Asked Questions

What is a "family" in FreeLLMAPI embeddings?

A family is a logical grouping defined by a specific model ID and vector dimension (e.g., 768 or 1536 dimensions). FreeLLMAPI routes requests within the same family to ensure vector compatibility, never mixing embeddings from different dimensional spaces because they occupy incompatible mathematical manifolds.

Can I use the standard OpenAI Python client with FreeLLMAPI?

Yes. Configure the client with base_url="http://localhost:3000/v1" (or your deployment URL) and use standard methods like client.embeddings.create(). The API returns identical response structures to OpenAI's official service, including the object: "embedding" type and float arrays.

How do I add a private embedding endpoint or local Ollama instance?

Use the POST /api/embeddings/custom endpoint to register your provider. Supply parameters including keyId, modelId, family, dimensions, and maxInputTokens as shown in the code examples. This registers the model in server/src/services/embeddings.ts for subsequent routing by the family selector.

Does FreeLLMAPI failover between different embedding providers?

The system only fails over between providers within the same family (matching dimensions). It never fails over across families because vectors from incompatible dimensional spaces cannot be meaningfully compared, searched, or averaged in vector databases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →