How to Use Embeddings with FreeLLMAPI: A Complete Guide to the OpenAI-Compatible API
FreeLLMAPI exposes a unified OpenAI-compatible embeddings endpoint that routes requests to free providers while guaranteeing vector-space consistency through a family-based architecture.
FreeLLMAPI is an open-source gateway that aggregates free LLM and embedding tiers from multiple providers into a single standardized interface. When you use embeddings with FreeLLMAPI, you interact with standard REST endpoints or SDKs while the service handles provider selection, credential encryption, and dimensional compatibility behind the scenes.
Understanding the Embedding Architecture
The core embedding logic resides in server/src/services/embeddings.ts. This service implements a family-based routing system where each model belongs to a specific family defined by its model ID and vector dimensions.
Family-Based Routing and Vector Consistency
FreeLLMAPI organizes models into families to ensure mathematical compatibility. The router strictly avoids failover across families because vectors from different dimensional spaces cannot be meaningfully compared. When you send a request, the service looks up encrypted credentials in the api_keys table, then forwards the request via proxyFetch (implemented in server/src/lib/proxy.ts) to the provider's /embeddings endpoint.
Supported Provider Platforms
The system supports major platforms defined in EMBEDDING_PLATFORMS, including Google, NVIDIA, OpenRouter, Cloudflare, HuggingFace, and Cohere. You can view the current catalog by calling GET /api/embeddings, which returns all available embedding families configured in your instance.
Core API Endpoints
The HTTP interface defined in server/src/routes/embeddings.ts provides four primary endpoints for embedding operations.
Retrieve Available Models
Send a GET request to /api/embeddings to list all configured embedding families. This returns metadata including model IDs, dimensions, and provider information.
Generate Embeddings
The POST /api/embeddings endpoint accepts OpenAI-standard JSON payloads with model and input fields. Set model to "auto" to use the configured default provider, or specify a concrete family name like "text-embedding-3-small" for deterministic routing.
Example structure:
curl -X POST http://localhost:3000/api/embeddings \
-H "Content-Type: application/json" \
-d '{
"model":"auto",
"input":["Free LLMAPI aggregates free tiers from many providers."]
}'
Register Custom Providers
For private or local embedding endpoints (such as Ollama), use POST /api/embeddings/custom to register a new family. This endpoint invokes registerCustomEmbeddingModel and stores your configuration with parameters including keyId, modelId, family, dimensions, and maxInputTokens.
Implementation Examples
cURL Commands
Default provider selection:
curl -X POST http://localhost:3000/api/embeddings \
-H "Content-Type: application/json" \
-d '{
"model":"auto",
"input":["Free LLMAPI aggregates free tiers from many providers."]
}'
Specific family selection:
curl -X POST http://localhost:3000/api/embeddings \
-H "Content-Type: application/json" \
-d '{
"model":"text-embedding-3-small",
"input":["Embedding vectors are useful for similarity search."]
}'
Custom provider registration:
curl -X POST http://localhost:3000/api/embeddings/custom \
-H "Content-Type: application/json" \
-d '{
"keyId": 42,
"modelId": "my-ollama-embed",
"displayName": "Ollama Embedding",
"family": "ollama-embed",
"dimensions": 768,
"maxInputTokens": 4096,
"quotaLabel": "default"
}'
Python Client Integration
Since FreeLLMAPI is OpenAI-compatible, use the official openai Python SDK by setting the base_url to your local instance:
import openai
client = openai.OpenAI(base_url="http://localhost:3000/v1")
resp = client.embeddings.create(
model="auto", # or a concrete family name
input=["Free LLMAPI makes embeddings free!"]
)
vectors = resp.data[0].embedding # list of floats
JavaScript and TypeScript Usage
The JavaScript implementation follows identical patterns using the OpenAI SDK:
import { OpenAI } from "openai";
const client = new OpenAI({ baseURL: "http://localhost:3000/v1" });
const result = await client.embeddings.create({
model: "auto",
input: ["Embedding vectors power RAG pipelines."]
});
console.log(result.data[0].embedding);
Key Files and Implementation Details
Understanding the source structure helps debug and extend embedding functionality:
server/src/services/embeddings.ts– Contains the core routing logic, provider adapters, andregisterCustomEmbeddingModelfunction.server/src/routes/embeddings.ts– Defines the HTTP API surface including route handlers for GET and POST endpoints.server/src/lib/proxy.ts– ImplementsproxyFetch, which handles request forwarding to upstream providers with timeout management and error translation.
Summary
- FreeLLMAPI provides an OpenAI-compatible
/api/embeddingsendpoint that abstracts multiple free providers. - The family-based routing system ensures vector consistency by preventing cross-dimensional failover.
- Use
POST /api/embeddingswithmodel: "auto"for default routing, or specify exact families for deterministic behavior. - Register custom embedding endpoints via
POST /api/embeddings/customto integrate local models like Ollama. - All responses follow the standard OpenAI embeddings format, ensuring compatibility with existing SDKs and tools.
Frequently Asked Questions
What is a "family" in FreeLLMAPI embeddings?
A family is a logical grouping defined by a specific model ID and vector dimension (e.g., 768 or 1536 dimensions). FreeLLMAPI routes requests within the same family to ensure vector compatibility, never mixing embeddings from different dimensional spaces because they occupy incompatible mathematical manifolds.
Can I use the standard OpenAI Python client with FreeLLMAPI?
Yes. Configure the client with base_url="http://localhost:3000/v1" (or your deployment URL) and use standard methods like client.embeddings.create(). The API returns identical response structures to OpenAI's official service, including the object: "embedding" type and float arrays.
How do I add a private embedding endpoint or local Ollama instance?
Use the POST /api/embeddings/custom endpoint to register your provider. Supply parameters including keyId, modelId, family, dimensions, and maxInputTokens as shown in the code examples. This registers the model in server/src/services/embeddings.ts for subsequent routing by the family selector.
Does FreeLLMAPI failover between different embedding providers?
The system only fails over between providers within the same family (matching dimensions). It never fails over across families because vectors from incompatible dimensional spaces cannot be meaningfully compared, searched, or averaged in vector databases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →