# How to Use Embeddings with FreeLLMAPI: A Complete Guide to the OpenAI-Compatible API

> Learn how to use embeddings with FreeLLMAPI's OpenAI-compatible API. Access free embedding providers and maintain vector-space consistency effortlessly.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-08-29

---

**FreeLLMAPI exposes a unified OpenAI-compatible embeddings endpoint that routes requests to free providers while guaranteeing vector-space consistency through a family-based architecture.**

FreeLLMAPI is an open-source gateway that aggregates free LLM and embedding tiers from multiple providers into a single standardized interface. When you use embeddings with FreeLLMAPI, you interact with standard REST endpoints or SDKs while the service handles provider selection, credential encryption, and dimensional compatibility behind the scenes.

## Understanding the Embedding Architecture

The core embedding logic resides in [`server/src/services/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/embeddings.ts). This service implements a **family-based routing system** where each model belongs to a specific family defined by its model ID and vector dimensions.

### Family-Based Routing and Vector Consistency

FreeLLMAPI organizes models into **families** to ensure mathematical compatibility. The router strictly avoids failover across families because vectors from different dimensional spaces cannot be meaningfully compared. When you send a request, the service looks up encrypted credentials in the `api_keys` table, then forwards the request via **proxyFetch** (implemented in [`server/src/lib/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/proxy.ts)) to the provider's `/embeddings` endpoint.

### Supported Provider Platforms

The system supports major platforms defined in **EMBEDDING_PLATFORMS**, including Google, NVIDIA, OpenRouter, Cloudflare, HuggingFace, and Cohere. You can view the current catalog by calling `GET /api/embeddings`, which returns all available embedding families configured in your instance.

## Core API Endpoints

The HTTP interface defined in [`server/src/routes/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/embeddings.ts) provides four primary endpoints for embedding operations.

### Retrieve Available Models

Send a GET request to `/api/embeddings` to list all configured embedding families. This returns metadata including model IDs, dimensions, and provider information.

### Generate Embeddings

The `POST /api/embeddings` endpoint accepts OpenAI-standard JSON payloads with `model` and `input` fields. Set `model` to `"auto"` to use the configured default provider, or specify a concrete family name like `"text-embedding-3-small"` for deterministic routing.

Example structure:

```bash
curl -X POST http://localhost:3000/api/embeddings \
  -H "Content-Type: application/json" \
  -d '{
        "model":"auto",
        "input":["Free LLMAPI aggregates free tiers from many providers."]
      }'

```

### Register Custom Providers

For private or local embedding endpoints (such as Ollama), use `POST /api/embeddings/custom` to register a new family. This endpoint invokes `registerCustomEmbeddingModel` and stores your configuration with parameters including `keyId`, `modelId`, `family`, `dimensions`, and `maxInputTokens`.

## Implementation Examples

### cURL Commands

Default provider selection:

```bash
curl -X POST http://localhost:3000/api/embeddings \
  -H "Content-Type: application/json" \
  -d '{
        "model":"auto",
        "input":["Free LLMAPI aggregates free tiers from many providers."]
      }'

```

Specific family selection:

```bash
curl -X POST http://localhost:3000/api/embeddings \
  -H "Content-Type: application/json" \
  -d '{
        "model":"text-embedding-3-small",
        "input":["Embedding vectors are useful for similarity search."]
      }'

```

Custom provider registration:

```bash
curl -X POST http://localhost:3000/api/embeddings/custom \
  -H "Content-Type: application/json" \
  -d '{
        "keyId": 42,
        "modelId": "my-ollama-embed",
        "displayName": "Ollama Embedding",
        "family": "ollama-embed",
        "dimensions": 768,
        "maxInputTokens": 4096,
        "quotaLabel": "default"
      }'

```

### Python Client Integration

Since FreeLLMAPI is OpenAI-compatible, use the official `openai` Python SDK by setting the `base_url` to your local instance:

```python
import openai

client = openai.OpenAI(base_url="http://localhost:3000/v1")
resp = client.embeddings.create(
    model="auto",                     # or a concrete family name

    input=["Free LLMAPI makes embeddings free!"]
)

vectors = resp.data[0].embedding   # list of floats

```

### JavaScript and TypeScript Usage

The JavaScript implementation follows identical patterns using the OpenAI SDK:

```javascript
import { OpenAI } from "openai";

const client = new OpenAI({ baseURL: "http://localhost:3000/v1" });
const result = await client.embeddings.create({
  model: "auto",
  input: ["Embedding vectors power RAG pipelines."]
});
console.log(result.data[0].embedding);

```

## Key Files and Implementation Details

Understanding the source structure helps debug and extend embedding functionality:

- **[`server/src/services/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/embeddings.ts)** – Contains the core routing logic, provider adapters, and `registerCustomEmbeddingModel` function.
- **[`server/src/routes/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/embeddings.ts)** – Defines the HTTP API surface including route handlers for GET and POST endpoints.
- **[`server/src/lib/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/proxy.ts)** – Implements `proxyFetch`, which handles request forwarding to upstream providers with timeout management and error translation.

## Summary

- **FreeLLMAPI** provides an OpenAI-compatible `/api/embeddings` endpoint that abstracts multiple free providers.
- The **family-based routing system** ensures vector consistency by preventing cross-dimensional failover.
- Use `POST /api/embeddings` with `model: "auto"` for default routing, or specify exact families for deterministic behavior.
- Register custom embedding endpoints via `POST /api/embeddings/custom` to integrate local models like Ollama.
- All responses follow the standard OpenAI embeddings format, ensuring compatibility with existing SDKs and tools.

## Frequently Asked Questions

### What is a "family" in FreeLLMAPI embeddings?

A **family** is a logical grouping defined by a specific model ID and vector dimension (e.g., 768 or 1536 dimensions). FreeLLMAPI routes requests within the same family to ensure vector compatibility, never mixing embeddings from different dimensional spaces because they occupy incompatible mathematical manifolds.

### Can I use the standard OpenAI Python client with FreeLLMAPI?

Yes. Configure the client with `base_url="http://localhost:3000/v1"` (or your deployment URL) and use standard methods like `client.embeddings.create()`. The API returns identical response structures to OpenAI's official service, including the `object: "embedding"` type and float arrays.

### How do I add a private embedding endpoint or local Ollama instance?

Use the `POST /api/embeddings/custom` endpoint to register your provider. Supply parameters including `keyId`, `modelId`, `family`, `dimensions`, and `maxInputTokens` as shown in the code examples. This registers the model in [`server/src/services/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/embeddings.ts) for subsequent routing by the family selector.

### Does FreeLLMAPI failover between different embedding providers?

The system only fails over between providers within the same **family** (matching dimensions). It never fails over across families because vectors from incompatible dimensional spaces cannot be meaningfully compared, searched, or averaged in vector databases.