# How to Use the MLX Omni Server Text Embeddings Endpoint for Vector Similarity

> Learn to use the MLX Omni Server text embeddings endpoint for powerful vector similarity searches. Generate dense text representations for semantic comparisons.

- Repository: [madroid/mlx-omni-server](https://github.com/madroidmaq/mlx-omni-server)
- Tags: how-to-guide
- Published: 2026-03-06

---

**The MLX Omni Server exposes an OpenAI-compatible `/v1/embeddings` endpoint that returns dense vector representations for text, enabling semantic search and similarity comparisons using cosine distance or dot product metrics.**

The **mlx-omni-server** repository provides a high-performance local inference server for Apple Silicon, implementing OpenAI-compatible APIs for generating text embeddings. This article explains how to use the **`/v1/embeddings`** endpoint to generate dense vectors and compute **vector similarity** for semantic search applications.

## How the Text Embeddings Endpoint Works

The endpoint follows a standard request-response flow through the FastAPI application layer down to the MLX-based model inference engine.

### Request Routing and Validation

Incoming HTTP `POST` requests to `/v1/embeddings` (or the alias `/embeddings`) are handled by the `create_embeddings` function in **[`src/mlx_omni_server/embeddings/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/embeddings/router.py)**. This router validates the request body against the `EmbeddingRequest` Pydantic model defined in **[`src/mlx_omni_server/embeddings/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/embeddings/schema.py)**, which expects parameters including `model`, `input`, and optional `encoding_format`, `user`, and `dimensions`.

### Model Loading and Caching

The `EmbeddingsService` class in **[`src/mlx_omni_server/embeddings/embeddings_service.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/embeddings/embeddings_service.py)** manages the core logic. When a request arrives, `_get_model` loads the requested transformer model using `mlx_embeddings.load` and caches the `(model, processor)` tuple in the `_models` dictionary. This in-memory caching prevents reloading heavy model weights on subsequent requests, significantly improving latency for batch processing.

### Embedding Extraction Strategy

The service implements a two-tier extraction strategy in `_get_bert_embeddings`. For MiniLM-style models, it returns the **CLS token** representation (`last_hidden_state[:,0,:]`). For other BERT-like architectures, it applies **mean-pooling** across the hidden states. If BERT-specific extraction fails, the service falls back to `mlx_embeddings.generate`, extracting either the CLS token or raw output depending on the model architecture. The `_ensure_float_list` utility guarantees a flat `List[float]` output regardless of whether the raw data originates as a Python list, MLX array, or NumPy array.

### Token Counting and Usage Metadata

The `_count_tokens` method computes input token counts for the response's `usage` field. It attempts to use `tiktoken` when available, falling back to simple whitespace splitting for models without specific tokenizers.

## Generating Text Embeddings with Python

You can interact with the endpoint using the official OpenAI Python SDK or standard HTTP clients.

### Single and Batch Embeddings

Configure the client to point at your local server instance and specify the model identifier:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:10240/v1",
    api_key="not-needed"  # Local inference requires no authentication

)

# Single text embedding

response = client.embeddings.create(
    model="mlx-community/all-MiniLM-L6-v2-4bit",
    input="MLX Omni Server provides efficient local AI inference"
)
vector = response.data[0].embedding
print(f"Vector dimension: {len(vector)}")

# Batch processing for multiple texts

batch_response = client.embeddings.create(
    model="mlx-community/all-MiniLM-L6-v2-4bit",
    input=["First document", "Second document", "Third document"]
)
vectors = [item.embedding for item in batch_response.data]

```

### Raw HTTP Requests with cURL

For lightweight integrations or testing, send raw JSON payloads directly:

```bash
curl http://localhost:10240/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mlx-community/all-MiniLM-L6-v2-4bit",
    "input": ["Query text", "Document to compare"]
  }'

```

The server returns a JSON object containing a `data` array, where each element includes the `embedding` list and the original input index.

## Computing Vector Similarity from Embeddings

Once you retrieve the embedding vectors, calculate similarity using standard metrics. The most common approach for semantic search is **cosine similarity**, which measures the cosine of the angle between two vectors regardless of magnitude.

```python
import numpy as np
from scipy.spatial.distance import cosine

def cosine_similarity(a, b):
    return 1 - cosine(a, b)

# Compare first two vectors from batch

similarity = cosine_similarity(
    np.array(vectors[0]), 
    np.array(vectors[1])
)
print(f"Cosine similarity: {similarity:.4f}")

```

For production semantic search pipelines, store these vectors in a vector database like ChromaDB, Pinecone, or FAISS, then perform nearest-neighbor queries using the same similarity metrics.

## Key Implementation Files

Understanding the source architecture helps when debugging or extending functionality:

- **[`src/mlx_omni_server/embeddings/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/embeddings/router.py)** – FastAPI endpoint definitions and request routing
- **[`src/mlx_omni_server/embeddings/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/embeddings/schema.py)** – Pydantic models for `EmbeddingRequest` and `EmbeddingResponse` 
- **[`src/mlx_omni_server/embeddings/embeddings_service.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/embeddings/embeddings_service.py)** – Core business logic including `_get_model`, `_get_bert_embeddings`, and `_count_tokens`
- **[`src/mlx_omni_server/main.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/main.py)** – Application startup and router registration

## Summary

- The **MLX Omni Server** provides an OpenAI-compatible **`/v1/embeddings`** endpoint for generating text embeddings locally on Apple Silicon.
- The **`EmbeddingsService`** caches loaded models in memory and implements BERT-specific extraction (CLS token or mean-pooling) with a generic fallback.
- Request validation uses Pydantic schemas in **[`schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/schema.py)**, while routing logic resides in **[`router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/router.py)**.
- You can generate embeddings using the OpenAI Python SDK, raw HTTP requests, or any OpenAI-compatible client by pointing the base URL to `http://localhost:10240/v1`.
- Compute **vector similarity** using cosine distance, dot product, or Euclidean distance on the returned float arrays for semantic search and clustering tasks.

## Frequently Asked Questions

### What embedding models are supported by the MLX Omni Server text endpoint?

The endpoint supports any MLX-compatible transformer model loadable via `mlx_embeddings.load`, including BERT-based architectures like `mlx-community/all-MiniLM-L6-v2-4bit` and other sentence-transformers. The service automatically detects model architecture to apply the correct extraction strategy (CLS token for MiniLM-style models, mean-pooling for others).

### How does the endpoint handle token counting for the usage field?

The `_count_tokens` method in **[`embeddings_service.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/embeddings_service.py)** first attempts to use `tiktoken` for accurate OpenAI-compatible token counts. If `tiktoken` is unavailable or incompatible with the model, it falls back to simple whitespace tokenization to provide approximate counts for the `usage` metadata in the response.

### What is the difference between the BERT-specific and fallback embedding extraction?

The BERT-specific path in `_get_bert_embeddings` extracts contextualized representations using either the **[CLS] token** (for MiniLM-style models) or **mean-pooled hidden states** (for other BERT variants). If this extraction fails, the service falls back to `mlx_embeddings.generate`, which returns raw model outputs or CLS tokens depending on the specific model implementation, ensuring compatibility with non-BERT architectures.

### How do I calculate cosine similarity between two embedding vectors?

Convert the returned float lists to NumPy arrays and use the formula `cosine_sim = 1 - cosine(a, b)` from `scipy.spatial.distance`, or compute manually using `np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))`. Values range from -1 (opposite) to 1 (identical), with higher values indicating greater semantic similarity.