How to Use the MLX Omni Server Text Embeddings Endpoint for Vector Similarity

The MLX Omni Server exposes an OpenAI-compatible /v1/embeddings endpoint that returns dense vector representations for text, enabling semantic search and similarity comparisons using cosine distance or dot product metrics.

The mlx-omni-server repository provides a high-performance local inference server for Apple Silicon, implementing OpenAI-compatible APIs for generating text embeddings. This article explains how to use the /v1/embeddings endpoint to generate dense vectors and compute vector similarity for semantic search applications.

How the Text Embeddings Endpoint Works

The endpoint follows a standard request-response flow through the FastAPI application layer down to the MLX-based model inference engine.

Request Routing and Validation

Incoming HTTP POST requests to /v1/embeddings (or the alias /embeddings) are handled by the create_embeddings function in src/mlx_omni_server/embeddings/router.py. This router validates the request body against the EmbeddingRequest Pydantic model defined in src/mlx_omni_server/embeddings/schema.py, which expects parameters including model, input, and optional encoding_format, user, and dimensions.

Model Loading and Caching

The EmbeddingsService class in src/mlx_omni_server/embeddings/embeddings_service.py manages the core logic. When a request arrives, _get_model loads the requested transformer model using mlx_embeddings.load and caches the (model, processor) tuple in the _models dictionary. This in-memory caching prevents reloading heavy model weights on subsequent requests, significantly improving latency for batch processing.

Embedding Extraction Strategy

The service implements a two-tier extraction strategy in _get_bert_embeddings. For MiniLM-style models, it returns the CLS token representation (last_hidden_state[:,0,:]). For other BERT-like architectures, it applies mean-pooling across the hidden states. If BERT-specific extraction fails, the service falls back to mlx_embeddings.generate, extracting either the CLS token or raw output depending on the model architecture. The _ensure_float_list utility guarantees a flat List[float] output regardless of whether the raw data originates as a Python list, MLX array, or NumPy array.

Token Counting and Usage Metadata

The _count_tokens method computes input token counts for the response's usage field. It attempts to use tiktoken when available, falling back to simple whitespace splitting for models without specific tokenizers.

Generating Text Embeddings with Python

You can interact with the endpoint using the official OpenAI Python SDK or standard HTTP clients.

Single and Batch Embeddings

Configure the client to point at your local server instance and specify the model identifier:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:10240/v1",
    api_key="not-needed"  # Local inference requires no authentication

)

# Single text embedding

response = client.embeddings.create(
    model="mlx-community/all-MiniLM-L6-v2-4bit",
    input="MLX Omni Server provides efficient local AI inference"
)
vector = response.data[0].embedding
print(f"Vector dimension: {len(vector)}")

# Batch processing for multiple texts

batch_response = client.embeddings.create(
    model="mlx-community/all-MiniLM-L6-v2-4bit",
    input=["First document", "Second document", "Third document"]
)
vectors = [item.embedding for item in batch_response.data]

Raw HTTP Requests with cURL

For lightweight integrations or testing, send raw JSON payloads directly:

curl http://localhost:10240/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mlx-community/all-MiniLM-L6-v2-4bit",
    "input": ["Query text", "Document to compare"]
  }'

The server returns a JSON object containing a data array, where each element includes the embedding list and the original input index.

Computing Vector Similarity from Embeddings

Once you retrieve the embedding vectors, calculate similarity using standard metrics. The most common approach for semantic search is cosine similarity, which measures the cosine of the angle between two vectors regardless of magnitude.

import numpy as np
from scipy.spatial.distance import cosine

def cosine_similarity(a, b):
    return 1 - cosine(a, b)

# Compare first two vectors from batch

similarity = cosine_similarity(
    np.array(vectors[0]), 
    np.array(vectors[1])
)
print(f"Cosine similarity: {similarity:.4f}")

For production semantic search pipelines, store these vectors in a vector database like ChromaDB, Pinecone, or FAISS, then perform nearest-neighbor queries using the same similarity metrics.

Key Implementation Files

Understanding the source architecture helps when debugging or extending functionality:

Summary

  • The MLX Omni Server provides an OpenAI-compatible /v1/embeddings endpoint for generating text embeddings locally on Apple Silicon.
  • The EmbeddingsService caches loaded models in memory and implements BERT-specific extraction (CLS token or mean-pooling) with a generic fallback.
  • Request validation uses Pydantic schemas in schema.py, while routing logic resides in router.py.
  • You can generate embeddings using the OpenAI Python SDK, raw HTTP requests, or any OpenAI-compatible client by pointing the base URL to http://localhost:10240/v1.
  • Compute vector similarity using cosine distance, dot product, or Euclidean distance on the returned float arrays for semantic search and clustering tasks.

Frequently Asked Questions

What embedding models are supported by the MLX Omni Server text endpoint?

The endpoint supports any MLX-compatible transformer model loadable via mlx_embeddings.load, including BERT-based architectures like mlx-community/all-MiniLM-L6-v2-4bit and other sentence-transformers. The service automatically detects model architecture to apply the correct extraction strategy (CLS token for MiniLM-style models, mean-pooling for others).

How does the endpoint handle token counting for the usage field?

The _count_tokens method in embeddings_service.py first attempts to use tiktoken for accurate OpenAI-compatible token counts. If tiktoken is unavailable or incompatible with the model, it falls back to simple whitespace tokenization to provide approximate counts for the usage metadata in the response.

What is the difference between the BERT-specific and fallback embedding extraction?

The BERT-specific path in _get_bert_embeddings extracts contextualized representations using either the [CLS] token (for MiniLM-style models) or mean-pooled hidden states (for other BERT variants). If this extraction fails, the service falls back to mlx_embeddings.generate, which returns raw model outputs or CLS tokens depending on the specific model implementation, ensuring compatibility with non-BERT architectures.

How do I calculate cosine similarity between two embedding vectors?

Convert the returned float lists to NumPy arrays and use the formula cosine_sim = 1 - cosine(a, b) from scipy.spatial.distance, or compute manually using np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)). Values range from -1 (opposite) to 1 (identical), with higher values indicating greater semantic similarity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →