How to Use MTPLX for RAG Setups with Embedding and Reranker Models

MTPLX provides OpenAI-compatible /v1/embeddings and /v1/rerank endpoints in a single daemon, allowing you to serve embedding and reranker models alongside chat models for fully integrated Retrieval-Augmented Generation pipelines.

The youssofal/MTPLX repository ships a unified inference server that eliminates the need for separate embedding and reranking services. By implementing the standard OpenAI API specification for retrieval tasks, MTPLX lets you run complete RAG workflows—document ingestion, semantic search, reranking, and answer generation—within a single process. This architecture reduces network overhead and simplifies deployment for Apple Silicon environments using MLX.

Architecture Overview

MTPLX handles retrieval through two primary endpoints implemented in mtplx/server/openai.py. The system lazily loads models on first request and keeps them resident in memory for subsequent calls.

Component Implementation Source Location
Embedding endpoint Converts text to dense vectors via POST /v1/embeddings mtplx/server/openai.py lines 30158-30230
Reranker endpoint Scores query-document pairs via POST /v1/rerank mtplx/server/openai.py lines 30239-30258
Model management Retrieval class handles lazy loading and weight sharing Internal mtplx/retrieval/ module
Descriptor API retrieval.descriptors() enumerates loaded models by capability mtplx/server/openai.py lines 30141-30155

When you specify the same checkpoint for both embedding and reranking, MTPLX shares the underlying weights rather than loading duplicate copies into memory.

Starting the Server with Retrieval Models

Launch the daemon with the --embedding-model and --reranker-model flags to expose the retrieval endpoints. The CLI implementation in mtplx/cli.py (lines 2619-2621) parses these arguments and initializes the retrieval subsystem.

mtplx serve \
  --embedding-model mlx-community/Qwen3-Embedding-8B-4bit-DWQ \
  --reranker-model vserifsaglam/Qwen3-Reranker-4B-4bit-MLX

This command starts the server on the default port (8000). Models load on demand when the first request arrives, not at startup. You can verify available models by checking the descriptor iterator, which exposes each model's id, capability ("embedding" or "rerank"), and metadata.

Security Configuration

By default, MTPLX blocks models containing executable Python code to prevent remote code execution. If your embedding or reranker model requires trusted remote code, enable it via:

mtplx serve \
  --embedding-model mlx-community/Qwen3-Embedding-8B-4bit-DWQ \
  --retrieval-trust-remote-code

Without this flag, requests to load models with custom code return 403 Forbidden as implemented in mtplx/server/openai.py (lines 30190-30192). Alternatively, set retrieval_trust_remote_code = true in ~/.mtplx/config.toml.

Generating Embeddings with /v1/embeddings

The embedding endpoint accepts batches of text and returns dense vectors. Located at lines 30166-30214 in mtplx/server/openai.py, the handler validates parameters and performs the forward pass.

Request Structure

  • model: Identifier matching the loaded embedding model
  • input: String or array of strings to embed
  • encoding_format: Either float (default) or base64
  • dimensions: Optional truncation or padding to specific vector size

Invalid parameter combinations return 400 Bad Request with descriptive error messages.

Example Requests

Using curl:

curl http://127.0.0.1:8000/v1/embeddings \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen3-Embedding-8B-4bit-DWQ",
    "input": ["Document one text", "Document two text"],
    "encoding_format": "float"
  }'

Using Python:

import openai

openai.api_base = "http://127.0.0.1:8000/v1"
openai.api_key = "ignored"

response = openai.Embedding.create(
    model="Qwen3-Embedding-8B-4bit-DWQ",
    input="What is the cache?",
    dimensions=1024
)
query_vector = response["data"][0]["embedding"]

Store the resulting vectors in any vector database (FAISS, Chroma, Pinecone, etc.) for similarity search.

Reranking Documents with /v1/rerank

After retrieving candidate documents via vector similarity, reranking improves result quality by scoring each document against the specific query. The reranker endpoint at lines 30239-30258 in mtplx/server/openai.py implements this functionality.

Request Parameters

  • model: Reranker model identifier
  • query: The search query string
  • documents: Array of candidate document strings
  • top_n: Optional limit for returned results

The endpoint returns scores aligned with the input document order, allowing you to sort candidates by relevance.

Example Usage

curl http://127.0.0.1:8000/v1/rerank \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen3-Reranker-4B-4bit-MLX",
    "query": "what is the cache?",
    "documents": ["doc 1 text", "doc 2 text", "doc 3 text"]
  }'

Building a Complete RAG Pipeline

Combine the embedding, reranking, and chat endpoints to build an end-to-end retrieval-augmented generation system.

Step-by-Step Implementation

  1. Index documents: Generate embeddings via /v1/embeddings and populate your vector store
  2. Query: Convert user questions to embeddings using the same endpoint
  3. Retrieve: Search the vector store for top-k candidates
  4. Rerank: Pass candidates to /v1/rerank for relevance scoring
  5. Generate: Submit the highest-ranked context to /v1/chat/completions

Python Integration Example

import openai

# Configure client for local MTPLX instance

openai.api_base = "http://127.0.0.1:8000/v1"
openai.api_key = "ignored"

# 1️⃣ Embed the query

emb_resp = openai.Embedding.create(
    model="Qwen3-Embedding-8B-4bit-DWQ",
    input="Explain how MTPLX performs token draft verification."
)
query_vec = emb_resp["data"][0]["embedding"]

# 2️⃣ Retrieve candidates from vector DB (pseudo-code)

candidates = vector_db.search(query_vec, k=5)

# 3️⃣ Rerank results

rerank_resp = openai.Rerank.create(
    model="Qwen3-Reranker-4B-4bit-MLX",
    query="Explain how MTPLX performs token draft verification.",
    documents=[c["text"] for c in candidates]
)

# Sort by score

sorted_indices = sorted(
    range(len(rerank_resp["data"])),
    key=lambda i: rerank_resp["data"][i]["score"],
    reverse=True
)
best_context = candidates[sorted_indices[0]]["text"]

# 4️⃣ Generate answer with context

chat_resp = openai.ChatCompletion.create(
    model="mtplx",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": f"{best_context}\n\nAnswer the question."}
    ]
)
print(chat_resp["choices"][0]["message"]["content"])

Because all components run in the same MTPLX process, you eliminate network hops between embedding, reranking, and generation services.

Summary

  • MTPLX serves embedding and reranker models through OpenAI-compatible endpoints (/v1/embeddings and /v1/rerank) within the same daemon used for chat inference.
  • Lazy loading and weight sharing optimize memory usage when models are referenced multiple times.
  • Security defaults block remote code execution unless explicitly enabled via --retrieval-trust-remote-code or config file settings.
  • Flexible parameters support float/base64 encoding formats and custom vector dimensions for embedding outputs.
  • Unified architecture reduces infrastructure complexity by consolidating RAG components into a single Apple-optimized MLX server.

Frequently Asked Questions

Can I serve multiple embedding models simultaneously?

Yes. The server accepts multiple --embedding-model arguments and assigns each a unique identifier. The retrieval.descriptors() iterator exposes all loaded models with their capabilities, allowing clients to select specific models via the "model" field in request payloads.

How do I handle embedding models that require custom Python code?

By default, MTPLX returns a 403 Forbidden response if a model contains executable Python code. To enable these models, start the server with --retrieval-trust-remote-code or set retrieval_trust_remote_code = true in ~/.mtplx/config.toml as documented in mtplx/server/openai.py (lines 30190-30192).

Do I need separate inference servers for embeddings and chat completions?

No. MTPLX handles chat completions, embeddings, and reranking within a single process. This eliminates the network overhead and resource contention typical of microservice-based RAG architectures, allowing the chat model to access retrieval results with zero inter-service latency.

What vector stores work with MTPLX embeddings?

Any vector database that accepts standard float arrays works with MTPLX embeddings, including FAISS, Chroma, Pinecone, Weaviate, and pgvector. The /v1/embeddings endpoint returns vectors compatible with all major similarity search implementations, with optional dimensions parameter support for truncating or padding to specific sizes required by your chosen store.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →