# How to Use MTPLX for RAG Setups with Embedding and Reranker Models

> Learn how to set up RAG with embedding and reranker models using MTPLX. Integrate OpenAI-compatible endpoints for seamless Retrieval-Augmented Generation pipelines.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-08

---

**MTPLX provides OpenAI-compatible `/v1/embeddings` and `/v1/rerank` endpoints in a single daemon, allowing you to serve embedding and reranker models alongside chat models for fully integrated Retrieval-Augmented Generation pipelines.**

The `youssofal/MTPLX` repository ships a unified inference server that eliminates the need for separate embedding and reranking services. By implementing the standard OpenAI API specification for retrieval tasks, MTPLX lets you run complete RAG workflows—document ingestion, semantic search, reranking, and answer generation—within a single process. This architecture reduces network overhead and simplifies deployment for Apple Silicon environments using MLX.

## Architecture Overview

MTPLX handles retrieval through two primary endpoints implemented in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py). The system lazily loads models on first request and keeps them resident in memory for subsequent calls.

| Component | Implementation | Source Location |
|-----------|---------------|-----------------|
| **Embedding endpoint** | Converts text to dense vectors via `POST /v1/embeddings` | [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) lines 30158-30230 |
| **Reranker endpoint** | Scores query-document pairs via `POST /v1/rerank` | [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) lines 30239-30258 |
| **Model management** | `Retrieval` class handles lazy loading and weight sharing | Internal `mtplx/retrieval/` module |
| **Descriptor API** | `retrieval.descriptors()` enumerates loaded models by capability | [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) lines 30141-30155 |

When you specify the same checkpoint for both embedding and reranking, MTPLX shares the underlying weights rather than loading duplicate copies into memory.

## Starting the Server with Retrieval Models

Launch the daemon with the `--embedding-model` and `--reranker-model` flags to expose the retrieval endpoints. The CLI implementation in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) (lines 2619-2621) parses these arguments and initializes the retrieval subsystem.

```bash
mtplx serve \
  --embedding-model mlx-community/Qwen3-Embedding-8B-4bit-DWQ \
  --reranker-model vserifsaglam/Qwen3-Reranker-4B-4bit-MLX

```

This command starts the server on the default port (`8000`). Models load on demand when the first request arrives, not at startup. You can verify available models by checking the descriptor iterator, which exposes each model's `id`, `capability` (`"embedding"` or `"rerank"`), and metadata.

### Security Configuration

By default, MTPLX blocks models containing executable Python code to prevent remote code execution. If your embedding or reranker model requires trusted remote code, enable it via:

```bash
mtplx serve \
  --embedding-model mlx-community/Qwen3-Embedding-8B-4bit-DWQ \
  --retrieval-trust-remote-code

```

Without this flag, requests to load models with custom code return **403 Forbidden** as implemented in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (lines 30190-30192). Alternatively, set `retrieval_trust_remote_code = true` in `~/.mtplx/config.toml`.

## Generating Embeddings with `/v1/embeddings`

The embedding endpoint accepts batches of text and returns dense vectors. Located at lines 30166-30214 in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), the handler validates parameters and performs the forward pass.

### Request Structure

- **model**: Identifier matching the loaded embedding model
- **input**: String or array of strings to embed
- **encoding_format**: Either `float` (default) or `base64`
- **dimensions**: Optional truncation or padding to specific vector size

Invalid parameter combinations return **400 Bad Request** with descriptive error messages.

### Example Requests

Using curl:

```bash
curl http://127.0.0.1:8000/v1/embeddings \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen3-Embedding-8B-4bit-DWQ",
    "input": ["Document one text", "Document two text"],
    "encoding_format": "float"
  }'

```

Using Python:

```python
import openai

openai.api_base = "http://127.0.0.1:8000/v1"
openai.api_key = "ignored"

response = openai.Embedding.create(
    model="Qwen3-Embedding-8B-4bit-DWQ",
    input="What is the cache?",
    dimensions=1024
)
query_vector = response["data"][0]["embedding"]

```

Store the resulting vectors in any vector database (FAISS, Chroma, Pinecone, etc.) for similarity search.

## Reranking Documents with `/v1/rerank`

After retrieving candidate documents via vector similarity, reranking improves result quality by scoring each document against the specific query. The reranker endpoint at lines 30239-30258 in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) implements this functionality.

### Request Parameters

- **model**: Reranker model identifier
- **query**: The search query string
- **documents**: Array of candidate document strings
- **top_n**: Optional limit for returned results

The endpoint returns scores aligned with the input document order, allowing you to sort candidates by relevance.

### Example Usage

```bash
curl http://127.0.0.1:8000/v1/rerank \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen3-Reranker-4B-4bit-MLX",
    "query": "what is the cache?",
    "documents": ["doc 1 text", "doc 2 text", "doc 3 text"]
  }'

```

## Building a Complete RAG Pipeline

Combine the embedding, reranking, and chat endpoints to build an end-to-end retrieval-augmented generation system.

### Step-by-Step Implementation

1. **Index documents**: Generate embeddings via `/v1/embeddings` and populate your vector store
2. **Query**: Convert user questions to embeddings using the same endpoint
3. **Retrieve**: Search the vector store for top-k candidates
4. **Rerank**: Pass candidates to `/v1/rerank` for relevance scoring
5. **Generate**: Submit the highest-ranked context to `/v1/chat/completions`

### Python Integration Example

```python
import openai

# Configure client for local MTPLX instance

openai.api_base = "http://127.0.0.1:8000/v1"
openai.api_key = "ignored"

# 1️⃣ Embed the query

emb_resp = openai.Embedding.create(
    model="Qwen3-Embedding-8B-4bit-DWQ",
    input="Explain how MTPLX performs token draft verification."
)
query_vec = emb_resp["data"][0]["embedding"]

# 2️⃣ Retrieve candidates from vector DB (pseudo-code)

candidates = vector_db.search(query_vec, k=5)

# 3️⃣ Rerank results

rerank_resp = openai.Rerank.create(
    model="Qwen3-Reranker-4B-4bit-MLX",
    query="Explain how MTPLX performs token draft verification.",
    documents=[c["text"] for c in candidates]
)

# Sort by score

sorted_indices = sorted(
    range(len(rerank_resp["data"])),
    key=lambda i: rerank_resp["data"][i]["score"],
    reverse=True
)
best_context = candidates[sorted_indices[0]]["text"]

# 4️⃣ Generate answer with context

chat_resp = openai.ChatCompletion.create(
    model="mtplx",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": f"{best_context}\n\nAnswer the question."}
    ]
)
print(chat_resp["choices"][0]["message"]["content"])

```

Because all components run in the same MTPLX process, you eliminate network hops between embedding, reranking, and generation services.

## Summary

- **MTPLX serves embedding and reranker models** through OpenAI-compatible endpoints (`/v1/embeddings` and `/v1/rerank`) within the same daemon used for chat inference.
- **Lazy loading and weight sharing** optimize memory usage when models are referenced multiple times.
- **Security defaults** block remote code execution unless explicitly enabled via `--retrieval-trust-remote-code` or config file settings.
- **Flexible parameters** support `float`/`base64` encoding formats and custom vector dimensions for embedding outputs.
- **Unified architecture** reduces infrastructure complexity by consolidating RAG components into a single Apple-optimized MLX server.

## Frequently Asked Questions

### Can I serve multiple embedding models simultaneously?

Yes. The server accepts multiple `--embedding-model` arguments and assigns each a unique identifier. The `retrieval.descriptors()` iterator exposes all loaded models with their capabilities, allowing clients to select specific models via the `"model"` field in request payloads.

### How do I handle embedding models that require custom Python code?

By default, MTPLX returns a **403 Forbidden** response if a model contains executable Python code. To enable these models, start the server with `--retrieval-trust-remote-code` or set `retrieval_trust_remote_code = true` in `~/.mtplx/config.toml` as documented in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (lines 30190-30192).

### Do I need separate inference servers for embeddings and chat completions?

No. MTPLX handles chat completions, embeddings, and reranking within a single process. This eliminates the network overhead and resource contention typical of microservice-based RAG architectures, allowing the chat model to access retrieval results with zero inter-service latency.

### What vector stores work with MTPLX embeddings?

Any vector database that accepts standard float arrays works with MTPLX embeddings, including FAISS, Chroma, Pinecone, Weaviate, and pgvector. The `/v1/embeddings` endpoint returns vectors compatible with all major similarity search implementations, with optional `dimensions` parameter support for truncating or padding to specific sizes required by your chosen store.