# How to Use Embedding Models and Pooling Parameters in vLLM: Complete Technical Guide

> Learn to use embedding models and pooling parameters in vLLM. This guide details how vLLM's pooling runner generates text embeddings efficiently by routing hidden states through the model's pooler.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**vLLM serves embedding models through a dedicated pooling runner that converts HTTP requests into `PoolingParams` instances, routing hidden states through the model's pooler instead of the language-model decoder to generate text embeddings.**

vLLM supports high-performance inference for text embedding models through a specialized pooling architecture distinct from its generation pipeline. This guide explains how to configure embedding models and pooling parameters in vLLM based on the `vllm-project/vllm` source code, covering the complete request lifecycle from HTTP API to GPU execution.

## Understanding the Pooling Runner Architecture

When you start vLLM with the `--runner pooling` flag, the engine initializes a **pooling runner** rather than a standard language-model decoder. In [`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py) (line 408), the worker checks `self.is_pooling_model` (determined by `model_config.runner_type == "pooling"`) to route tensors through the encoder and pooler while bypassing the generation head.

This architecture is essential for models like `BAAI/bge-m3` that lack a language-model head. The engine detects this configuration via `model_config.pooler_config` and sets `max_output_tokens` to zero, treating the request as **embedding-only** inference.

## The PoolingParams Configuration Object

The `PoolingParams` class defined in [`vllm/pooling_params.py`](https://github.com/vllm-project/vllm/blob/main/vllm/pooling_params.py) controls how the pooler processes hidden states. When a request arrives at the `/pooling` endpoint, the `to_pooling_params()` method in [`vllm/entrypoints/pooling/embed/protocol.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/pooling/embed/protocol.py) (lines 59-64) constructs this object with the following fields:

- **`task`**: Fixed to `"embed"` for text embedding models. Set automatically by the protocol layer.
- **`dimensions`**: Target output dimensionality. Only honored when the model supports **Matryoshka** (nested-vector) representations.
- **`use_activation`**: Controls whether the pooler applies its activation function (typically `Tanh`). Defaults to `True` when set to `None`.
- **`step_tag_id` and `returned_token_ids`**: Reserved for step-wise pooling tasks such as token-wise reranking, not used for standard embeddings.

The `_merge_default_parameters` function in [`vllm/pooling_params.py`](https://github.com/vllm-project/vllm/blob/main/vllm/pooling_params.py) (line 85) merges user-provided values with model-level defaults from `ModelConfig.pooler_config`. Invalid parameter combinations are rejected by `_verify_valid_parameters` (line 68), ensuring that unsupported options (like `dimensions` on non-Matryoshka models) raise errors before execution.

## Request Flow: From HTTP to Embedding Tensor

### HTTP Request Parsing

Incoming requests hit [`vllm/entrypoints/pooling/embed/protocol.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/pooling/embed/protocol.py), where `EmbeddingCompletionRequest` or `EmbeddingChatRequest` objects parse the JSON payload. The `to_pooling_params()` method extracts `dimensions` and `use_activation` fields to build the `PoolingParams` instance.

### Parameter Validation

In [`vllm/v1/engine/input_processor.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/input_processor.py), the engine calls `PoolingParams.verify()` to validate and normalize parameters against the model's capabilities. This verification step ensures that the pooling task is compatible with the provided modifiers before dispatching to the GPU worker.

### Engine Orchestration

The `LLMEngine` in [`vllm/v1/engine/llm_engine.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/llm_engine.py) forwards validated `PoolingParams` to the worker. For programmatic usage, the `LLM.encode()` method internally constructs and validates these parameters before execution.

### GPU Execution

The [`gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/gpu_model_runner.py) worker checks `is_pooling_model` and processes inputs through the encoder network. Hidden states pass through the model's pooler rather than the decoder, producing a `PoolingOutput` tensor with shape `[batch, dim]`.

### Response Encoding

The serving layer in [`vllm/entrypoints/pooling/embed/serving.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/pooling/embed/serving.py) converts `PoolingRequestOutput` objects into OpenAI-compatible responses. Utility functions in [`vllm/entrypoints/pooling/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/pooling/utils.py) handle encoding:

- **`encode_pooling_output_float`**: Returns JSON arrays of floating-point values.
- **`encode_pooling_output_base64`**: Returns base64-encoded binary for efficient large-scale retrieval.

## Practical Implementation Examples

### Starting a Pooling Server

Launch the server with the pooling runner flag to enable embedding endpoints:

```bash
vllm serve BAAI/bge-m3 --runner pooling --port 8000

```

The `--runner pooling` argument is required to activate the embedding pipeline instead of the generation path.

### Sending HTTP Requests with Pooling Parameters

Send POST requests to the `/pooling` endpoint with optional dimension and activation controls:

```python
import json
import requests

url = "http://localhost:8000/pooling"
payload = {
    "model": "BAAI/bge-m3",
    "input": [
        "Hello, my name is",
        "The capital of France is Paris."
    ],
    "dimensions": 256,
    "use_activation": True,
    "encoding_format": "float"
}

resp = requests.post(
    url, 
    json=payload, 
    headers={"User-Agent": "vllm-client"}
)

data = resp.json()["data"]
embeddings = [entry["embedding"] for entry in data]
print(json.dumps(embeddings, indent=2))

```

The `dimensions` and `use_activation` fields map directly to `PoolingParams` attributes, allowing dynamic control over the output vector size and activation application.

### Using the Python Client Library

For integrated Python applications, use the `LLM` class with explicit `PoolingParams`:

```python
from vllm import LLM, PoolingParams

# Initialize engine with pooling runner

engine = LLM(
    model="BAAI/bge-m3", 
    runner_type="pooling"
)

# Configure pooling parameters

pool_params = PoolingParams(
    task="embed",
    dimensions=256,
    use_activation=False
)

# Generate embeddings

outputs = engine.encode(
    ["I love AI.", "vLLM makes inference fast."],
    pooling_params=pool_params
)

# Access tensor output

print(outputs[0].embedding)  # Shape: (256,)

```

The `engine.encode()` method handles `PoolingParams` validation internally and returns `PoolingRequestOutput` objects containing the embedding tensors.

### Binary Encoding for Large-Scale Retrieval

For production retrieval systems requiring minimal bandwidth, request base64-encoded binary responses:

```python
import requests
import base64
import struct

url = "http://localhost:8000/pooling"
payload = {
    "model": "BAAI/bge-m3",
    "input": ["query 1", "query 2"],
    "encoding_format": "base64",
    "embed_dtype": "float32",
    "endianness": "little"
}

resp = requests.post(url, json=payload)
metadata = resp.headers.get("metadata")

# Decode binary payload (example: 256-dim float32)

binary = resp.content
first_vec = struct.unpack("<256f", binary[:1024])
print(first_vec[:5])

```

Binary encoding is processed by `encode_pooling_output_base64` in [`vllm/entrypoints/pooling/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/pooling/utils.py), significantly reducing response size compared to JSON arrays.

## Summary

- **vLLM uses a pooling runner** (activated via `--runner pooling`) to serve embedding models, routing tensors through the pooler rather than the language-model decoder.
- **`PoolingParams`** in [`vllm/pooling_params.py`](https://github.com/vllm-project/vllm/blob/main/vllm/pooling_params.py) controls embedding generation with fields for `dimensions` (Matryoshka support), `use_activation` (Tanh control), and task specification.
- **Request validation** occurs in the input processor and pooling params verifier, merging user inputs with model defaults and rejecting incompatible configurations.
- **Binary encoding** via `encoding_format: "base64"` optimizes large-scale retrieval throughput compared to standard JSON responses.

## Frequently Asked Questions

### What is the difference between the pooling runner and generation runner in vLLM?

The **pooling runner** processes inputs through the encoder and pooler to produce embedding vectors, while the **generation runner** routes tensors through the language-model decoder to generate text tokens. According to [`vllm/v1/worker/gpu_model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py), the engine checks `model_config.runner_type` to determine which path to execute, with pooling models setting `is_pooling_model = True` and bypassing the decoder entirely.

### How do I specify output dimensions for Matryoshka embeddings?

Pass the `dimensions` parameter in your request body or `PoolingParams` constructor. The `_merge_default_parameters` function in [`vllm/pooling_params.py`](https://github.com/vllm-project/vllm/blob/main/vllm/pooling_params.py) applies this value only if the model supports Matryoshka representations (nested vectors). For standard embedding models, specifying dimensions outside the model's fixed output size will raise a validation error during `PoolingParams.verify()`.

### Can I disable the activation function in the embedding pooler?

Yes. Set `use_activation: false` in your HTTP request or `use_activation=False` in the `PoolingParams` object. When set to `None` (default), the pooler uses its built-in default (typically `True` applying Tanh). Disabling activation returns raw pooler outputs, which may be necessary for specific downstream similarity calculations or when the model definition handles normalization internally.

### How does vLLM handle binary embedding responses?

When `encoding_format` is set to `"base64"`, the `encode_pooling_output_base64` function in [`vllm/entrypoints/pooling/utils.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/pooling/utils.py) serializes the `PoolingOutput` tensor into a compact binary stream. The response includes metadata headers describing the dtype and shape, while the body contains the raw bytes. Clients must decode this using appropriate struct unpacking (e.g., `struct.unpack("<256f", bytes)`) to reconstruct the float vectors.