How to Use Embedding Models and Pooling Parameters in vLLM: Complete Technical Guide
vLLM serves embedding models through a dedicated pooling runner that converts HTTP requests into PoolingParams instances, routing hidden states through the model's pooler instead of the language-model decoder to generate text embeddings.
vLLM supports high-performance inference for text embedding models through a specialized pooling architecture distinct from its generation pipeline. This guide explains how to configure embedding models and pooling parameters in vLLM based on the vllm-project/vllm source code, covering the complete request lifecycle from HTTP API to GPU execution.
Understanding the Pooling Runner Architecture
When you start vLLM with the --runner pooling flag, the engine initializes a pooling runner rather than a standard language-model decoder. In vllm/v1/worker/gpu_model_runner.py (line 408), the worker checks self.is_pooling_model (determined by model_config.runner_type == "pooling") to route tensors through the encoder and pooler while bypassing the generation head.
This architecture is essential for models like BAAI/bge-m3 that lack a language-model head. The engine detects this configuration via model_config.pooler_config and sets max_output_tokens to zero, treating the request as embedding-only inference.
The PoolingParams Configuration Object
The PoolingParams class defined in vllm/pooling_params.py controls how the pooler processes hidden states. When a request arrives at the /pooling endpoint, the to_pooling_params() method in vllm/entrypoints/pooling/embed/protocol.py (lines 59-64) constructs this object with the following fields:
task: Fixed to"embed"for text embedding models. Set automatically by the protocol layer.dimensions: Target output dimensionality. Only honored when the model supports Matryoshka (nested-vector) representations.use_activation: Controls whether the pooler applies its activation function (typicallyTanh). Defaults toTruewhen set toNone.step_tag_idandreturned_token_ids: Reserved for step-wise pooling tasks such as token-wise reranking, not used for standard embeddings.
The _merge_default_parameters function in vllm/pooling_params.py (line 85) merges user-provided values with model-level defaults from ModelConfig.pooler_config. Invalid parameter combinations are rejected by _verify_valid_parameters (line 68), ensuring that unsupported options (like dimensions on non-Matryoshka models) raise errors before execution.
Request Flow: From HTTP to Embedding Tensor
HTTP Request Parsing
Incoming requests hit vllm/entrypoints/pooling/embed/protocol.py, where EmbeddingCompletionRequest or EmbeddingChatRequest objects parse the JSON payload. The to_pooling_params() method extracts dimensions and use_activation fields to build the PoolingParams instance.
Parameter Validation
In vllm/v1/engine/input_processor.py, the engine calls PoolingParams.verify() to validate and normalize parameters against the model's capabilities. This verification step ensures that the pooling task is compatible with the provided modifiers before dispatching to the GPU worker.
Engine Orchestration
The LLMEngine in vllm/v1/engine/llm_engine.py forwards validated PoolingParams to the worker. For programmatic usage, the LLM.encode() method internally constructs and validates these parameters before execution.
GPU Execution
The gpu_model_runner.py worker checks is_pooling_model and processes inputs through the encoder network. Hidden states pass through the model's pooler rather than the decoder, producing a PoolingOutput tensor with shape [batch, dim].
Response Encoding
The serving layer in vllm/entrypoints/pooling/embed/serving.py converts PoolingRequestOutput objects into OpenAI-compatible responses. Utility functions in vllm/entrypoints/pooling/utils.py handle encoding:
encode_pooling_output_float: Returns JSON arrays of floating-point values.encode_pooling_output_base64: Returns base64-encoded binary for efficient large-scale retrieval.
Practical Implementation Examples
Starting a Pooling Server
Launch the server with the pooling runner flag to enable embedding endpoints:
vllm serve BAAI/bge-m3 --runner pooling --port 8000
The --runner pooling argument is required to activate the embedding pipeline instead of the generation path.
Sending HTTP Requests with Pooling Parameters
Send POST requests to the /pooling endpoint with optional dimension and activation controls:
import json
import requests
url = "http://localhost:8000/pooling"
payload = {
"model": "BAAI/bge-m3",
"input": [
"Hello, my name is",
"The capital of France is Paris."
],
"dimensions": 256,
"use_activation": True,
"encoding_format": "float"
}
resp = requests.post(
url,
json=payload,
headers={"User-Agent": "vllm-client"}
)
data = resp.json()["data"]
embeddings = [entry["embedding"] for entry in data]
print(json.dumps(embeddings, indent=2))
The dimensions and use_activation fields map directly to PoolingParams attributes, allowing dynamic control over the output vector size and activation application.
Using the Python Client Library
For integrated Python applications, use the LLM class with explicit PoolingParams:
from vllm import LLM, PoolingParams
# Initialize engine with pooling runner
engine = LLM(
model="BAAI/bge-m3",
runner_type="pooling"
)
# Configure pooling parameters
pool_params = PoolingParams(
task="embed",
dimensions=256,
use_activation=False
)
# Generate embeddings
outputs = engine.encode(
["I love AI.", "vLLM makes inference fast."],
pooling_params=pool_params
)
# Access tensor output
print(outputs[0].embedding) # Shape: (256,)
The engine.encode() method handles PoolingParams validation internally and returns PoolingRequestOutput objects containing the embedding tensors.
Binary Encoding for Large-Scale Retrieval
For production retrieval systems requiring minimal bandwidth, request base64-encoded binary responses:
import requests
import base64
import struct
url = "http://localhost:8000/pooling"
payload = {
"model": "BAAI/bge-m3",
"input": ["query 1", "query 2"],
"encoding_format": "base64",
"embed_dtype": "float32",
"endianness": "little"
}
resp = requests.post(url, json=payload)
metadata = resp.headers.get("metadata")
# Decode binary payload (example: 256-dim float32)
binary = resp.content
first_vec = struct.unpack("<256f", binary[:1024])
print(first_vec[:5])
Binary encoding is processed by encode_pooling_output_base64 in vllm/entrypoints/pooling/utils.py, significantly reducing response size compared to JSON arrays.
Summary
- vLLM uses a pooling runner (activated via
--runner pooling) to serve embedding models, routing tensors through the pooler rather than the language-model decoder. PoolingParamsinvllm/pooling_params.pycontrols embedding generation with fields fordimensions(Matryoshka support),use_activation(Tanh control), and task specification.- Request validation occurs in the input processor and pooling params verifier, merging user inputs with model defaults and rejecting incompatible configurations.
- Binary encoding via
encoding_format: "base64"optimizes large-scale retrieval throughput compared to standard JSON responses.
Frequently Asked Questions
What is the difference between the pooling runner and generation runner in vLLM?
The pooling runner processes inputs through the encoder and pooler to produce embedding vectors, while the generation runner routes tensors through the language-model decoder to generate text tokens. According to vllm/v1/worker/gpu_model_runner.py, the engine checks model_config.runner_type to determine which path to execute, with pooling models setting is_pooling_model = True and bypassing the decoder entirely.
How do I specify output dimensions for Matryoshka embeddings?
Pass the dimensions parameter in your request body or PoolingParams constructor. The _merge_default_parameters function in vllm/pooling_params.py applies this value only if the model supports Matryoshka representations (nested vectors). For standard embedding models, specifying dimensions outside the model's fixed output size will raise a validation error during PoolingParams.verify().
Can I disable the activation function in the embedding pooler?
Yes. Set use_activation: false in your HTTP request or use_activation=False in the PoolingParams object. When set to None (default), the pooler uses its built-in default (typically True applying Tanh). Disabling activation returns raw pooler outputs, which may be necessary for specific downstream similarity calculations or when the model definition handles normalization internally.
How does vLLM handle binary embedding responses?
When encoding_format is set to "base64", the encode_pooling_output_base64 function in vllm/entrypoints/pooling/utils.py serializes the PoolingOutput tensor into a compact binary stream. The response includes metadata headers describing the dtype and shape, while the body contains the raw bytes. Clients must decode this using appropriate struct unpacking (e.g., struct.unpack("<256f", bytes)) to reconstruct the float vectors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →