# How to Use Multi-Modal Models (LLaVA) with vLLM: A Complete Guide

> Integrate LLaVA multi-modal models seamlessly with vLLM. Our guide simplifies image preprocessing and vision encoder inference for efficient LLM deployment.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**vLLM supports LLaVA and other vision-language models through the `MULTIMODAL_REGISTRY`, which automatically handles image preprocessing, vision encoder inference, and placeholder token insertion when you pass image data via the `image` parameter in Python or the OpenAI-compatible API.**

Running multi-modal models like LLaVA with vLLM enables high-throughput inference for vision-language tasks without manual preprocessing of image tensors. The vllm-project/vllm repository implements this through a registry-based architecture that wires model-specific processors to the inference engine.

## How vLLM Handles Multi-Modal Inputs

vLLM treats vision-augmented language models as **multimodal** models through the `MULTIMODAL_REGISTRY` defined in [`vllm/multimodal/registry.py`](https://github.com/vllm-project/vllm/blob/main/vllm/multimodal/registry.py). This registry maps model architectures to processors that execute three critical steps: loading the vision encoder, converting raw images into hidden-state embeddings, and inserting placeholder tokens into the text prompt.

### Model Registration and Processing Info

`LlavaForConditionalGeneration` registers itself with the multimodal registry in [`vllm/model_executor/models/llava.py`](https://github.com/vllm-project/vllm/blob/main/vllm/model_executor/models/llava.py)【^1†L500-L506】. When you request a LLaVA checkpoint (e.g., `llava-hf/llava-1.5-7b-hf`), vLLM constructs a `LlavaProcessingInfo` object that extracts the HuggingFace `LlavaConfig` and initializes the vision tower via `init_vision_tower_for_llava`.

The vision tower uses either `CLIPVisionModel` or `SiglipVisionModel` depending on the checkpoint configuration, producing feature tensors of shape `(batch*num_images, hidden_dim, height, width)`.

### Projection and Placeholder Handling

Raw vision features pass through `LlavaMultiModalProjector`, a two-layer MLP defined in [`vllm/model_executor/models/llava.py`](https://github.com/vllm-project/vllm/blob/main/vllm/model_executor/models/llava.py)【^1†L29-L61】. This projects features to the language model's hidden size.

LLaVA's tokenizer reserves a single `<image>` token. During preprocessing, the processor expands this token to *N* placeholders, where *N* equals `self.info.get_num_image_tokens`. This ensures the KV cache allocates correct dimensions before the first forward pass【^1†L78-L92】.

## Loading and Configuring LLaVA Models

To run LLaVA with vLLM, initialize the `LLM` class with a HuggingFace-compatible checkpoint and configure multimodal limits.

```python
from vllm import LLM

# Initialize with LLaVA 1.5 7B

llm = LLM(
    model="llava-hf/llava-1.5-7b-hf",
    max_model_len=4096,
    limit_mm_per_prompt={"image": 1}  # Enable one image per request

)

```

The `limit_mm_per_prompt` parameter controls how many images or videos the model accepts per prompt. For LLaVA-1.5, this is typically `{"image": 1}`, while LLaVA-NeXT and LLaVA-OneVision support multiple images or video frames.

## Passing Images to LLaVA: Three Methods

vLLM exposes three ways to feed visual data to LLaVA models, defined in [`vllm/multimodal/inputs.py`](https://github.com/vllm-project/vllm/blob/main/vllm/multimodal/inputs.py) and documented in [`docs/features/multimodal_inputs.md`](https://github.com/vllm-project/vllm/blob/main/docs/features/multimodal_inputs.md).

### 1. Raw Pixel Values (Tensors)

Pass a `torch.Tensor` via the `image` keyword argument in `generate()` or `chat()`. The tensor must match the vision model's expected shape `[batch*num_images, 3, H, W]`.

```python
from vllm import SamplingParams
import torch
from PIL import Image

# Load and convert image

pil_image = Image.open("cat.jpg")

# Convert to tensor (C, H, W) then add batch dimension if needed

tensor = torch.from_numpy(np.array(pil_image)).permute(2, 0, 1).float() / 255.0

sampling_params = SamplingParams(max_tokens=64)
output = llm.generate(
    "USER: <image>\nWhat is shown in this picture?\nASSISTANT:",
    sampling_params,
    image=tensor.unsqueeze(0)  # Shape: (1, 3, H, W)

)

```

### 2. Pre-Computed Image Embeddings

If you have already run the vision encoder offline (e.g., for caching), supply the projected embeddings directly via `image_embeds`. This bypasses the vision tower and projector inside vLLM.

```python

# Assuming image_embeds is a tensor of shape (num_image_tokens, hidden_size)

output = llm.generate(
    prompt,
    sampling_params,
    image_embeds=pre_computed_embeddings
)

```

### 3. File Paths and URLs

When using the OpenAI-compatible HTTP server, vLLM can fetch remote images or local files (when `--allowed-local-media-path` is configured) and automatically convert them to tensors. See the multimodal API documentation for the exact JSON schema【^1†L697-L704】.

## Complete Code Examples

### Python API with Local Image

This example follows the pattern in [`examples/offline_inference/vision_language.py`](https://github.com/vllm-project/vllm/blob/main/examples/offline_inference/vision_language.py):

```python
from vllm import LLM, SamplingParams
from PIL import Image
from vllm.assets.image import ImageAsset

# Initialize model

llm = LLM(
    model="llava-hf/llava-1.5-7b-hf",
    max_model_len=4096,
    limit_mm_per_prompt={"image": 1}
)

# Load image

img = Image.open("examples/vision_language/cat.jpg")
tensor = ImageAsset.from_pil(img).as_tensor()  # shape (3, H, W)

# Prompt with automatic placeholder handling

prompt = "USER: <image>\nWhat is shown in this picture?\nASSISTANT:"

# Generate

sampling_params = SamplingParams(max_tokens=64)
output = llm.generate(
    prompt,
    sampling_params,
    image=tensor  # passes pixel_values to processor

)

print(output[0].outputs[0].text)

```

### OpenAI-Compatible Server

Launch the server:

```bash
vllm serve llava-hf/llava-1.5-7b-hf \
    --max-model-len 4096 \
    --limit-mm-per-prompt '{"image":1}'

```

Client request:

```python
import os
from openai import OpenAI

client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")

resp = client.chat.completions.create(
    model="llava-hf/llava-1.5-7b-hf",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "What animal is this?"},
                {
                    "type": "image_url",
                    "image_url": {"url": "https://example.com/cat.jpg"}
                }
            ]
        }
    ],
    max_tokens=64,
)

print(resp.choices[0].message.content)

```

### Video Input with LLaVA-OneVision

LLaVA-OneVision is the only LLaVA variant that supports video. Configure the server for video:

```bash
vllm serve llava-hf/llava-onevision-qwen2-7b-ov-hf \
    --max-model-len 16384 \
    --limit-mm-per-prompt '{"video":1}'

```

Client request:

```python
resp = client.chat.completions.create(
    model="llava-hf/llava-onevision-qwen2-7b-ov-hf",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "What is happening in this clip?"},
                {
                    "type": "video_url",
                    "video_url": {"url": "https://example.com/clip.mp4"}
                }
            ]
        }
    ],
    max_tokens=128,
)

```

## Summary

- **vLLM supports LLaVA through the `MULTIMODAL_REGISTRY`** in [`vllm/multimodal/registry.py`](https://github.com/vllm-project/vllm/blob/main/vllm/multimodal/registry.py), which automatically wires vision processors to language models.
- **The LLaVA architecture** in [`vllm/model_executor/models/llava.py`](https://github.com/vllm-project/vllm/blob/main/vllm/model_executor/models/llava.py) combines a vision tower (`CLIPVisionModel` or `SiglipVisionModel`), a two-layer MLP projector, and placeholder token expansion to fuse image embeddings with text.
- **Three input methods** are supported: raw `pixel_values` tensors, pre-computed `image_embeds`, and file paths/URLs via the HTTP server.
- **Configuration** requires setting `limit_mm_per_prompt` to specify how many images or videos the model accepts per request.
- **LLaVA-OneVision** uniquely supports video inputs, while standard LLaVA 1.5 and NeXT focus on single or multiple image understanding.

## Frequently Asked Questions

### How does vLLM handle the image token placeholder expansion for LLaVA?

vLLM automatically expands the single `<image>` token reserved in LLaVA's tokenizer into the correct number of placeholder tokens based on the vision encoder output. The `LlavaProcessingInfo` class calculates the number of image tokens via `get_num_image_tokens`, ensuring the KV cache is correctly sized before the first forward pass in `LlavaForConditionalGeneration`【^1†L78-L92】.

### Can I use pre-computed image embeddings instead of raw pixels with vLLM?

Yes. vLLM accepts pre-computed embeddings via the `image_embeds` parameter in the `generate()` method. This bypasses the internal vision tower (`CLIPVisionModel` or `SiglipVisionModel`) and the `LlavaMultiModalProjector`, allowing you to inject features directly into the language model's embedding layer. This is useful for caching vision features or using custom encoders.

### What is the difference between LLaVA, LLaVA-NeXT, and LLaVA-OneVision in vLLM?

All three variants share the same core architecture in [`vllm/model_executor/models/llava.py`](https://github.com/vllm-project/vllm/blob/main/vllm/model_executor/models/llava.py) but differ in capabilities and configuration. **LLaVA 1.5** typically supports single images with a 336px or 224px resolution. **LLaVA-NeXT** (1.6) supports higher resolutions and multiple images per prompt. **LLaVA-OneVision** uniquely supports video inputs in addition to images, requiring the `limit_mm_per_prompt={"video": 1}` configuration and using `video_url` in the OpenAI-compatible API.

### How do I configure the vLLM server to accept local image files?

When using the OpenAI-compatible HTTP server, you must explicitly allow local media paths using the `--allowed-local-media-path` flag. For example: `vllm serve llava-hf/llava-1.5-7b-hf --allowed-local-media-path /path/to/images`. Without this flag, the server only accepts remote URLs or base64-encoded images for security reasons, as documented in [`docs/features/multimodal_inputs.md`](https://github.com/vllm-project/vllm/blob/main/docs/features/multimodal_inputs.md)【^1†L697-L704】.