How to Use Multi-Modal Models (LLaVA) with vLLM: A Complete Guide

vLLM supports LLaVA and other vision-language models through the MULTIMODAL_REGISTRY, which automatically handles image preprocessing, vision encoder inference, and placeholder token insertion when you pass image data via the image parameter in Python or the OpenAI-compatible API.

Running multi-modal models like LLaVA with vLLM enables high-throughput inference for vision-language tasks without manual preprocessing of image tensors. The vllm-project/vllm repository implements this through a registry-based architecture that wires model-specific processors to the inference engine.

How vLLM Handles Multi-Modal Inputs

vLLM treats vision-augmented language models as multimodal models through the MULTIMODAL_REGISTRY defined in vllm/multimodal/registry.py. This registry maps model architectures to processors that execute three critical steps: loading the vision encoder, converting raw images into hidden-state embeddings, and inserting placeholder tokens into the text prompt.

Model Registration and Processing Info

LlavaForConditionalGeneration registers itself with the multimodal registry in vllm/model_executor/models/llava.py【^1†L500-L506】. When you request a LLaVA checkpoint (e.g., llava-hf/llava-1.5-7b-hf), vLLM constructs a LlavaProcessingInfo object that extracts the HuggingFace LlavaConfig and initializes the vision tower via init_vision_tower_for_llava.

The vision tower uses either CLIPVisionModel or SiglipVisionModel depending on the checkpoint configuration, producing feature tensors of shape (batch*num_images, hidden_dim, height, width).

Projection and Placeholder Handling

Raw vision features pass through LlavaMultiModalProjector, a two-layer MLP defined in vllm/model_executor/models/llava.py【^1†L29-L61】. This projects features to the language model's hidden size.

LLaVA's tokenizer reserves a single <image> token. During preprocessing, the processor expands this token to N placeholders, where N equals self.info.get_num_image_tokens. This ensures the KV cache allocates correct dimensions before the first forward pass【^1†L78-L92】.

Loading and Configuring LLaVA Models

To run LLaVA with vLLM, initialize the LLM class with a HuggingFace-compatible checkpoint and configure multimodal limits.

from vllm import LLM

# Initialize with LLaVA 1.5 7B

llm = LLM(
    model="llava-hf/llava-1.5-7b-hf",
    max_model_len=4096,
    limit_mm_per_prompt={"image": 1}  # Enable one image per request

)

The limit_mm_per_prompt parameter controls how many images or videos the model accepts per prompt. For LLaVA-1.5, this is typically {"image": 1}, while LLaVA-NeXT and LLaVA-OneVision support multiple images or video frames.

Passing Images to LLaVA: Three Methods

vLLM exposes three ways to feed visual data to LLaVA models, defined in vllm/multimodal/inputs.py and documented in docs/features/multimodal_inputs.md.

1. Raw Pixel Values (Tensors)

Pass a torch.Tensor via the image keyword argument in generate() or chat(). The tensor must match the vision model's expected shape [batch*num_images, 3, H, W].

from vllm import SamplingParams
import torch
from PIL import Image

# Load and convert image

pil_image = Image.open("cat.jpg")

# Convert to tensor (C, H, W) then add batch dimension if needed

tensor = torch.from_numpy(np.array(pil_image)).permute(2, 0, 1).float() / 255.0

sampling_params = SamplingParams(max_tokens=64)
output = llm.generate(
    "USER: <image>\nWhat is shown in this picture?\nASSISTANT:",
    sampling_params,
    image=tensor.unsqueeze(0)  # Shape: (1, 3, H, W)

)

2. Pre-Computed Image Embeddings

If you have already run the vision encoder offline (e.g., for caching), supply the projected embeddings directly via image_embeds. This bypasses the vision tower and projector inside vLLM.


# Assuming image_embeds is a tensor of shape (num_image_tokens, hidden_size)

output = llm.generate(
    prompt,
    sampling_params,
    image_embeds=pre_computed_embeddings
)

3. File Paths and URLs

When using the OpenAI-compatible HTTP server, vLLM can fetch remote images or local files (when --allowed-local-media-path is configured) and automatically convert them to tensors. See the multimodal API documentation for the exact JSON schema【^1†L697-L704】.

Complete Code Examples

Python API with Local Image

This example follows the pattern in examples/offline_inference/vision_language.py:

from vllm import LLM, SamplingParams
from PIL import Image
from vllm.assets.image import ImageAsset

# Initialize model

llm = LLM(
    model="llava-hf/llava-1.5-7b-hf",
    max_model_len=4096,
    limit_mm_per_prompt={"image": 1}
)

# Load image

img = Image.open("examples/vision_language/cat.jpg")
tensor = ImageAsset.from_pil(img).as_tensor()  # shape (3, H, W)

# Prompt with automatic placeholder handling

prompt = "USER: <image>\nWhat is shown in this picture?\nASSISTANT:"

# Generate

sampling_params = SamplingParams(max_tokens=64)
output = llm.generate(
    prompt,
    sampling_params,
    image=tensor  # passes pixel_values to processor

)

print(output[0].outputs[0].text)

OpenAI-Compatible Server

Launch the server:

vllm serve llava-hf/llava-1.5-7b-hf \
    --max-model-len 4096 \
    --limit-mm-per-prompt '{"image":1}'

Client request:

import os
from openai import OpenAI

client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")

resp = client.chat.completions.create(
    model="llava-hf/llava-1.5-7b-hf",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "What animal is this?"},
                {
                    "type": "image_url",
                    "image_url": {"url": "https://example.com/cat.jpg"}
                }
            ]
        }
    ],
    max_tokens=64,
)

print(resp.choices[0].message.content)

Video Input with LLaVA-OneVision

LLaVA-OneVision is the only LLaVA variant that supports video. Configure the server for video:

vllm serve llava-hf/llava-onevision-qwen2-7b-ov-hf \
    --max-model-len 16384 \
    --limit-mm-per-prompt '{"video":1}'

Client request:

resp = client.chat.completions.create(
    model="llava-hf/llava-onevision-qwen2-7b-ov-hf",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "What is happening in this clip?"},
                {
                    "type": "video_url",
                    "video_url": {"url": "https://example.com/clip.mp4"}
                }
            ]
        }
    ],
    max_tokens=128,
)

Summary

  • vLLM supports LLaVA through the MULTIMODAL_REGISTRY in vllm/multimodal/registry.py, which automatically wires vision processors to language models.
  • The LLaVA architecture in vllm/model_executor/models/llava.py combines a vision tower (CLIPVisionModel or SiglipVisionModel), a two-layer MLP projector, and placeholder token expansion to fuse image embeddings with text.
  • Three input methods are supported: raw pixel_values tensors, pre-computed image_embeds, and file paths/URLs via the HTTP server.
  • Configuration requires setting limit_mm_per_prompt to specify how many images or videos the model accepts per request.
  • LLaVA-OneVision uniquely supports video inputs, while standard LLaVA 1.5 and NeXT focus on single or multiple image understanding.

Frequently Asked Questions

How does vLLM handle the image token placeholder expansion for LLaVA?

vLLM automatically expands the single <image> token reserved in LLaVA's tokenizer into the correct number of placeholder tokens based on the vision encoder output. The LlavaProcessingInfo class calculates the number of image tokens via get_num_image_tokens, ensuring the KV cache is correctly sized before the first forward pass in LlavaForConditionalGeneration【^1†L78-L92】.

Can I use pre-computed image embeddings instead of raw pixels with vLLM?

Yes. vLLM accepts pre-computed embeddings via the image_embeds parameter in the generate() method. This bypasses the internal vision tower (CLIPVisionModel or SiglipVisionModel) and the LlavaMultiModalProjector, allowing you to inject features directly into the language model's embedding layer. This is useful for caching vision features or using custom encoders.

What is the difference between LLaVA, LLaVA-NeXT, and LLaVA-OneVision in vLLM?

All three variants share the same core architecture in vllm/model_executor/models/llava.py but differ in capabilities and configuration. LLaVA 1.5 typically supports single images with a 336px or 224px resolution. LLaVA-NeXT (1.6) supports higher resolutions and multiple images per prompt. LLaVA-OneVision uniquely supports video inputs in addition to images, requiring the limit_mm_per_prompt={"video": 1} configuration and using video_url in the OpenAI-compatible API.

How do I configure the vLLM server to accept local image files?

When using the OpenAI-compatible HTTP server, you must explicitly allow local media paths using the --allowed-local-media-path flag. For example: vllm serve llava-hf/llava-1.5-7b-hf --allowed-local-media-path /path/to/images. Without this flag, the server only accepts remote URLs or base64-encoded images for security reasons, as documented in docs/features/multimodal_inputs.md【^1†L697-L704】.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →