vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

21 articles 71.8k View on GitHub ↗
21 articles
How to Implement Custom Logit Processors and Sampling Parameters in vLLM

Learn to implement custom logit processors and sampling parameters in vLLM. Extend vLLM's capabilities by subclassing LogitsProcessor and registering your custom logic for advanced text generation control.

how-to-guide
Mar 3, 2026
How to Configure Compilation Modes (Eager, CUDAGraph, Inductor) in vLLM

Master vLLM compilation modes eager CUDAGraph and Inductor. Learn to configure settings via CLI or Python API for optimized performance.

deep-dive
Mar 3, 2026
How to Debug Memory Issues and OOM Errors in vLLM: A Complete Guide

Debug vLLM memory issues and OOM errors with our guide. Isolate problems in workspace buffers, KV-cache, or activations using memory profiling tools.

how-to-guide
Mar 3, 2026
How to Use the gRPC API for High-Performance Serving in vLLM

Learn to use the vLLM gRPC API for high-performance serving. Stream tokens efficiently with unlimited message sizes and structured output support for powerful language model inference.

how-to-guide
Mar 3, 2026
How to Configure KV Cache Allocation and Management in vLLM

Master KV cache allocation and management in vLLM. Learn to set CacheConfig parameters like block_size and gpu_memory_utilization for optimized tensor allocation via CLI or Python.

how-to-guide
Mar 3, 2026
How to Use Multi-Modal Models (LLaVA) with vLLM: A Complete Guide

Integrate LLaVA multi-modal models seamlessly with vLLM. Our guide simplifies image preprocessing and vision encoder inference for efficient LLM deployment.

how-to-guide
Mar 3, 2026
vLLM v1 Engine vs Legacy Engine: Key Architectural Differences Explained

Explore the vLLM v1 engine vs legacy engine architectural differences. Discover its unified scheduler, `EngineCore` loop, and removal of legacy features for enhanced performance.

architecture
Mar 3, 2026
How to Configure Attention Backends (FlashAttention and FlashInfer) in vLLM

Learn to configure attention backends like FlashAttention and FlashInfer in vLLM using CLI flags API or config files for faster inference. Optimize your LLM performance effortlessly.

how-to-guide
Mar 3, 2026
How to Use Embedding Models and Pooling Parameters in vLLM: Complete Technical Guide

Learn to use embedding models and pooling parameters in vLLM. This guide details how vLLM's pooling runner generates text embeddings efficiently by routing hidden states through the model's pooler.

how-to-guide
Mar 3, 2026
How Expert Parallelism Works for Mixture-of-Experts Models in vLLM

Discover how expert parallelism in vLLM shards expert weights across GPUs for scalable distributed inference of Mixture-of-Experts models. Optimize your MoE inference.

deep-dive
Mar 3, 2026
How to Optimize GPU Memory Utilization with gpu_memory_utilization in vLLM

Optimize GPU memory utilization in vLLM by setting the gpu_memory_utilization parameter. Balance throughput and OOM risk by adjusting this value. Learn how now.

how-to-guide
Mar 3, 2026
How to Implement Beam Search and Parallel Sampling in vLLM

Learn to implement beam search and parallel sampling in vLLM. Optimize your LLM inference with efficient batching and GPU acceleration for faster requests.

how-to-guide
Mar 3, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →