How to Use the gRPC API for High-Performance Serving in vLLM
vLLM provides a native gRPC server via VllmEngineServicer that exposes the AsyncLLM engine through six RPC methods, enabling binary-efficient, token-streaming inference with unlimited message sizes and full structured output support.
The vLLM project ships with a production-ready gRPC entrypoint that bypasses HTTP overhead for high-throughput workloads. This implementation leverages protobuf-defined contracts and asynchronous I/O to serve thousands of concurrent generation streams with minimal latency.
Architecture of the vLLM gRPC Server
The gRPC stack in vLLM consists of tightly integrated layers that bridge the protobuf interface with the core inference engine.
Core Components
vllm/grpc/vllm_engine.proto defines the service contract. This file specifies the VllmEngine service, request/response messages (like GenerateRequest and GenerateResponse), and the SamplingParams structure. It provides a language-agnostic binary protocol that clients can use from Python, Go, Rust, or any gRPC-compatible runtime.
vllm/entrypoints/grpc_server.py serves as the CLI entry point. This module parses arguments, instantiates the AsyncLLM engine, and constructs a grpc.aio server backed by uvloop. It registers the VllmEngineServicer implementation and configures unlimited message sizes via grpc.max_send_message_length and grpc.max_receive_message_length set to -1.
VllmEngineServicer (defined within grpc_server.py) implements six RPC methods:
Generate– Streams token chunks or returns complete responsesEmbed– Handles embedding requestsHealthCheck– Reports engine readinessAbort– Cancels in-flight request IDsGetModelInfo– Exposes model path and context length limitsGetServerInfo– Returns active request counts and uptime
AsyncLLM (from vllm/v1/engine/async_llm.py) powers the backend. The servicer forwards decoded prompts to async_llm.generate(), which yields RequestOutput objects that the servicer converts into protobuf responses.
vllm/grpc/compile_protos.py regenerates Python bindings. This utility compiles vllm_engine.proto into vllm_engine_pb2.py and vllm_engine_pb2_grpc.py, which both the server and clients import.
Request-Response Flow for Generation
When a client calls the Generate RPC, the system executes the following pipeline:
-
Request Decoding: The servicer receives a
GenerateRequestcontaining either aTokenizedInputor raw text, plusSamplingParams(temperature, top-p, stop strings). -
Engine Invocation: The servicer calls
self.async_llm.generate(prompt, sampling_params, ...)and receives an asynchronous generator. -
Streaming Emission: For each
RequestOutputfrom the engine:- If streaming is enabled,
_chunk_responseemits aGenerateResponse.chunkcontaining only newly generated token IDs (completion.token_ids) and statistics (prompt_tokens,completion_tokens,cached_tokens). - When
output.finishedisTrue,_complete_responseemits a finalGenerateResponse.completemessage.
- If streaming is enabled,
-
Client Consumption: The client iterates over the
Generateresponse stream, concatenating token IDs or applying detokenization logic to reconstruct text.
Starting the gRPC Server
Launch the server using the module entry point. The following command starts the servicer on all interfaces with the specified model:
python -m vllm.entrypoints.grpc_server \
--model meta-llama/Llama-2-7b-chat-hf \
--host 0.0.0.0 \
--port 50051
The server initializes the AsyncLLM engine, loads the model weights, and begins listening for gRPC connections on port 50051.
Client Implementation Examples
The generated protobuf modules in vllm/grpc/ provide the necessary stubs for client development.
Non-Streaming Generation
For synchronous-style completion where the server returns a single final response:
import asyncio
import grpc
from vllm.grpc import vllm_engine_pb2, vllm_engine_pb2_grpc
async def generate_once():
async with grpc.aio.insecure_channel("localhost:50051") as channel:
stub = vllm_engine_pb2_grpc.VllmEngineStub(channel)
request = vllm_engine_pb2.GenerateRequest(
request_id="demo-1",
text="The capital of Italy is",
sampling_params=vllm_engine_pb2.SamplingParams(
temperature=0.0,
max_tokens=5,
n=1,
),
stream=False,
)
async for resp in stub.Generate(request):
if resp.HasField("complete"):
output_ids = resp.complete.output_ids
print("Token IDs:", output_ids)
asyncio.run(generate_once())
Streaming Token Generation
For real-time applications requiring per-token latency, enable streaming to receive delta updates:
import asyncio
import grpc
from vllm.grpc import vllm_engine_pb2, vllm_engine_pb2_grpc
async def generate_stream():
async with grpc.aio.insecure_channel("localhost:50051") as channel:
stub = vllm_engine_pb2_grpc.VllmEngineStub(channel)
request = vllm_engine_pb2.GenerateRequest(
request_id="stream-1",
text="The quick brown fox",
sampling_params=vllm_engine_pb2.SamplingParams(
temperature=0.7,
max_tokens=12,
n=1,
),
stream=True,
)
async for resp in stub.Generate(request):
if resp.HasField("chunk"):
chunk = resp.chunk
print(f"New tokens: {list(chunk.token_ids)} "
f"(prompt={chunk.prompt_tokens}, "
f"completion={chunk.completion_tokens})")
elif resp.HasField("complete"):
print("Generation finished")
break
asyncio.run(generate_stream())
Health and Metadata RPCs
Use utility RPCs for orchestration and load balancing:
import asyncio
import grpc
from vllm.grpc import vllm_engine_pb2, vllm_engine_pb2_grpc
async def meta():
async with grpc.aio.insecure_channel("localhost:50051") as ch:
stub = vllm_engine_pb2_grpc.VllmEngineStub(ch)
health = await stub.HealthCheck(vllm_engine_pb2.HealthCheckRequest())
print("Healthy:", health.healthy, health.message)
model_info = await stub.GetModelInfo(vllm_engine_pb2.GetModelInfoRequest())
print("Model path:", model_info.model_path,
"max ctx:", model_info.max_context_length)
server_info = await stub.GetServerInfo(vllm_engine_pb2.GetServerInfoRequest())
print("Active requests:", server_info.active_requests,
"uptime:", server_info.uptime_seconds)
asyncio.run(meta())
Aborting In-Flight Requests
Cancel specific request IDs to free up engine capacity:
abort_req = vllm_engine_pb2.AbortRequest(request_ids=["stream-1"])
await stub.Abort(abort_req)
Key Design Features for Performance
The vLLM gRPC implementation optimizes for high-throughput serving through several architectural decisions.
AsyncIO-First Execution: All RPC methods in VllmEngineServicer use grpc.aio and async/await syntax. This allows the server to manage thousands of concurrent generation streams without blocking threads, maximizing GPU utilization.
Delta-Mode Streaming: When stream=True, the server transmits only new token IDs (delta) rather than the full generated sequence. This reduces bandwidth consumption proportionally to sequence length, making it ideal for long-form generation.
Unlimited Message Sizes: The server explicitly sets grpc.max_send_message_length and grpc.max_receive_message_length to -1, removing protobuf size limits. This accommodates very long contexts and large batch requests without truncation.
Structured Output Constraints: The protobuf schema defines a oneof constraint field supporting JSON schema, regex, and grammar constraints. The servicer maps these to StructuredOutputsParams via _sampling_params_from_proto, enabling type-safe generation without client-side validation.
Detokenization Control: SamplingParams.detokenize is forced to True only when stop strings are present. Otherwise, the server returns raw token IDs, allowing clients to handle detokenization or downstream processing with zero server overhead.
Summary
- Entry Point: Launch the server via
python -m vllm.entrypoints.grpc_serverwith standard vLLM arguments. - Service Definition: The
VllmEngineservice is defined invllm/grpc/vllm_engine.protoand implemented byVllmEngineServicer. - Core Methods: Use
Generatefor text completion (streaming or non-streaming),Embedfor embeddings, andAbortto cancel requests. - Performance: The architecture uses
AsyncLLM,grpc.aio, delta streaming, and unlimited message sizes to maximize throughput. - Client Access: Import generated modules from
vllm.grpcto build clients in Python, or use the.protofile for other languages.
Frequently Asked Questions
How do I regenerate the protobuf Python files if I modify the proto schema?
Run vllm/grpc/compile_protos.py. This script invokes the protobuf compiler to regenerate vllm_engine_pb2.py and vllm_engine_pb2_grpc.py from vllm_engine.proto, ensuring the server and clients stay synchronized with your schema changes.
Can I use the gRPC API with tensor parallelism or pipeline parallelism?
Yes. The AsyncLLM engine initialized in grpc_server.py respects all standard vLLM arguments including --tensor-parallel-size and --pipeline-parallel-size. The gRPC servicer operates above the engine layer and is agnostic to the underlying distributed configuration.
What is the difference between the HTTP OpenAI-compatible server and the gRPC server?
The HTTP server (vllm.entrypoints.openai.api_server) provides OpenAI-compatible REST endpoints for easy integration with existing tools. The gRPC server (vllm.entrypoints.grpc_server) offers binary protobuf encoding, bidirectional streaming, and unlimited message sizes, making it superior for high-throughput, low-latency internal microservices.
How does the server handle very long contexts or large batches?
The gRPC server sets maximum message length parameters to -1 (unlimited) and relies on the AsyncLLM engine's continuous batching and PagedAttention memory management. This combination allows processing of extremely long sequences without the size limitations typically imposed by HTTP JSON payloads.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →