How to Use the gRPC API for High-Performance Serving in vLLM

vLLM provides a native gRPC server via VllmEngineServicer that exposes the AsyncLLM engine through six RPC methods, enabling binary-efficient, token-streaming inference with unlimited message sizes and full structured output support.

The vLLM project ships with a production-ready gRPC entrypoint that bypasses HTTP overhead for high-throughput workloads. This implementation leverages protobuf-defined contracts and asynchronous I/O to serve thousands of concurrent generation streams with minimal latency.

Architecture of the vLLM gRPC Server

The gRPC stack in vLLM consists of tightly integrated layers that bridge the protobuf interface with the core inference engine.

Core Components

vllm/grpc/vllm_engine.proto defines the service contract. This file specifies the VllmEngine service, request/response messages (like GenerateRequest and GenerateResponse), and the SamplingParams structure. It provides a language-agnostic binary protocol that clients can use from Python, Go, Rust, or any gRPC-compatible runtime.

vllm/entrypoints/grpc_server.py serves as the CLI entry point. This module parses arguments, instantiates the AsyncLLM engine, and constructs a grpc.aio server backed by uvloop. It registers the VllmEngineServicer implementation and configures unlimited message sizes via grpc.max_send_message_length and grpc.max_receive_message_length set to -1.

VllmEngineServicer (defined within grpc_server.py) implements six RPC methods:

  • Generate – Streams token chunks or returns complete responses
  • Embed – Handles embedding requests
  • HealthCheck – Reports engine readiness
  • Abort – Cancels in-flight request IDs
  • GetModelInfo – Exposes model path and context length limits
  • GetServerInfo – Returns active request counts and uptime

AsyncLLM (from vllm/v1/engine/async_llm.py) powers the backend. The servicer forwards decoded prompts to async_llm.generate(), which yields RequestOutput objects that the servicer converts into protobuf responses.

vllm/grpc/compile_protos.py regenerates Python bindings. This utility compiles vllm_engine.proto into vllm_engine_pb2.py and vllm_engine_pb2_grpc.py, which both the server and clients import.

Request-Response Flow for Generation

When a client calls the Generate RPC, the system executes the following pipeline:

  1. Request Decoding: The servicer receives a GenerateRequest containing either a TokenizedInput or raw text, plus SamplingParams (temperature, top-p, stop strings).

  2. Engine Invocation: The servicer calls self.async_llm.generate(prompt, sampling_params, ...) and receives an asynchronous generator.

  3. Streaming Emission: For each RequestOutput from the engine:

    • If streaming is enabled, _chunk_response emits a GenerateResponse.chunk containing only newly generated token IDs (completion.token_ids) and statistics (prompt_tokens, completion_tokens, cached_tokens).
    • When output.finished is True, _complete_response emits a final GenerateResponse.complete message.
  4. Client Consumption: The client iterates over the Generate response stream, concatenating token IDs or applying detokenization logic to reconstruct text.

Starting the gRPC Server

Launch the server using the module entry point. The following command starts the servicer on all interfaces with the specified model:

python -m vllm.entrypoints.grpc_server \
    --model meta-llama/Llama-2-7b-chat-hf \
    --host 0.0.0.0 \
    --port 50051

The server initializes the AsyncLLM engine, loads the model weights, and begins listening for gRPC connections on port 50051.

Client Implementation Examples

The generated protobuf modules in vllm/grpc/ provide the necessary stubs for client development.

Non-Streaming Generation

For synchronous-style completion where the server returns a single final response:

import asyncio
import grpc
from vllm.grpc import vllm_engine_pb2, vllm_engine_pb2_grpc

async def generate_once():
    async with grpc.aio.insecure_channel("localhost:50051") as channel:
        stub = vllm_engine_pb2_grpc.VllmEngineStub(channel)

        request = vllm_engine_pb2.GenerateRequest(
            request_id="demo-1",
            text="The capital of Italy is",
            sampling_params=vllm_engine_pb2.SamplingParams(
                temperature=0.0,
                max_tokens=5,
                n=1,
            ),
            stream=False,
        )

        async for resp in stub.Generate(request):
            if resp.HasField("complete"):
                output_ids = resp.complete.output_ids
                print("Token IDs:", output_ids)

asyncio.run(generate_once())

Streaming Token Generation

For real-time applications requiring per-token latency, enable streaming to receive delta updates:

import asyncio
import grpc
from vllm.grpc import vllm_engine_pb2, vllm_engine_pb2_grpc

async def generate_stream():
    async with grpc.aio.insecure_channel("localhost:50051") as channel:
        stub = vllm_engine_pb2_grpc.VllmEngineStub(channel)

        request = vllm_engine_pb2.GenerateRequest(
            request_id="stream-1",
            text="The quick brown fox",
            sampling_params=vllm_engine_pb2.SamplingParams(
                temperature=0.7,
                max_tokens=12,
                n=1,
            ),
            stream=True,
        )

        async for resp in stub.Generate(request):
            if resp.HasField("chunk"):
                chunk = resp.chunk
                print(f"New tokens: {list(chunk.token_ids)} "
                      f"(prompt={chunk.prompt_tokens}, "
                      f"completion={chunk.completion_tokens})")
            elif resp.HasField("complete"):
                print("Generation finished")
                break

asyncio.run(generate_stream())

Health and Metadata RPCs

Use utility RPCs for orchestration and load balancing:

import asyncio
import grpc
from vllm.grpc import vllm_engine_pb2, vllm_engine_pb2_grpc

async def meta():
    async with grpc.aio.insecure_channel("localhost:50051") as ch:
        stub = vllm_engine_pb2_grpc.VllmEngineStub(ch)

        health = await stub.HealthCheck(vllm_engine_pb2.HealthCheckRequest())
        print("Healthy:", health.healthy, health.message)

        model_info = await stub.GetModelInfo(vllm_engine_pb2.GetModelInfoRequest())
        print("Model path:", model_info.model_path,
              "max ctx:", model_info.max_context_length)

        server_info = await stub.GetServerInfo(vllm_engine_pb2.GetServerInfoRequest())
        print("Active requests:", server_info.active_requests,
              "uptime:", server_info.uptime_seconds)

asyncio.run(meta())

Aborting In-Flight Requests

Cancel specific request IDs to free up engine capacity:

abort_req = vllm_engine_pb2.AbortRequest(request_ids=["stream-1"])
await stub.Abort(abort_req)

Key Design Features for Performance

The vLLM gRPC implementation optimizes for high-throughput serving through several architectural decisions.

AsyncIO-First Execution: All RPC methods in VllmEngineServicer use grpc.aio and async/await syntax. This allows the server to manage thousands of concurrent generation streams without blocking threads, maximizing GPU utilization.

Delta-Mode Streaming: When stream=True, the server transmits only new token IDs (delta) rather than the full generated sequence. This reduces bandwidth consumption proportionally to sequence length, making it ideal for long-form generation.

Unlimited Message Sizes: The server explicitly sets grpc.max_send_message_length and grpc.max_receive_message_length to -1, removing protobuf size limits. This accommodates very long contexts and large batch requests without truncation.

Structured Output Constraints: The protobuf schema defines a oneof constraint field supporting JSON schema, regex, and grammar constraints. The servicer maps these to StructuredOutputsParams via _sampling_params_from_proto, enabling type-safe generation without client-side validation.

Detokenization Control: SamplingParams.detokenize is forced to True only when stop strings are present. Otherwise, the server returns raw token IDs, allowing clients to handle detokenization or downstream processing with zero server overhead.

Summary

  • Entry Point: Launch the server via python -m vllm.entrypoints.grpc_server with standard vLLM arguments.
  • Service Definition: The VllmEngine service is defined in vllm/grpc/vllm_engine.proto and implemented by VllmEngineServicer.
  • Core Methods: Use Generate for text completion (streaming or non-streaming), Embed for embeddings, and Abort to cancel requests.
  • Performance: The architecture uses AsyncLLM, grpc.aio, delta streaming, and unlimited message sizes to maximize throughput.
  • Client Access: Import generated modules from vllm.grpc to build clients in Python, or use the .proto file for other languages.

Frequently Asked Questions

How do I regenerate the protobuf Python files if I modify the proto schema?

Run vllm/grpc/compile_protos.py. This script invokes the protobuf compiler to regenerate vllm_engine_pb2.py and vllm_engine_pb2_grpc.py from vllm_engine.proto, ensuring the server and clients stay synchronized with your schema changes.

Can I use the gRPC API with tensor parallelism or pipeline parallelism?

Yes. The AsyncLLM engine initialized in grpc_server.py respects all standard vLLM arguments including --tensor-parallel-size and --pipeline-parallel-size. The gRPC servicer operates above the engine layer and is agnostic to the underlying distributed configuration.

What is the difference between the HTTP OpenAI-compatible server and the gRPC server?

The HTTP server (vllm.entrypoints.openai.api_server) provides OpenAI-compatible REST endpoints for easy integration with existing tools. The gRPC server (vllm.entrypoints.grpc_server) offers binary protobuf encoding, bidirectional streaming, and unlimited message sizes, making it superior for high-throughput, low-latency internal microservices.

How does the server handle very long contexts or large batches?

The gRPC server sets maximum message length parameters to -1 (unlimited) and relies on the AsyncLLM engine's continuous batching and PagedAttention memory management. This combination allows processing of extremely long sequences without the size limitations typically imposed by HTTP JSON payloads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →