# How to Use the gRPC API for High-Performance Serving in vLLM

> Learn to use the vLLM gRPC API for high-performance serving. Stream tokens efficiently with unlimited message sizes and structured output support for powerful language model inference.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**vLLM provides a native gRPC server via `VllmEngineServicer` that exposes the `AsyncLLM` engine through six RPC methods, enabling binary-efficient, token-streaming inference with unlimited message sizes and full structured output support.**

The vLLM project ships with a production-ready gRPC entrypoint that bypasses HTTP overhead for high-throughput workloads. This implementation leverages protobuf-defined contracts and asynchronous I/O to serve thousands of concurrent generation streams with minimal latency.

## Architecture of the vLLM gRPC Server

The gRPC stack in vLLM consists of tightly integrated layers that bridge the protobuf interface with the core inference engine.

### Core Components

**`vllm/grpc/vllm_engine.proto`** defines the service contract. This file specifies the `VllmEngine` service, request/response messages (like `GenerateRequest` and `GenerateResponse`), and the `SamplingParams` structure. It provides a language-agnostic binary protocol that clients can use from Python, Go, Rust, or any gRPC-compatible runtime.

**[`vllm/entrypoints/grpc_server.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/grpc_server.py)** serves as the CLI entry point. This module parses arguments, instantiates the `AsyncLLM` engine, and constructs a `grpc.aio` server backed by `uvloop`. It registers the `VllmEngineServicer` implementation and configures unlimited message sizes via `grpc.max_send_message_length` and `grpc.max_receive_message_length` set to `-1`.

**`VllmEngineServicer`** (defined within [`grpc_server.py`](https://github.com/vllm-project/vllm/blob/main/grpc_server.py)) implements six RPC methods:
- `Generate` – Streams token chunks or returns complete responses
- `Embed` – Handles embedding requests
- `HealthCheck` – Reports engine readiness
- `Abort` – Cancels in-flight request IDs
- `GetModelInfo` – Exposes model path and context length limits
- `GetServerInfo` – Returns active request counts and uptime

**`AsyncLLM`** (from [`vllm/v1/engine/async_llm.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/async_llm.py)) powers the backend. The servicer forwards decoded prompts to `async_llm.generate()`, which yields `RequestOutput` objects that the servicer converts into protobuf responses.

**[`vllm/grpc/compile_protos.py`](https://github.com/vllm-project/vllm/blob/main/vllm/grpc/compile_protos.py)** regenerates Python bindings. This utility compiles `vllm_engine.proto` into [`vllm_engine_pb2.py`](https://github.com/vllm-project/vllm/blob/main/vllm_engine_pb2.py) and [`vllm_engine_pb2_grpc.py`](https://github.com/vllm-project/vllm/blob/main/vllm_engine_pb2_grpc.py), which both the server and clients import.

### Request-Response Flow for Generation

When a client calls the `Generate` RPC, the system executes the following pipeline:

1. **Request Decoding**: The servicer receives a `GenerateRequest` containing either a `TokenizedInput` or raw text, plus `SamplingParams` (temperature, top-p, stop strings).

2. **Engine Invocation**: The servicer calls `self.async_llm.generate(prompt, sampling_params, ...)` and receives an asynchronous generator.

3. **Streaming Emission**: For each `RequestOutput` from the engine:
   - If streaming is enabled, `_chunk_response` emits a `GenerateResponse.chunk` containing only newly generated token IDs (`completion.token_ids`) and statistics (`prompt_tokens`, `completion_tokens`, `cached_tokens`).
   - When `output.finished` is `True`, `_complete_response` emits a final `GenerateResponse.complete` message.

4. **Client Consumption**: The client iterates over the `Generate` response stream, concatenating token IDs or applying detokenization logic to reconstruct text.

## Starting the gRPC Server

Launch the server using the module entry point. The following command starts the servicer on all interfaces with the specified model:

```bash
python -m vllm.entrypoints.grpc_server \
    --model meta-llama/Llama-2-7b-chat-hf \
    --host 0.0.0.0 \
    --port 50051

```

The server initializes the `AsyncLLM` engine, loads the model weights, and begins listening for gRPC connections on port 50051.

## Client Implementation Examples

The generated protobuf modules in `vllm/grpc/` provide the necessary stubs for client development.

### Non-Streaming Generation

For synchronous-style completion where the server returns a single final response:

```python
import asyncio
import grpc
from vllm.grpc import vllm_engine_pb2, vllm_engine_pb2_grpc

async def generate_once():
    async with grpc.aio.insecure_channel("localhost:50051") as channel:
        stub = vllm_engine_pb2_grpc.VllmEngineStub(channel)

        request = vllm_engine_pb2.GenerateRequest(
            request_id="demo-1",
            text="The capital of Italy is",
            sampling_params=vllm_engine_pb2.SamplingParams(
                temperature=0.0,
                max_tokens=5,
                n=1,
            ),
            stream=False,
        )

        async for resp in stub.Generate(request):
            if resp.HasField("complete"):
                output_ids = resp.complete.output_ids
                print("Token IDs:", output_ids)

asyncio.run(generate_once())

```

### Streaming Token Generation

For real-time applications requiring per-token latency, enable streaming to receive delta updates:

```python
import asyncio
import grpc
from vllm.grpc import vllm_engine_pb2, vllm_engine_pb2_grpc

async def generate_stream():
    async with grpc.aio.insecure_channel("localhost:50051") as channel:
        stub = vllm_engine_pb2_grpc.VllmEngineStub(channel)

        request = vllm_engine_pb2.GenerateRequest(
            request_id="stream-1",
            text="The quick brown fox",
            sampling_params=vllm_engine_pb2.SamplingParams(
                temperature=0.7,
                max_tokens=12,
                n=1,
            ),
            stream=True,
        )

        async for resp in stub.Generate(request):
            if resp.HasField("chunk"):
                chunk = resp.chunk
                print(f"New tokens: {list(chunk.token_ids)} "
                      f"(prompt={chunk.prompt_tokens}, "
                      f"completion={chunk.completion_tokens})")
            elif resp.HasField("complete"):
                print("Generation finished")
                break

asyncio.run(generate_stream())

```

### Health and Metadata RPCs

Use utility RPCs for orchestration and load balancing:

```python
import asyncio
import grpc
from vllm.grpc import vllm_engine_pb2, vllm_engine_pb2_grpc

async def meta():
    async with grpc.aio.insecure_channel("localhost:50051") as ch:
        stub = vllm_engine_pb2_grpc.VllmEngineStub(ch)

        health = await stub.HealthCheck(vllm_engine_pb2.HealthCheckRequest())
        print("Healthy:", health.healthy, health.message)

        model_info = await stub.GetModelInfo(vllm_engine_pb2.GetModelInfoRequest())
        print("Model path:", model_info.model_path,
              "max ctx:", model_info.max_context_length)

        server_info = await stub.GetServerInfo(vllm_engine_pb2.GetServerInfoRequest())
        print("Active requests:", server_info.active_requests,
              "uptime:", server_info.uptime_seconds)

asyncio.run(meta())

```

### Aborting In-Flight Requests

Cancel specific request IDs to free up engine capacity:

```python
abort_req = vllm_engine_pb2.AbortRequest(request_ids=["stream-1"])
await stub.Abort(abort_req)

```

## Key Design Features for Performance

The vLLM gRPC implementation optimizes for high-throughput serving through several architectural decisions.

**AsyncIO-First Execution**: All RPC methods in `VllmEngineServicer` use `grpc.aio` and async/await syntax. This allows the server to manage thousands of concurrent generation streams without blocking threads, maximizing GPU utilization.

**Delta-Mode Streaming**: When `stream=True`, the server transmits only new token IDs (delta) rather than the full generated sequence. This reduces bandwidth consumption proportionally to sequence length, making it ideal for long-form generation.

**Unlimited Message Sizes**: The server explicitly sets `grpc.max_send_message_length` and `grpc.max_receive_message_length` to `-1`, removing protobuf size limits. This accommodates very long contexts and large batch requests without truncation.

**Structured Output Constraints**: The protobuf schema defines a `oneof constraint` field supporting JSON schema, regex, and grammar constraints. The servicer maps these to `StructuredOutputsParams` via `_sampling_params_from_proto`, enabling type-safe generation without client-side validation.

**Detokenization Control**: `SamplingParams.detokenize` is forced to `True` only when stop strings are present. Otherwise, the server returns raw token IDs, allowing clients to handle detokenization or downstream processing with zero server overhead.

## Summary

- **Entry Point**: Launch the server via `python -m vllm.entrypoints.grpc_server` with standard vLLM arguments.
- **Service Definition**: The `VllmEngine` service is defined in `vllm/grpc/vllm_engine.proto` and implemented by `VllmEngineServicer`.
- **Core Methods**: Use `Generate` for text completion (streaming or non-streaming), `Embed` for embeddings, and `Abort` to cancel requests.
- **Performance**: The architecture uses `AsyncLLM`, `grpc.aio`, delta streaming, and unlimited message sizes to maximize throughput.
- **Client Access**: Import generated modules from `vllm.grpc` to build clients in Python, or use the `.proto` file for other languages.

## Frequently Asked Questions

### How do I regenerate the protobuf Python files if I modify the proto schema?

Run [`vllm/grpc/compile_protos.py`](https://github.com/vllm-project/vllm/blob/main/vllm/grpc/compile_protos.py). This script invokes the protobuf compiler to regenerate [`vllm_engine_pb2.py`](https://github.com/vllm-project/vllm/blob/main/vllm_engine_pb2.py) and [`vllm_engine_pb2_grpc.py`](https://github.com/vllm-project/vllm/blob/main/vllm_engine_pb2_grpc.py) from `vllm_engine.proto`, ensuring the server and clients stay synchronized with your schema changes.

### Can I use the gRPC API with tensor parallelism or pipeline parallelism?

Yes. The `AsyncLLM` engine initialized in [`grpc_server.py`](https://github.com/vllm-project/vllm/blob/main/grpc_server.py) respects all standard vLLM arguments including `--tensor-parallel-size` and `--pipeline-parallel-size`. The gRPC servicer operates above the engine layer and is agnostic to the underlying distributed configuration.

### What is the difference between the HTTP OpenAI-compatible server and the gRPC server?

The HTTP server (`vllm.entrypoints.openai.api_server`) provides OpenAI-compatible REST endpoints for easy integration with existing tools. The gRPC server (`vllm.entrypoints.grpc_server`) offers binary protobuf encoding, bidirectional streaming, and unlimited message sizes, making it superior for high-throughput, low-latency internal microservices.

### How does the server handle very long contexts or large batches?

The gRPC server sets maximum message length parameters to `-1` (unlimited) and relies on the `AsyncLLM` engine's continuous batching and PagedAttention memory management. This combination allows processing of extremely long sequences without the size limitations typically imposed by HTTP JSON payloads.