# Implementing Streaming Responses with GLM-5 in vLLM: A Complete Guide

> Implement GLM-5 streaming responses in vLLM with this comprehensive guide. Discover how setting "stream": true enables server-sent events for token generation.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-19

---

**Enable streaming by setting `"stream": true` in your vLLM request payload when using the GLM-5-S checkpoint, which emits server-sent events (SSE) containing individual tokens as they are generated.**

The **zai-org/GLM-5** repository provides a specialized streaming variant called **GLM-5-S** that delivers token-wise output rather than waiting for complete generation. When deployed with **vLLM**, a high-throughput inference engine, you can implement real-time streaming responses that improve user experience for interactive applications.

## What is GLM-5-S and Why Use Streaming?

**GLM-5-S** is the streaming-optimized variant of the GLM-5 family, identified by the `-S` suffix in model identifiers (e.g., `zai-org/GLM-5.2-S`). Unlike standard batch generation, this variant leverages vLLM's **Prefill-Decode** pipeline to emit tokens immediately as they are generated.

Streaming responses provide three critical advantages for production deployments:

- **Reduced perceived latency**: Users see the first token within milliseconds rather than waiting for the full completion
- **Interactive experiences**: Chatbots and IDE assistants feel responsive and alive
- **Efficient resource utilization**: The **Sparse Mixture-of-Experts (MoE)** architecture in GLM-5 reduces per-token computation through the **MoE Mega-Fusion** operator, which eliminates redundant reads and writes of intermediate tensors during streaming generation

## Prerequisites and Model Configuration

Before implementing streaming, ensure your environment meets the following requirements specified in the repository's [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md):

- **vLLM version**: Install vLLM ≥ 0.23.0 as recommended in the [README](https://github.com/zai-org/GLM-5/blob/main/README.md#L75) for GLM-5 series compatibility
- **Model identifier**: Use the checkpoint ending with `-S` from the [model download table](https://github.com/zai-org/GLM-5/blob/main/README.md#L61-L66) (e.g., `zai-org/GLM-5.2-S`)
- **Streaming flag**: Include `"stream": true` in your JSON payload to activate the `StreamingEngine` in vLLM

The repository also recommends keeping the `reasoning_effort` parameter at its default value of `max` unless you specifically require higher-quality generation with different latency characteristics.

## How Streaming Works in vLLM with GLM-5

When vLLM receives a streaming request for GLM-5-S, it executes a sophisticated pipeline that maintains low latency while handling the model's sparse architecture:

1. **Request handling**: The vLLM HTTP server creates a `StreamingEngine` that forwards the request to the underlying scheduler
2. **Prefill phase**: The **Prefill-Decode** pipeline processes the prompt once, caching attention states using **IndexCache** and prefix caching mechanisms
3. **Token generation**: The model generates tokens one-by-one, with the MoE Mega-Fusion operator optimizing expert routing for each individual token
4. **Response flushing**: The HTTP handler writes each token to the response body immediately, using either newline-delimited JSON for SSE or HTTP-chunked transfer encoding

This architecture ensures that long-context generation remains smooth through aggressive prefix caching, as documented in the repository's performance notes.

## Implementation Steps

### Launching the vLLM Server

Start the vLLM OpenAI-compatible API server with the GLM-5-S checkpoint:

```bash
python -m vllm.entrypoints.openai.api_server \
    --model zai-org/GLM-5.2-S \
    --port 8000

```

For **Ascend NPU** deployments, refer to the hardware-specific instructions in [[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) to optimize streaming performance on specialized accelerators.

### Sending Streaming Requests

When constructing your request payload, include the `stream` parameter to enable SSE output:

```json
{
  "model": "zai-org/GLM-5.2-S",
  "messages": [{"role": "user", "content": "Your prompt here"}],
  "stream": true,
  "max_tokens": 256,
  "temperature": 0.7
}

```

## Code Examples

### Python Client with Server-Sent Events

This implementation consumes the SSE stream and displays tokens in real-time:

```python
import requests

def stream_glm_s(prompt):
    url = "http://localhost:8000/v1/chat/completions"
    payload = {
        "model": "zai-org/GLM-5.2-S",
        "messages": [{"role": "user", "content": prompt}],
        "stream": True,
        "max_tokens": 256,
        "temperature": 0.7,
    }
    
    with requests.post(url, json=payload, stream=True) as resp:
        for line in resp.iter_lines():
            if line:
                data = line.decode("utf-8")
                if data.startswith("data:"):
                    token_json = data[5:].strip()
                    token_data = eval(token_json)
                    token = token_data['choices'][0]['delta'].get('content')
                    if token:
                        print(token, end='', flush=True)
    print()

# Example usage

stream_glm_s("Explain the concept of sparse attention in GLM-5.")

```

The client uses `stream=True` in the `requests` library to prevent buffering, ensuring each token appears immediately as vLLM generates it.

### curl Command for Testing

Test your streaming endpoint using curl with the `-N` flag to disable buffering:

```bash
curl -N -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "zai-org/GLM-5.2-S",
        "messages": [{"role": "user", "content": "Write a short poem about AI."}],
        "stream": true,
        "max_tokens": 128
      }'

```

Each `data:` line contains a JSON fragment with the next token in the `choices[0].delta.content` field.

## Optimization Techniques for Low-Latency Streaming

To maximize streaming performance with GLM-5 in vLLM, leverage these architectural optimizations:

- **MoE Mega-Fusion**: The sparse architecture uses fused operators to minimize memory bandwidth during token generation, critical for maintaining throughput in streaming mode
- **Prefix Caching**: vLLM's `IndexCache` stores precomputed attention keys and values, eliminating redundant computation for long contexts
- **Connection Management**: For production services, implement back-pressure handling on the client side; vLLM automatically respects the client's `Accept-Encoding` header and pauses generation if the network stalls

For additional deployment strategies and community troubleshooting resources, consult [[`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md)](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md) and [[`resources/WECHAT.md`](https://github.com/zai-org/GLM-5/blob/main/resources/WECHAT.md)](https://github.com/zai-org/GLM-5/blob/main/resources/WECHAT.md).

## Summary

- **GLM-5-S** is the streaming variant of GLM-5, enabled by using model identifiers ending in `-S` and setting `"stream": true` in vLLM requests
- **vLLM ≥ 0.23.0** provides the `StreamingEngine` and `Prefill-Decode` pipeline required for efficient token-by-token generation
- **Server-sent events (SSE)** deliver JSON-formatted tokens immediately upon generation, creating responsive user experiences
- **MoE Mega-Fusion** and **IndexCache** optimizations reduce latency and memory overhead during streaming inference
- **Ascend NPU** support is available through specialized deployment guides in the repository

## Frequently Asked Questions

### What is the difference between GLM-5 and GLM-5-S?

**GLM-5-S** is specifically optimized for streaming inference, using the same base weights as GLM-5 but configured to work with vLLM's `StreamingEngine` to emit tokens individually rather than waiting for complete generation. According to the repository's [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md), you identify the streaming variant by the `-S` suffix in the model name (e.g., `GLM-5.2-S`).

### Why does my streaming response buffer instead of displaying tokens immediately?

Buffering typically occurs on the **client side** rather than in vLLM. Ensure your HTTP client disables buffering—for example, use `stream=True` in Python's `requests` library or the `-N` flag with `curl`. The vLLM server emits tokens immediately as SSE messages, but client-side buffering can delay display until the connection closes.

### How does vLLM handle the GLM-5 Sparse Mixture-of-Experts architecture during streaming?

vLLM optimizes GLM-5's **MoE** layers through the **Mega-Fusion** operator, which fuses expert routing and computation to eliminate redundant memory reads during token generation. This is particularly important for streaming, where each token must be processed and emitted with minimal latency. The sparse architecture actually benefits streaming by reducing per-token computational cost compared to dense models.

### Can I use GLM-5 streaming with Ascend NPUs?

Yes, the repository includes specific deployment instructions for **Ascend NPUs** in [[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md). These hardware-accelerated deployments can further reduce streaming latency through optimized kernel implementations for the GLM-5 architecture, though you should verify compatibility with your specific vLLM version and NPU driver stack.