Implementing Streaming Responses with GLM-5 in vLLM: A Complete Guide

Enable streaming by setting "stream": true in your vLLM request payload when using the GLM-5-S checkpoint, which emits server-sent events (SSE) containing individual tokens as they are generated.

The zai-org/GLM-5 repository provides a specialized streaming variant called GLM-5-S that delivers token-wise output rather than waiting for complete generation. When deployed with vLLM, a high-throughput inference engine, you can implement real-time streaming responses that improve user experience for interactive applications.

What is GLM-5-S and Why Use Streaming?

GLM-5-S is the streaming-optimized variant of the GLM-5 family, identified by the -S suffix in model identifiers (e.g., zai-org/GLM-5.2-S). Unlike standard batch generation, this variant leverages vLLM's Prefill-Decode pipeline to emit tokens immediately as they are generated.

Streaming responses provide three critical advantages for production deployments:

  • Reduced perceived latency: Users see the first token within milliseconds rather than waiting for the full completion
  • Interactive experiences: Chatbots and IDE assistants feel responsive and alive
  • Efficient resource utilization: The Sparse Mixture-of-Experts (MoE) architecture in GLM-5 reduces per-token computation through the MoE Mega-Fusion operator, which eliminates redundant reads and writes of intermediate tensors during streaming generation

Prerequisites and Model Configuration

Before implementing streaming, ensure your environment meets the following requirements specified in the repository's README.md:

  • vLLM version: Install vLLM ≥ 0.23.0 as recommended in the README for GLM-5 series compatibility
  • Model identifier: Use the checkpoint ending with -S from the model download table (e.g., zai-org/GLM-5.2-S)
  • Streaming flag: Include "stream": true in your JSON payload to activate the StreamingEngine in vLLM

The repository also recommends keeping the reasoning_effort parameter at its default value of max unless you specifically require higher-quality generation with different latency characteristics.

How Streaming Works in vLLM with GLM-5

When vLLM receives a streaming request for GLM-5-S, it executes a sophisticated pipeline that maintains low latency while handling the model's sparse architecture:

  1. Request handling: The vLLM HTTP server creates a StreamingEngine that forwards the request to the underlying scheduler
  2. Prefill phase: The Prefill-Decode pipeline processes the prompt once, caching attention states using IndexCache and prefix caching mechanisms
  3. Token generation: The model generates tokens one-by-one, with the MoE Mega-Fusion operator optimizing expert routing for each individual token
  4. Response flushing: The HTTP handler writes each token to the response body immediately, using either newline-delimited JSON for SSE or HTTP-chunked transfer encoding

This architecture ensures that long-context generation remains smooth through aggressive prefix caching, as documented in the repository's performance notes.

Implementation Steps

Launching the vLLM Server

Start the vLLM OpenAI-compatible API server with the GLM-5-S checkpoint:

python -m vllm.entrypoints.openai.api_server \
    --model zai-org/GLM-5.2-S \
    --port 8000

For Ascend NPU deployments, refer to the hardware-specific instructions in [example/ascend.md](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) to optimize streaming performance on specialized accelerators.

Sending Streaming Requests

When constructing your request payload, include the stream parameter to enable SSE output:

{
  "model": "zai-org/GLM-5.2-S",
  "messages": [{"role": "user", "content": "Your prompt here"}],
  "stream": true,
  "max_tokens": 256,
  "temperature": 0.7
}

Code Examples

Python Client with Server-Sent Events

This implementation consumes the SSE stream and displays tokens in real-time:

import requests

def stream_glm_s(prompt):
    url = "http://localhost:8000/v1/chat/completions"
    payload = {
        "model": "zai-org/GLM-5.2-S",
        "messages": [{"role": "user", "content": prompt}],
        "stream": True,
        "max_tokens": 256,
        "temperature": 0.7,
    }
    
    with requests.post(url, json=payload, stream=True) as resp:
        for line in resp.iter_lines():
            if line:
                data = line.decode("utf-8")
                if data.startswith("data:"):
                    token_json = data[5:].strip()
                    token_data = eval(token_json)
                    token = token_data['choices'][0]['delta'].get('content')
                    if token:
                        print(token, end='', flush=True)
    print()

# Example usage

stream_glm_s("Explain the concept of sparse attention in GLM-5.")

The client uses stream=True in the requests library to prevent buffering, ensuring each token appears immediately as vLLM generates it.

curl Command for Testing

Test your streaming endpoint using curl with the -N flag to disable buffering:

curl -N -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "zai-org/GLM-5.2-S",
        "messages": [{"role": "user", "content": "Write a short poem about AI."}],
        "stream": true,
        "max_tokens": 128
      }'

Each data: line contains a JSON fragment with the next token in the choices[0].delta.content field.

Optimization Techniques for Low-Latency Streaming

To maximize streaming performance with GLM-5 in vLLM, leverage these architectural optimizations:

  • MoE Mega-Fusion: The sparse architecture uses fused operators to minimize memory bandwidth during token generation, critical for maintaining throughput in streaming mode
  • Prefix Caching: vLLM's IndexCache stores precomputed attention keys and values, eliminating redundant computation for long contexts
  • Connection Management: For production services, implement back-pressure handling on the client side; vLLM automatically respects the client's Accept-Encoding header and pauses generation if the network stalls

For additional deployment strategies and community troubleshooting resources, consult [skills/glm-master-skill/SKILL.md](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md) and [resources/WECHAT.md](https://github.com/zai-org/GLM-5/blob/main/resources/WECHAT.md).

Summary

  • GLM-5-S is the streaming variant of GLM-5, enabled by using model identifiers ending in -S and setting "stream": true in vLLM requests
  • vLLM ≥ 0.23.0 provides the StreamingEngine and Prefill-Decode pipeline required for efficient token-by-token generation
  • Server-sent events (SSE) deliver JSON-formatted tokens immediately upon generation, creating responsive user experiences
  • MoE Mega-Fusion and IndexCache optimizations reduce latency and memory overhead during streaming inference
  • Ascend NPU support is available through specialized deployment guides in the repository

Frequently Asked Questions

What is the difference between GLM-5 and GLM-5-S?

GLM-5-S is specifically optimized for streaming inference, using the same base weights as GLM-5 but configured to work with vLLM's StreamingEngine to emit tokens individually rather than waiting for complete generation. According to the repository's README.md, you identify the streaming variant by the -S suffix in the model name (e.g., GLM-5.2-S).

Why does my streaming response buffer instead of displaying tokens immediately?

Buffering typically occurs on the client side rather than in vLLM. Ensure your HTTP client disables buffering—for example, use stream=True in Python's requests library or the -N flag with curl. The vLLM server emits tokens immediately as SSE messages, but client-side buffering can delay display until the connection closes.

How does vLLM handle the GLM-5 Sparse Mixture-of-Experts architecture during streaming?

vLLM optimizes GLM-5's MoE layers through the Mega-Fusion operator, which fuses expert routing and computation to eliminate redundant memory reads during token generation. This is particularly important for streaming, where each token must be processed and emitted with minimal latency. The sparse architecture actually benefits streaming by reducing per-token computational cost compared to dense models.

Can I use GLM-5 streaming with Ascend NPUs?

Yes, the repository includes specific deployment instructions for Ascend NPUs in [example/ascend.md](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md). These hardware-accelerated deployments can further reduce streaming latency through optimized kernel implementations for the GLM-5 architecture, though you should verify compatibility with your specific vLLM version and NPU driver stack.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →