How to Handle Prefill-Decode Disaggregation for Mixed Loads in GLM-5.2

Prefill-decode disaggregation separates the prefill and decode stages into distinct execution streams, while prefill delay scheduling and chunked prefill prevent long-prompt requests from blocking short-prompt decode operations in high-concurrency GLM-5.2 deployments.

GLM-5.2 (also referred to as GLM-S.2) processes inference requests through two distinct computational stages that become bottlenecks under mixed loads—simultaneous requests with varying prompt lengths. When deployed at high concurrency, the default synchronous pipeline causes head-of-line blocking that degrades throughput and increases latency, requiring specific architectural optimizations documented in the zai-org/GLM-5 repository.

Understanding the Mixed-Load Bottleneck

The inference pipeline divides each request into discrete phases:

  • Prefill: The initial forward pass that computes key-value (KV) caches for new input tokens
  • Decode: Autoregressive generation steps that repeatedly attend to cached KV states

According to example/ascend.md, long-prompt prefill operations can monopolize GPU or NPU resources for extended periods. This occupation forces shorter decode-bound requests to wait in queue, causing latency spikes and underutilization of compute resources in production environments.

Core Disaggregation Techniques

PD Separation Strategy

The PD Separation technique splits the processing graph so prefill and decode operations run on separate execution streams. As documented in the "PD Separation and Prefix Caching" section of example/ascend.md, this isolation prevents long-running prefill jobs from interfering with the fast decode pipeline, enabling parallel processing of different request types without resource contention.

Prefix Caching

Prefix Caching stores the KV states of common prompt prefixes—such as system prompts or repeated instruction templates—in memory. Future requests sharing these prefixes can reuse cached states instead of recomputing them during prefill. This optimization significantly reduces computational overhead for frequently occurring prompt patterns.

IndexCache and Sparse Index Retrieval

For sparse Mixture-of-Experts (MoE) variants, IndexCache stores high-frequency expert routing paths. This technique accelerates both prefill and decode stages by eliminating redundant routing computations, as detailed in the "Intelligent Caching and Index Optimization" section of example/ascend.md.

Prefill Delay Scheduling

When the system detects request surges, Prefill Delay Scheduling deliberately postpones new prefill jobs for a configurable interval (typically 4-5 milliseconds). As implemented in the "High-Concurrency Scheduling with Prefill Delay" section of example/ascend.md, this smoothing technique allows decode workers to process queued tokens while the prefill queue drains, preventing congestion collapse.

Chunked Prefill

Chunked Prefill breaks large prefill computations into smaller sub-units that interleave with decode steps. This approach prevents any single prefill operation from blocking the entire pipeline, reducing worst-case stall times for short-prompt requests waiting behind long-prompt prefill jobs.

Configuration Examples for Production Deployment

vLLM Configuration

The vLLM framework supports GLM-5.2 disaggregation through specific initialization parameters:

from vllm import LLM, SamplingParams

# Enable Prefill-Decode disaggregation and set a prefill delay (ms)

engine = LLM(
    model="zai-org/GLM-5.2",
    tensor_parallel_size=2,
    enable_prefill_decode_disaggregation=True,   # <‑‑ separates PD stages

    prefill_delay_ms=5,                         # <‑‑ short delay to smooth spikes

    max_num_batched_tokens=512,                 # chunk size for large prefills

)

sampling_params = SamplingParams(temperature=0.7, max_new_tokens=256)

# Simple request

prompt = "Explain the impact of quantum computing on cryptography."
outputs = engine.generate(prompt, sampling_params)
print(outputs[0].text)

SGLang Configuration

For Ascend NPU deployments using SGLang, activate disaggregation through the engine configuration:

from sglang import SGLangEngine, GenerationConfig

engine = SGLangEngine(
    model="zai-org/GLM-5.2",
    device="ascend",
    pd_separate=True,               # <‑‑ activates PD separation

    prefill_delay=4,                # <‑‑ milliseconds

    prefix_caching=True,            # <‑‑ enable prefix cache

    chunked_prefill=True,           # <‑‑ split large prefills

)

gen_cfg = GenerationConfig(temperature=0.8, max_new_tokens=200)

prompt = "Write a short story about a traveler in a desert."
result = engine.generate(prompt, gen_cfg)
print(result.text)

Essential Documentation References

Three critical files in the zai-org/GLM-5 repository provide implementation details:

  • example/ascend.md: Documents the complete inference optimization suite including PD separation, prefill delay, prefix caching, and chunked prefill for Ascend NPU hardware
  • README.md: Lists framework support for prefill-decode disaggregation across vLLM, SGLang, and xLLM serving engines
  • skills/glm-master-skill/SKILL.md: Indexes official GLM skills and links to detailed deployment documentation

Summary

  • Prefill-decode disaggregation isolates resource-intensive prefill operations from latency-sensitive decode steps by running them on separate execution streams
  • Prefix caching eliminates redundant computation for shared prompt prefixes, reducing prefill time for common request patterns
  • Prefill delay scheduling smooths traffic spikes by temporarily deferring new prefill jobs for 4-5 milliseconds, allowing decode queues to drain
  • Chunked prefill interleaves large prefill computations with decode steps to prevent pipeline stalls and head-of-line blocking
  • IndexCache optimizes expert routing for sparse MoE models during both prefill and decode phases
  • Both vLLM and SGLang support these optimizations through specific configuration flags for GLM-5.2 deployments

Frequently Asked Questions

What is Prefill-Decode disaggregation?

Prefill-Decode disaggregation is an architectural optimization that separates the two-stage inference pipeline of large language models. By running prefill (KV cache computation) and decode (token generation) on distinct execution streams, the system prevents long-prompt requests from blocking short-prompt generation, significantly improving throughput in high-concurrency scenarios.

How does prefill delay scheduling improve throughput?

Prefill delay scheduling introduces a configurable delay (typically 4-5 milliseconds) before starting new prefill jobs during request surges. According to example/ascend.md, this brief pause allows the decode queue to drain while smoothing the demand curve, preventing the system from accepting more prefill work than it can efficiently process.

Why is prefix caching important for mixed loads?

Prefix caching stores the KV cache states of frequently used prompt prefixes, such as system instructions. When handling mixed loads in GLM-5.2, this technique eliminates the need to recompute expensive prefill operations for common prefixes, reducing both latency and computational overhead for subsequent requests sharing identical starting sequences.

Which serving frameworks support GLM-5.2 disaggregation?

The zai-org/GLM-5 repository officially supports prefill-decode disaggregation in vLLM, SGLang, and xLLM. The README.md provides high-level deployment guidance, while example/ascend.md contains specific configuration examples for Ascend NPU hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →