How to Implement PD Separation for Throughput Stability in GLM-5.2
Implement PD separation in GLM-5.2 by running prefill and decode phases on independent execution streams with prefix caching enabled, eliminating resource contention that causes latency jitter in high-concurrency inference deployments.
GLM-5.2 (the inference variant of the GLM-5 series, also referred to as GLM-S.2) follows a classic two-stage generation pipeline where prefill and decode phases compete for compute resources. When executed on the same queue, long prefills block short decodes, causing throughput instability. The zai-org/GLM-5 repository provides specific optimizations to separate these phases, documented in the Ascend deployment guide and model configuration files.
Understanding PD Separation in GLM-5.2
The inference process consists of two distinct stages:
- Prefill stage – The model consumes the prompt and builds the KV-cache for already-seen tokens.
- Decode stage – The model autoregressively generates new tokens, reusing cached KV-states from the prefill phase.
When these stages share execution streams, decode operations stall while waiting for prefill computations to complete. PD separation removes this contention by running prefill and decode on independent threads or devices, ensuring decode progresses at a steady rate regardless of prefill workload.
Core Mechanisms for Throughput Stability
According to the Ascend deployment guide in example/ascend.md, GLM-5.2 implements PD separation through four optimization mechanisms:
Prefill-Decode Disaggregation
This mechanism splits the two phases into separate kernels or threads. By eliminating resource-stealing that occurs when a long prefill blocks a short decode, the system guarantees that decode operations progress at a steady rate, smoothing the latency curve.
Prefix Caching
After prefill completes, the KV-cache for the prompt prefix is retained in fast-access memory (GPU/CPU or dedicated index cache). Subsequent decode calls reuse this cached prefix without recomputation, reducing per-token costs and suppressing jitter caused by redundant computation.
Prefill Delay Scheduling
In high-concurrency mixed-load scenarios, the system adds a tiny artificial delay (typically 2ms) to the prefill pipeline. This spreads prefill workload evenly over time, preventing sudden spikes that would otherwise stall the decode stream.
IndexCache and Expert-Path Caching
For MoE-based models, routing decisions for the prefix are cached in IndexCache. This avoids re-routing the same set of experts on every generation, lowering compute overhead and allowing the decode stream to remain fully occupied.
Implementation by Inference Engine
The repository supports PD separation across multiple inference backends. Configure your engine using the following patterns.
vLLM (Python)
In vLLM, enable PD separation by setting prefill_decode_separate and enable_prefix_caching when initializing the engine:
from vllm import LLM, SamplingParams
# Enable Prefill-Decode separation + prefix caching
engine = LLM(
model="zai-org/GLM-5.2",
tensor_parallel_size=2,
enable_prefix_caching=True, # keep KV-cache for prompt prefixes
prefill_decode_separate=True, # run prefill and decode on independent streams
prefill_delay=2, # optional delay (ms) for high-concurrency
)
sampling_params = SamplingParams(max_tokens=128, temperature=0.7)
prompt = "Explain the benefits of PD separation in large language models."
# The first call performs prefill; subsequent calls reuse the prefix cache
outputs = engine.generate([prompt], sampling_params)
print(outputs[0].text)
The key flags enable_prefix_caching, prefill_decode_separate, and prefill_delay map directly to the optimization mechanisms described in the Ascend guide.
SGLang (YAML Configuration)
For SGLang deployments, specify PD separation in your configuration file:
model:
name: "zai-org/GLM-5.2"
# Activate PD separation
prefill_decode_separate: true
# Keep prompt KV-cache in fast memory
enable_prefix_caching: true
# Optional prefill delay (ms) for mixed-load workloads
prefill_delay_ms: 2
# MoE index cache (size in MB)
index_cache_size: 256
Load this configuration when starting the engine:
sglang --config pd_separated.yaml serve
The engine spins up two worker threads—one for prefill and one for decode—and automatically caches prefix KV-states.
xLLM (C++ API)
For xLLM integrations, use the builder pattern to enable separation:
auto engine = xllm::Engine::Builder()
.model("zai-org/GLM-5.2")
.prefill_decode_separate(true)
.enable_prefix_caching(true)
.prefill_delay_ms(2)
.build();
std::string prompt = "What are the design trade-offs of PD separation?";
auto result = engine->generate(prompt, max_tokens=128);
std::cout << result.text << std::endl;
All three backends expose the same core knobs: separation flags, prefix caching, and optional delay scheduling.
Key Source Files and Configuration
The following files in the zai-org/GLM-5 repository contain the implementation details and configuration schemas for PD separation:
example/ascend.md– Documents the six optimization techniques, including PD separation and prefix caching strategies for Ascend hardware.README.md(serve section) – Lists supported inference frameworks and points to the configuration flags required for PD separation.skills/glm-master-skill/SKILL.md– References PD separation as a capability in the skill registry for downstream agent integration.
Summary
- PD separation eliminates throughput instability by isolating prefill and decode execution streams.
- Prefix caching stores prompt KV-cache in fast memory, preventing recomputation during decode phases.
- Prefill delay scheduling (2ms) smooths high-concurrency workloads to prevent decode stalls.
- Expert-path caching optimizes MoE routing decisions for the prefill stage.
- Configure separation using
prefill_decode_separateandenable_prefix_cachingflags in vLLM, SGLang, or xLLM.
Frequently Asked Questions
What causes throughput instability in GLM-5.2 without PD separation?
Without separation, prefill and decode operations compete for the same compute resources. Long prefills block short decodes, creating latency jitter and uneven request processing times that degrade service level objectives in high-concurrency deployments.
How does prefix caching improve decode performance?
Prefix caching retains the KV-cache computed during the initial prefill in fast-access memory. Subsequent decode calls reuse these cached states instead of recomputing them, reducing per-token latency and eliminating jitter caused by redundant computation of the same prompt prefixes.
What configuration flags enable PD separation in GLM-5.2?
Set prefill_decode_separate=true to enable stream separation, enable_prefix_caching=true to retain prefix KV-cache, and optionally prefill_delay_ms=2 to add artificial delay for load smoothing. For MoE models, configure index_cache_size to cache expert routing decisions.
Can PD separation be used with mixture-of-experts (MoE) architectures?
Yes. GLM-5.2 implements IndexCache specifically for MoE models to cache expert routing decisions made during the prefill phase. This prevents re-routing overhead during decode, ensuring the separated streams maintain high throughput even with complex expert architectures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →