# How to Implement PD Separation for Throughput Stability in GLM-5.2

> Stabilize GLM-5.2 throughput using PD separation. Run prefill and decode on separate streams with prefix caching to eliminate latency jitter in high-concurrency inference.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-21

---

**Implement PD separation in GLM-5.2 by running prefill and decode phases on independent execution streams with prefix caching enabled, eliminating resource contention that causes latency jitter in high-concurrency inference deployments.**

GLM-5.2 (the inference variant of the GLM-5 series, also referred to as GLM-S.2) follows a classic two-stage generation pipeline where prefill and decode phases compete for compute resources. When executed on the same queue, long prefills block short decodes, causing throughput instability. The `zai-org/GLM-5` repository provides specific optimizations to separate these phases, documented in the Ascend deployment guide and model configuration files.

## Understanding PD Separation in GLM-5.2

The inference process consists of two distinct stages:

1. **Prefill stage** – The model consumes the prompt and builds the KV-cache for already-seen tokens.
2. **Decode stage** – The model autoregressively generates new tokens, reusing cached KV-states from the prefill phase.

When these stages share execution streams, decode operations stall while waiting for prefill computations to complete. **PD separation** removes this contention by running prefill and decode on independent threads or devices, ensuring decode progresses at a steady rate regardless of prefill workload.

## Core Mechanisms for Throughput Stability

According to the Ascend deployment guide in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md), GLM-5.2 implements PD separation through four optimization mechanisms:

### Prefill-Decode Disaggregation

This mechanism splits the two phases into separate kernels or threads. By eliminating resource-stealing that occurs when a long prefill blocks a short decode, the system guarantees that decode operations progress at a steady rate, smoothing the latency curve.

### Prefix Caching

After prefill completes, the KV-cache for the prompt prefix is retained in fast-access memory (GPU/CPU or dedicated index cache). Subsequent decode calls reuse this cached prefix without recomputation, reducing per-token costs and suppressing jitter caused by redundant computation.

### Prefill Delay Scheduling

In high-concurrency mixed-load scenarios, the system adds a tiny artificial delay (typically 2ms) to the prefill pipeline. This spreads prefill workload evenly over time, preventing sudden spikes that would otherwise stall the decode stream.

### IndexCache and Expert-Path Caching

For MoE-based models, routing decisions for the prefix are cached in `IndexCache`. This avoids re-routing the same set of experts on every generation, lowering compute overhead and allowing the decode stream to remain fully occupied.

## Implementation by Inference Engine

The repository supports PD separation across multiple inference backends. Configure your engine using the following patterns.

### vLLM (Python)

In `vLLM`, enable PD separation by setting `prefill_decode_separate` and `enable_prefix_caching` when initializing the engine:

```python
from vllm import LLM, SamplingParams

# Enable Prefill-Decode separation + prefix caching

engine = LLM(
    model="zai-org/GLM-5.2",
    tensor_parallel_size=2,
    enable_prefix_caching=True,          # keep KV-cache for prompt prefixes

    prefill_decode_separate=True,        # run prefill and decode on independent streams

    prefill_delay=2,                     # optional delay (ms) for high-concurrency

)

sampling_params = SamplingParams(max_tokens=128, temperature=0.7)

prompt = "Explain the benefits of PD separation in large language models."

# The first call performs prefill; subsequent calls reuse the prefix cache

outputs = engine.generate([prompt], sampling_params)
print(outputs[0].text)

```

The key flags `enable_prefix_caching`, `prefill_decode_separate`, and `prefill_delay` map directly to the optimization mechanisms described in the Ascend guide.

### SGLang (YAML Configuration)

For `SGLang` deployments, specify PD separation in your configuration file:

```yaml
model:
  name: "zai-org/GLM-5.2"
  # Activate PD separation

  prefill_decode_separate: true
  # Keep prompt KV-cache in fast memory

  enable_prefix_caching: true
  # Optional prefill delay (ms) for mixed-load workloads

  prefill_delay_ms: 2
  # MoE index cache (size in MB)

  index_cache_size: 256

```

Load this configuration when starting the engine:

```bash
sglang --config pd_separated.yaml serve

```

The engine spins up two worker threads—one for prefill and one for decode—and automatically caches prefix KV-states.

### xLLM (C++ API)

For `xLLM` integrations, use the builder pattern to enable separation:

```cpp
auto engine = xllm::Engine::Builder()
                  .model("zai-org/GLM-5.2")
                  .prefill_decode_separate(true)
                  .enable_prefix_caching(true)
                  .prefill_delay_ms(2)
                  .build();

std::string prompt = "What are the design trade-offs of PD separation?";
auto result = engine->generate(prompt, max_tokens=128);
std::cout << result.text << std::endl;

```

All three backends expose the same core knobs: separation flags, prefix caching, and optional delay scheduling.

## Key Source Files and Configuration

The following files in the `zai-org/GLM-5` repository contain the implementation details and configuration schemas for PD separation:

- **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)** – Documents the six optimization techniques, including PD separation and prefix caching strategies for Ascend hardware.
- **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)** (serve section) – Lists supported inference frameworks and points to the configuration flags required for PD separation.
- **[`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md)** – References PD separation as a capability in the skill registry for downstream agent integration.

## Summary

- **PD separation** eliminates throughput instability by isolating prefill and decode execution streams.
- **Prefix caching** stores prompt KV-cache in fast memory, preventing recomputation during decode phases.
- **Prefill delay scheduling** (2ms) smooths high-concurrency workloads to prevent decode stalls.
- **Expert-path caching** optimizes MoE routing decisions for the prefill stage.
- Configure separation using `prefill_decode_separate` and `enable_prefix_caching` flags in vLLM, SGLang, or xLLM.

## Frequently Asked Questions

### What causes throughput instability in GLM-5.2 without PD separation?

Without separation, prefill and decode operations compete for the same compute resources. Long prefills block short decodes, creating latency jitter and uneven request processing times that degrade service level objectives in high-concurrency deployments.

### How does prefix caching improve decode performance?

Prefix caching retains the KV-cache computed during the initial prefill in fast-access memory. Subsequent decode calls reuse these cached states instead of recomputing them, reducing per-token latency and eliminating jitter caused by redundant computation of the same prompt prefixes.

### What configuration flags enable PD separation in GLM-5.2?

Set `prefill_decode_separate=true` to enable stream separation, `enable_prefix_caching=true` to retain prefix KV-cache, and optionally `prefill_delay_ms=2` to add artificial delay for load smoothing. For MoE models, configure `index_cache_size` to cache expert routing decisions.

### Can PD separation be used with mixture-of-experts (MoE) architectures?

Yes. GLM-5.2 implements `IndexCache` specifically for MoE models to cache expert routing decisions made during the prefill phase. This prevents re-routing overhead during decode, ensuring the separated streams maintain high throughput even with complex expert architectures.