# How to Optimize Memory Usage with Chunked Prefill in GLM-5.2

> Optimize GLM-5.2 memory usage with Chunked Prefill. Learn how to set flags and tune chunk size to drastically reduce KV-cache memory for large prompts.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: performance
- Published: 2026-06-21

---

**Enable Chunked Prefill by setting the `--enable-chunked-prefill` flag and tuning `--prefill-chunk-size` to split the prefill computation into smaller chunks, dramatically reducing peak KV-cache memory when processing 1M-token prompts with GLM-5.2.**

GLM-5.2 (also known as *GLM-S.2* in community discussions) implements a **prefill-decode** execution model that can exhaust GPU memory when handling long-context inputs. By configuring Chunked Prefill from the `zai-org/GLM-5` repository, you can shard the prefill stage into manageable pieces, allowing the model to run on limited hardware without out-of-memory (OOM) failures.

## Understanding the Memory Challenge

GLM-5.2 processes prompts using a two-phase approach: first, it "prefills" the KV-cache for the entire prompt, then "decodes" new tokens autoregressively. When prompts approach the 1M-token range, the KV-cache allocation dominates GPU memory, causing single-GPU deployments to crash. The Ascend-NPU optimization guide in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) (lines 9-10) identifies this as a critical bottleneck for high-concurrency scheduling.

## How Chunked Prefill Reduces Memory Usage

**Chunked Prefill** mitigates memory pressure by splitting the prefill stage into sequential chunks rather than allocating cache for the entire prompt at once. As implemented in GLM-5.2, this technique delivers two primary benefits:

- **Peak-memory reduction** – Only the KV-cache for the active chunk resides in GPU memory at any moment, lowering the maximum footprint required for long sequences.
- **Improved concurrency** – Independent chunk processing allows the scheduler to overlap prefill work with other decoding jobs, smoothing GPU utilization curves.

## Configuration Steps

Activate Chunked Prefill in your inference server by following these configuration steps:

### 1. Enable the Runtime Flag

Start the server with `--enable-chunked-prefill` to switch from monolithic prefill to chunked mode. This flag tells the engine to process prompts in segments rather than attempting to cache the full sequence at once.

### 2. Tune the Prefill Chunk Size

Set `--prefill-chunk-size` to control the maximum number of tokens processed per chunk. The default value is approximately 8,192 tokens, but you should adjust this based on your GPU memory budget. Smaller chunks reduce memory usage but may increase scheduling overhead.

### 3. Adjust Batching Limits

Lower `--max-num-batched-tokens` if running concurrent requests. This prevents the scheduler from over-committing memory across multiple simultaneous prefill operations, ensuring each chunk has sufficient headroom.

### 4. Leverage Prefix Caching

Enable prefix caching (often active by default) to retain KV entries for recent prefixes while discarding older chunks. This "PD separation" strategy—documented in the Ascend-NPU guide as "High-Concurrency Scheduling with Prefill Delay"—keeps decoding latency low while freeing memory from processed segments.

## Implementation Examples

Configure Chunked Prefill using popular inference frameworks that support GLM-5.2. Replace `MODEL_ID` with your checkpoint path (e.g., `ZhipuAI/GLM-5.2`).

### vLLM Configuration

```bash
python -m vllm.entrypoints.api_server \
  --model MODEL_ID \
  --max-model-len 1048576 \
  --enable-chunked-prefill \
  --prefill-chunk-size 8192

```

### SGLang Configuration

```bash
python -m sglang.launch_server \
  --model MODEL_ID \
  --max-context-len 1048576 \
  --chunked-prefill true \
  --prefill-chunk-size 8192

```

### xLLM Configuration

```bash
python -m xllm.run_server \
  --model MODEL_ID \
  --max_seq_len 1048576 \
  --chunked_prefill true \
  --prefill_chunk_size 8192

```

## Key Source Files

The GLM-5 repository references Chunked Prefill in the following locations:

- **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)** (lines 9-10) – Describes Chunked Prefill alongside *IndexCache* and *sparse index retrieval* as part of the "Intelligent Caching and Index Optimization" block for Ascend-NPU hardware.
- **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)** – Lists compatible inference frameworks (vLLM, SGLang, KTransformers) in the "Serve GLM-5 Series Locally" section, indicating where Chunked Prefill can be activated.

## Summary

- **Chunked Prefill** splits the GLM-5.2 prefill stage into smaller chunks to prevent OOM errors on 1M-token prompts.
- Configure the feature using `--enable-chunked-prefill` and tune `--prefill-chunk-size` to match your GPU memory capacity.
- Combine with prefix caching (PD separation) to maintain low latency while minimizing memory footprint.
- Reference [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) in the `zai-org/GLM-5` repository for Ascend-NPU specific optimization guidance.

## Frequently Asked Questions

### What is the default prefill chunk size in GLM-5.2?

The default `--prefill-chunk-size` is approximately 8,192 tokens across supported inference engines. You should reduce this value if experiencing memory pressure, or increase it if you have abundant GPU memory and want to reduce scheduling overhead.

### Does Chunked Prefill affect generation latency?

Chunked Prefill may slightly increase prefill latency due to scheduling overhead, but when combined with prefix caching, it maintains low decoding latency. The technique prevents OOM crashes that would otherwise halt generation entirely, making it essential for long-context workloads.

### Can I use Chunked Prefill with multi-GPU setups?

Yes. Chunked Prefill works with tensor-parallel and pipeline-parallel configurations. When using multiple GPUs, ensure `--max-num-batched-tokens` accounts for the distributed memory pool, as each chunk may be sharded across devices.

### Where is Chunked Prefill documented in the GLM-5 repository?

The technique is documented in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) at lines 9-10, where it appears under the "Intelligent Caching and Index Optimization" section alongside other memory optimization strategies for Ascend-NPU hardware.