How to Optimize Memory Usage with Chunked Prefill in GLM-5.2
Enable Chunked Prefill by setting the --enable-chunked-prefill flag and tuning --prefill-chunk-size to split the prefill computation into smaller chunks, dramatically reducing peak KV-cache memory when processing 1M-token prompts with GLM-5.2.
GLM-5.2 (also known as GLM-S.2 in community discussions) implements a prefill-decode execution model that can exhaust GPU memory when handling long-context inputs. By configuring Chunked Prefill from the zai-org/GLM-5 repository, you can shard the prefill stage into manageable pieces, allowing the model to run on limited hardware without out-of-memory (OOM) failures.
Understanding the Memory Challenge
GLM-5.2 processes prompts using a two-phase approach: first, it "prefills" the KV-cache for the entire prompt, then "decodes" new tokens autoregressively. When prompts approach the 1M-token range, the KV-cache allocation dominates GPU memory, causing single-GPU deployments to crash. The Ascend-NPU optimization guide in example/ascend.md (lines 9-10) identifies this as a critical bottleneck for high-concurrency scheduling.
How Chunked Prefill Reduces Memory Usage
Chunked Prefill mitigates memory pressure by splitting the prefill stage into sequential chunks rather than allocating cache for the entire prompt at once. As implemented in GLM-5.2, this technique delivers two primary benefits:
- Peak-memory reduction – Only the KV-cache for the active chunk resides in GPU memory at any moment, lowering the maximum footprint required for long sequences.
- Improved concurrency – Independent chunk processing allows the scheduler to overlap prefill work with other decoding jobs, smoothing GPU utilization curves.
Configuration Steps
Activate Chunked Prefill in your inference server by following these configuration steps:
1. Enable the Runtime Flag
Start the server with --enable-chunked-prefill to switch from monolithic prefill to chunked mode. This flag tells the engine to process prompts in segments rather than attempting to cache the full sequence at once.
2. Tune the Prefill Chunk Size
Set --prefill-chunk-size to control the maximum number of tokens processed per chunk. The default value is approximately 8,192 tokens, but you should adjust this based on your GPU memory budget. Smaller chunks reduce memory usage but may increase scheduling overhead.
3. Adjust Batching Limits
Lower --max-num-batched-tokens if running concurrent requests. This prevents the scheduler from over-committing memory across multiple simultaneous prefill operations, ensuring each chunk has sufficient headroom.
4. Leverage Prefix Caching
Enable prefix caching (often active by default) to retain KV entries for recent prefixes while discarding older chunks. This "PD separation" strategy—documented in the Ascend-NPU guide as "High-Concurrency Scheduling with Prefill Delay"—keeps decoding latency low while freeing memory from processed segments.
Implementation Examples
Configure Chunked Prefill using popular inference frameworks that support GLM-5.2. Replace MODEL_ID with your checkpoint path (e.g., ZhipuAI/GLM-5.2).
vLLM Configuration
python -m vllm.entrypoints.api_server \
--model MODEL_ID \
--max-model-len 1048576 \
--enable-chunked-prefill \
--prefill-chunk-size 8192
SGLang Configuration
python -m sglang.launch_server \
--model MODEL_ID \
--max-context-len 1048576 \
--chunked-prefill true \
--prefill-chunk-size 8192
xLLM Configuration
python -m xllm.run_server \
--model MODEL_ID \
--max_seq_len 1048576 \
--chunked_prefill true \
--prefill_chunk_size 8192
Key Source Files
The GLM-5 repository references Chunked Prefill in the following locations:
example/ascend.md(lines 9-10) – Describes Chunked Prefill alongside IndexCache and sparse index retrieval as part of the "Intelligent Caching and Index Optimization" block for Ascend-NPU hardware.README.md– Lists compatible inference frameworks (vLLM, SGLang, KTransformers) in the "Serve GLM-5 Series Locally" section, indicating where Chunked Prefill can be activated.
Summary
- Chunked Prefill splits the GLM-5.2 prefill stage into smaller chunks to prevent OOM errors on 1M-token prompts.
- Configure the feature using
--enable-chunked-prefilland tune--prefill-chunk-sizeto match your GPU memory capacity. - Combine with prefix caching (PD separation) to maintain low latency while minimizing memory footprint.
- Reference
example/ascend.mdin thezai-org/GLM-5repository for Ascend-NPU specific optimization guidance.
Frequently Asked Questions
What is the default prefill chunk size in GLM-5.2?
The default --prefill-chunk-size is approximately 8,192 tokens across supported inference engines. You should reduce this value if experiencing memory pressure, or increase it if you have abundant GPU memory and want to reduce scheduling overhead.
Does Chunked Prefill affect generation latency?
Chunked Prefill may slightly increase prefill latency due to scheduling overhead, but when combined with prefix caching, it maintains low decoding latency. The technique prevents OOM crashes that would otherwise halt generation entirely, making it essential for long-context workloads.
Can I use Chunked Prefill with multi-GPU setups?
Yes. Chunked Prefill works with tensor-parallel and pipeline-parallel configurations. When using multiple GPUs, ensure --max-num-batched-tokens accounts for the distributed memory pool, as each chunk may be sharded across devices.
Where is Chunked Prefill documented in the GLM-5 repository?
The technique is documented in example/ascend.md at lines 9-10, where it appears under the "Intelligent Caching and Index Optimization" section alongside other memory optimization strategies for Ascend-NPU hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →