How to Configure DeepSeek Sparse Attention for GLM-5.2 Long-Context Capacity
To enable DeepSeek Sparse Attention (DSA) in GLM-5.2, set the attention_type configuration field to "deepseek", which activates a block-sparse attention pattern that reduces per-token FLOPs by roughly two-thirds while preserving the ability to attend across 1 million tokens.
The GLM-5 series from zai-org/GLM-5 integrates DeepSeek Sparse Attention (DSA) to handle ultra-long contexts efficiently. This implementation replaces the standard dense self-attention matrix with a specialized block-sparse pattern that uses an IndexShare design, sharing a common indexer across every four attention layers to minimize memory overhead without sacrificing long-range retrieval capabilities.
Understanding DeepSeek Sparse Attention Architecture
DSA in GLM-5.2 employs a block-sparse pattern that strategically drops most of the quadratic computational cost associated with full self-attention. According to the repository's README.md, this design allows the model to maintain attention over contexts up to 1 million tokens while significantly reducing inference FLOPs.
The IndexShare mechanism is critical to this efficiency. By sharing the sparse indexer across every four attention layers, the model avoids redundant memory allocations and cache overheads that typically plague long-context transformers. This architecture is particularly essential for agentic tasks requiring long-horizon reasoning, where the model must retrieve relevant tokens from distant positions in the sequence.
Prerequisites
Before configuring DSA, ensure your environment meets the following requirements:
- Hugging Face Transformers version 0.5.12 or higher (
pip install "transformers>=0.5.12") - Compatible inference backend (Transformers, vLLM, SGLang, or KTransformers)
- Sufficient GPU memory for 1M token contexts (exact requirements vary by quantization)
Configuration Methods
The attention_type field in config.json controls whether DSA is active. You can configure this through several inference stacks.
Hugging Face Transformers
Modify the model configuration before loading weights to enable the sparse kernel:
from transformers import AutoConfig, AutoModelForCausalLM, AutoTokenizer
import torch
# Load the default GLM-5.2 configuration
cfg = AutoConfig.from_pretrained("zai-org/GLM-5.2")
# Enable DeepSeek Sparse Attention
cfg.attention_type = "deepseek"
# Load model with modified config
model = AutoModelForCausalLM.from_pretrained(
"zai-org/GLM-5.2",
config=cfg,
torch_dtype=torch.float16,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("zai-org/GLM-5.2")
The attention_type field is documented in the Transformers DSA documentation at glm_moe_dsa.md.
vLLM Backend
When using vLLM for high-throughput serving, pass the attention type via the Python API or command line:
import vllm
engine = vllm.LLM(
model="zai-org/GLM-5.2",
attention_type="deepseek"
)
Alternatively, use the CLI argument --attention-type deepseek when launching the server.
SGLang Backend
For SGLang deployments, include the attention configuration in your model's YAML specification:
model:
name: "zai-org/GLM-5.2"
attention_type: "deepseek"
KTransformers Backend
When using KTransformers for optimized inference, specify the attention type during object construction:
from ktransformers import KTransformer
kt = KTransformer(
"zai-org/GLM-5.2",
attention_type="deepseek"
)
Running Long-Context Inference
Once DSA is configured, the model can process contexts up to 1 million tokens. The following example demonstrates generating text with an ultra-long prompt:
# Create a long context (approximately 1M tokens)
prompt = "Explain the theory of relativity. " * 200_000
inputs = tokenizer(prompt, return_tensors="pt", truncation=False)
# Generate output
output_ids = model.generate(
input_ids=inputs["input_ids"],
max_new_tokens=128,
do_sample=False,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
With DSA active, this generation runs at approximately one-third the FLOPs of dense attention, while maintaining full access to the entire 1 million token context window. For additional optimization techniques including index caching on Ascend NPUs, refer to example/ascend.md in the repository.
Summary
- DeepSeek Sparse Attention reduces GLM-5.2 inference costs by implementing a block-sparse pattern with IndexShare across every four layers.
- Enable DSA by setting
attention_type="deepseek"in the model configuration before loading weights. - Compatible backends include Hugging Face Transformers (≥0.5.12), vLLM, SGLang, and KTransformers, all using the same configuration key.
- Configuration sources include
README.md,README_zh.md, and the Transformers documentation atglm_moe_dsa.md. - DSA unlocks 1 million token contexts with approximately 66% reduction in per-token FLOPs compared to dense attention.
Frequently Asked Questions
What is DeepSeek Sparse Attention?
DeepSeek Sparse Attention (DSA) is a block-sparse attention mechanism that replaces the full quadratic self-attention matrix with a structured sparse pattern. In GLM-5.2, DSA implements an IndexShare design where a common indexer is shared across every four attention layers, reducing memory overhead while allowing the model to attend to positions up to 1 million tokens away.
How much memory does DSA save compared to dense attention?
DSA reduces the per-token FLOPs by roughly two-thirds compared to dense attention, translating to significant memory savings during inference. While exact memory reduction depends on sequence length and batch size, the block-sparse pattern eliminates the need to materialize the full attention matrix, which is the primary memory bottleneck in long-context models.
Can I use DSA with existing GLM-5.2 checkpoints?
Yes, DSA is compatible with existing zai-org/GLM-5.2 checkpoints without requiring model retraining or weight conversion. The sparse attention pattern is activated at runtime through the attention_type configuration field, meaning the same checkpoint can be used for both dense and sparse attention modes depending on your inference requirements.
Which inference engines support DSA for GLM-5.2?
DSA is supported by Hugging Face Transformers (version 0.5.12+), vLLM, SGLang, and KTransformers. All four frameworks use the same underlying sparse kernel implementation and configuration pattern (attention_type="deepseek"), allowing seamless migration between serving stacks while maintaining the long-context capabilities of GLM-5.2.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →