# How to Configure DeepSeek Sparse Attention for GLM-5.2 Long-Context Capacity

> Configure DeepSeek Sparse Attention for GLM-5.2 to achieve 1 million token context with reduced FLOPs. Learn how to set attention_type deepseek for efficient long-context processing.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: how-to-guide
- Published: 2026-06-21

---

**To enable DeepSeek Sparse Attention (DSA) in GLM-5.2, set the `attention_type` configuration field to `"deepseek"`, which activates a block-sparse attention pattern that reduces per-token FLOPs by roughly two-thirds while preserving the ability to attend across 1 million tokens.**

The GLM-5 series from `zai-org/GLM-5` integrates DeepSeek Sparse Attention (DSA) to handle ultra-long contexts efficiently. This implementation replaces the standard dense self-attention matrix with a specialized block-sparse pattern that uses an **IndexShare** design, sharing a common indexer across every four attention layers to minimize memory overhead without sacrificing long-range retrieval capabilities.

## Understanding DeepSeek Sparse Attention Architecture

DSA in GLM-5.2 employs a **block-sparse pattern** that strategically drops most of the quadratic computational cost associated with full self-attention. According to the repository's [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md), this design allows the model to maintain attention over contexts up to 1 million tokens while significantly reducing inference FLOPs.

The **IndexShare** mechanism is critical to this efficiency. By sharing the sparse indexer across every four attention layers, the model avoids redundant memory allocations and cache overheads that typically plague long-context transformers. This architecture is particularly essential for agentic tasks requiring long-horizon reasoning, where the model must retrieve relevant tokens from distant positions in the sequence.

## Prerequisites

Before configuring DSA, ensure your environment meets the following requirements:

- **Hugging Face Transformers** version 0.5.12 or higher (`pip install "transformers>=0.5.12"`)
- Compatible inference backend (Transformers, vLLM, SGLang, or KTransformers)
- Sufficient GPU memory for 1M token contexts (exact requirements vary by quantization)

## Configuration Methods

The `attention_type` field in [`config.json`](https://github.com/zai-org/GLM-5/blob/main/config.json) controls whether DSA is active. You can configure this through several inference stacks.

### Hugging Face Transformers

Modify the model configuration before loading weights to enable the sparse kernel:

```python
from transformers import AutoConfig, AutoModelForCausalLM, AutoTokenizer
import torch

# Load the default GLM-5.2 configuration

cfg = AutoConfig.from_pretrained("zai-org/GLM-5.2")

# Enable DeepSeek Sparse Attention

cfg.attention_type = "deepseek"

# Load model with modified config

model = AutoModelForCausalLM.from_pretrained(
    "zai-org/GLM-5.2",
    config=cfg,
    torch_dtype=torch.float16,
    device_map="auto",
)

tokenizer = AutoTokenizer.from_pretrained("zai-org/GLM-5.2")

```

The `attention_type` field is documented in the Transformers DSA documentation at [`glm_moe_dsa.md`](https://github.com/zai-org/GLM-5/blob/main/glm_moe_dsa.md).

### vLLM Backend

When using vLLM for high-throughput serving, pass the attention type via the Python API or command line:

```python
import vllm

engine = vllm.LLM(
    model="zai-org/GLM-5.2",
    attention_type="deepseek"
)

```

Alternatively, use the CLI argument `--attention-type deepseek` when launching the server.

### SGLang Backend

For SGLang deployments, include the attention configuration in your model's YAML specification:

```yaml
model:
  name: "zai-org/GLM-5.2"
  attention_type: "deepseek"

```

### KTransformers Backend

When using KTransformers for optimized inference, specify the attention type during object construction:

```python
from ktransformers import KTransformer

kt = KTransformer(
    "zai-org/GLM-5.2",
    attention_type="deepseek"
)

```

## Running Long-Context Inference

Once DSA is configured, the model can process contexts up to 1 million tokens. The following example demonstrates generating text with an ultra-long prompt:

```python

# Create a long context (approximately 1M tokens)

prompt = "Explain the theory of relativity. " * 200_000
inputs = tokenizer(prompt, return_tensors="pt", truncation=False)

# Generate output

output_ids = model.generate(
    input_ids=inputs["input_ids"],
    max_new_tokens=128,
    do_sample=False,
)

print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

```

With DSA active, this generation runs at approximately one-third the FLOPs of dense attention, while maintaining full access to the entire 1 million token context window. For additional optimization techniques including index caching on Ascend NPUs, refer to [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) in the repository.

## Summary

- **DeepSeek Sparse Attention** reduces GLM-5.2 inference costs by implementing a block-sparse pattern with IndexShare across every four layers.
- Enable DSA by setting `attention_type="deepseek"` in the model configuration before loading weights.
- Compatible backends include Hugging Face Transformers (≥0.5.12), vLLM, SGLang, and KTransformers, all using the same configuration key.
- Configuration sources include [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md), [`README_zh.md`](https://github.com/zai-org/GLM-5/blob/main/README_zh.md), and the Transformers documentation at [`glm_moe_dsa.md`](https://github.com/zai-org/GLM-5/blob/main/glm_moe_dsa.md).
- DSA unlocks 1 million token contexts with approximately 66% reduction in per-token FLOPs compared to dense attention.

## Frequently Asked Questions

### What is DeepSeek Sparse Attention?

DeepSeek Sparse Attention (DSA) is a block-sparse attention mechanism that replaces the full quadratic self-attention matrix with a structured sparse pattern. In GLM-5.2, DSA implements an **IndexShare** design where a common indexer is shared across every four attention layers, reducing memory overhead while allowing the model to attend to positions up to 1 million tokens away.

### How much memory does DSA save compared to dense attention?

DSA reduces the per-token FLOPs by roughly two-thirds compared to dense attention, translating to significant memory savings during inference. While exact memory reduction depends on sequence length and batch size, the block-sparse pattern eliminates the need to materialize the full attention matrix, which is the primary memory bottleneck in long-context models.

### Can I use DSA with existing GLM-5.2 checkpoints?

Yes, DSA is compatible with existing `zai-org/GLM-5.2` checkpoints without requiring model retraining or weight conversion. The sparse attention pattern is activated at runtime through the `attention_type` configuration field, meaning the same checkpoint can be used for both dense and sparse attention modes depending on your inference requirements.

### Which inference engines support DSA for GLM-5.2?

DSA is supported by Hugging Face Transformers (version 0.5.12+), vLLM, SGLang, and KTransformers. All four frameworks use the same underlying sparse kernel implementation and configuration pattern (`attention_type="deepseek"`), allowing seamless migration between serving stacks while maintaining the long-context capabilities of GLM-5.2.