# How DeepSeek Sparse Attention (DSA) Reduces GLM‑5 Deployment Costs

> Discover how DeepSeek Sparse Attention (DSA) dramatically cuts GLM-5 deployment costs. Learn how it reduces FLOPs and enables inference on 40GB GPUs, saving you money.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: performance
- Published: 2026-06-19

---

**DeepSeek Sparse Attention (DSA) cuts GLM‑5 deployment costs by replacing dense quadratic attention with sparse O(N·k) patterns, reducing per‑token FLOPs by approximately 2.9× at 1‑million‑token contexts and enabling inference on 40 GB GPUs instead of 80 GB.**

DeepSeek Sparse Attention (DSA) is the specialized sparse‑attention mechanism integrated into the GLM‑5 architecture that preserves long‑context capabilities while dramatically lowering compute and memory requirements. According to the `zai-org/GLM-5` source code, DSA transforms the attention complexity from O(N²) to O(N·k), making it feasible to deploy 1‑million‑token context models on significantly cheaper hardware. This article examines the specific architectural innovations that drive these cost reductions and provides implementation examples using the official codebase.

## From Quadratic to Linear: The Sparse Attention Pattern

Standard dense attention mechanisms compute pairwise interactions between all tokens, resulting in **O(N²)** complexity where *N* is the sequence length. DSA replaces this with a pruned attention matrix that operates in **O(N·k)** time, where *k* ≪ *N* represents a small, fixed subset of attended tokens per position.

This complexity reduction directly translates to cost savings. As documented in [[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)](https://github.com/zai-org/GLM-5/blob/main/README.md#L45), the sparse pattern eliminates the majority of floating‑point operations (FLOPs) per token, which lowers GPU utilization, reduces power consumption, and decreases per‑token inference latency.

## Four Key Mechanisms That Lower Deployment Costs

### IndexShare Across Layers

**IndexShare** is a shared indexer reused across every four sparse‑attention layers. Rather than rebuilding a dense index for each layer, the same index‑lookup structure is used repeatedly throughout the forward pass.

This architectural choice cuts per‑token FLOPs by approximately **2.9×** when processing 1‑million‑token contexts, as noted in [[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)](https://github.com/zai-org/GLM-5/blob/main/README.md#L25). By eliminating redundant index computation, IndexShare minimizes the overhead typically associated with sparse attention implementations.

### Mixture‑of‑Token‑Parallel (MTP) Layer Improvements

The **MTP** layers implement speculative decoding that extends the accepted generation length. This mechanism enables the model to generate longer sequences without recomputing the entire context at each generation step.

Fewer forward passes directly reduce the number of required attention calculations, resulting in lower latency and decreased overall compute consumption during autoregressive generation.

### Chunked Prefill and Sparse Index Retrieval

The inference engine uses **Chunked Prefill** to process only active context chunks while fetching the sparse index on‑the‑fly. This approach avoids materializing the full attention matrix for the entire prompt.

Memory usage drops from **O(N²)** to **O(N·k)**, allowing deployment on 40 GB GPUs instead of the 80 GB hardware typically required for dense attention at 1‑million‑token contexts. This memory reduction is critical for cost‑effective deployment, as demonstrated in the Ascend NPU examples located in [[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md).

### Sparse‑Attention Pattern Pruning

By attending only to a subset of token pairs rather than the full quadratic matrix, DSA fundamentally reduces the number of FLOPs per token. This pruning preserves the model’s long‑context capability while dramatically cutting the computational resources required for inference.

## How DSA Integrates With the GLM‑5 Architecture

DSA operates as a drop‑in replacement for dense attention within the transformer blocks:

1. **Input embedding → Sparse‑attention blocks** – Each block receives a sparse attention mask generated by the shared indexer.
2. **Residual connections** remain dense, preserving the quality of representations and ensuring gradient flow is not compromised.
3. **Speculative decoding (MTP)** reuses existing context to predict future tokens, further shrinking the number of required attention calculations.

Because the bulk of inference cost in transformer models stems from the attention matrix, moving from dense to sparse attention yields the most significant savings. The indexed sharing across layers means the same sparsity pattern is reused, avoiding duplicated work and keeping the overhead of the sparsity logic minimal.

## Implementation Guide

### Enabling DSA in Hugging Face Transformers

The following example demonstrates how to load GLM‑5 with DSA support enabled. The `use_sparse_attention=True` flag activates the optimized kernels, while the underlying implementation automatically handles the sparse index management.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

# Load the GLM‑5 model that includes DSA support

model_name = "zai-org/GLM-5"
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Enable the sparse‑attention implementation

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="float16",            # Use bf16/float16 to further cut memory

    device_map="auto",                # Distribute across GPUs if available

    trust_remote_code=True,
    use_sparse_attention=True,        # Activates DSA‑aware kernels

)

# Simple generation with a 1 M‑token context (truncated for the example)

prompt = "Explain the principle of sparse attention in large language models."
input_ids = tokenizer(prompt, return_tensors="pt").input_ids.to(model.device)

outputs = model.generate(
    input_ids,
    max_new_tokens=200,
    do_sample=False,
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

```

### Deploying With vLLM

vLLM automatically detects DSA‑compatible models and switches to optimized kernels without manual configuration:

```python
from vllm import LLM, SamplingParams

model = LLM(
    model="zai-org/GLM-5",
    tensor_parallel_size=2,
    dtype="bfloat16",
    # vLLM automatically activates the DSA kernels when available

)

sampling_params = SamplingParams(max_tokens=150)
prompt = "What are the benefits of sparse attention for long‑context LLMs?"
outputs = model.generate(prompt, sampling_params)
print(outputs[0].text)

```

Both implementations illustrate that **the user does not need to manually manipulate sparse‑attention internals**; the underlying libraries detect the DSA‑compatible model and switch to optimized kernels, delivering lower latency and reduced GPU memory footprint.

## Key Source Files and References

The following files in the `zai-org/GLM-5` repository provide authoritative documentation on DSA implementation:

- **[[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)](https://github.com/zai-org/GLM-5/blob/main/README.md)** – Lines 25‑46 introduce the 2.9× FLOP reduction and cost‑saving impact of DSA.
- **[[`README_zh.md`](https://github.com/zai-org/GLM-5/blob/main/README_zh.md)](https://github.com/zai-org/GLM-5/blob/main/README_zh.md)** – Chinese counterpart describing the same sparse‑attention features.
- **[[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)** – Demonstrates how DSA enables deployment on Ascend NPU hardware with limited memory.
- **External Documentation** – The Transformers model page documents DSA‑aware kernel flags at [`huggingface/transformers/docs/source/en/model_doc/glm_moe_dsa.md`](https://github.com/zai-org/GLM-5/blob/main/huggingface/transformers/docs/source/en/model_doc/glm_moe_dsa.md).

## Summary

- **DeepSeek Sparse Attention (DSA)** reduces GLM‑5 attention complexity from **O(N²)** to **O(N·k)**, cutting per‑token compute costs.
- **IndexShare** reuse across layers delivers approximately **2.9× FLOP reduction** at 1‑million‑token contexts.
- **Memory efficiency** allows deployment on 40 GB GPUs instead of 80 GB, significantly lowering hardware requirements.
- **Automatic kernel selection** in Transformers and vLLM means developers do not manually configure sparse patterns.
- **MTP speculative decoding** and **Chunked Prefill** further reduce latency by minimizing forward passes and avoiding full attention matrix materialization.

## Frequently Asked Questions

### What is the computational complexity reduction achieved by DSA in GLM‑5?

DSA reduces the attention mechanism complexity from **O(N²)** to **O(N·k)**, where *k* represents a small, fixed subset of tokens attended to per position. This change cuts the number of floating‑point operations per token by approximately 2.9× when processing 1‑million‑token sequences, as documented in the [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) file.

### How does IndexShare contribute to cost reduction?

IndexShare is a shared indexer reused across every four sparse‑attention layers. By avoiding the need to rebuild a dense index for each layer, it eliminates redundant computation and reduces per‑token FLOPs. This shared structure keeps the overhead of sparse attention minimal while maintaining the model’s 1‑million‑token context capability.

### Can DSA run on consumer GPUs, or does it require specialized hardware?

DSA enables GLM‑5 to run on 40 GB GPUs rather than the 80 GB hardware typically required for dense attention models at 1‑million‑token contexts. The [[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) file additionally demonstrates deployment on Ascend NPU hardware, indicating broad compatibility with memory‑constrained environments.

### Do I need to manually configure sparse attention patterns when using GLM‑5?

No. When using supported backends like Hugging Face Transformers or vLLM, the libraries automatically detect DSA‑compatible models and switch to optimized kernels. Users can explicitly enable sparse attention with `use_sparse_attention=True` in Transformers, but the underlying sparse patterns and index management are handled automatically by the model implementation.