# How GLM-5.2's IndexShare Architecture Reduces FLOPs by 2.9× at 1M Context

> Discover how GLM-5.2's IndexShare architecture cuts FLOPs by 2.9× at 1M context. Learn how reusing attention indexers minimizes computational costs for efficient LLM processing.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: performance
- Published: 2026-06-21

---

**GLM-5.2's IndexShare architecture reduces FLOPs by reusing a single sparse-attention indexer across every four consecutive layers, eliminating three expensive index-building passes per layer group and cutting per-token computational costs by approximately 2.9× when processing 1 million tokens.**

The zai-org/GLM-5 repository introduces **IndexShare** as a core optimization in the GLM-S series, specifically engineered to handle ultra-long context windows without linear increases in computational overhead. By treating index construction as a shared resource rather than a per-layer operation, this sparse-attention scheme fundamentally changes how transformer models scale to million-token sequences.

## What Is IndexShare?

**IndexShare** is a sparse-attention mechanism that decouples the expensive index-building process from individual attention layers. In traditional sparse attention, each layer constructs its own index to identify which tokens to attend to—a process involving sorting and bucket-creation across the full sequence. IndexShare instead builds one **indexer** and reuses it across multiple consecutive layers, amortizing the computational cost.

According to the repository's [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md), this architecture "reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M-token context length." The design is detailed in the research paper (arXiv 2603.12201) referenced in [`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md), which provides the theoretical justification for the FLOP savings.

## How IndexShare Reduces Computational Cost

The FLOP reduction in GLM-5.2 stems from eliminating redundant index construction while maintaining attention quality. The mechanism operates through three key principles:

### Sharing the Indexer Across Layer Groups

Instead of building a fresh index for every sparse-attention layer, the model constructs one indexer and shares it across the next three layers. This means in every four-layer block, the expensive sorting and bucket-creation operations happen only once rather than four times.

The indexer construction dominates the computational cost of sparse-attention blocks because it requires processing the entire sequence. By sharing this resource, GLM-5.2 avoids doing this work three extra times per four-layer group, directly reducing the total operation count.

### Maintaining Constant Per-Token Operations

Each sparse-attention layer still attends to only a small subset of the full sequence—the "top-k" tokens identified by the shared index. Because the same index is used repeatedly across the four-layer group, the number of new token-to-token interactions does not increase with depth.

This design keeps per-token FLOPs roughly constant across the shared layers, preventing the linear growth in computational cost that typically accompanies increased model depth in standard sparse attention architectures.

### Measured FLOP Reduction at 1M Context

At 1 million tokens, the savings compound significantly. The authors measured the total operations for a 1M-token window and documented a **2.9-fold reduction** in per-token FLOPs compared to a baseline that builds a fresh index for every layer.

The saved FLOPs primarily come from eliminating three index-building passes per four-layer block, which dominate the cost when the context is very long. This makes it feasible to deploy 1M-token contexts on single GPU or NPU hardware without exhausting the FLOP budget.

## Implementation References in the GLM-5 Repository

The IndexShare architecture is documented across several key files in the zai-org/GLM-5 repository:

- **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)** – Contains the high-level description of the 2.9× FLOP reduction and the four-layer sharing pattern
- **[`README_zh.md`](https://github.com/zai-org/GLM-5/blob/main/README_zh.md)** – Provides the Chinese-language technical specifications confirming the same measurements
- **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)** – Demonstrates how IndexShare enables 1M-token deployment on Ascend NPUs by reducing computational overhead
- **[`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md)** – Lists the GLM-5 family and points to the technical report where IndexShare is introduced

These files collectively establish that IndexShare is not merely a theoretical optimization but a implemented feature in the GLM-5.2 production model.

## Running GLM-5.2 with IndexShare

When loading GLM-5.2 through hardware-accelerated inference engines like vLLM or SGLang, the IndexShare schedule activates automatically for long-context windows. The following examples demonstrate how to instantiate the model with 1M-token contexts where the FLOP reduction occurs.

### Using vLLM

```python
from vllm import LLM, SamplingParams

# Load GLM-5.2 with IndexShare optimization enabled for 1M tokens

llm = LLM(model="zai-org/GLM-5.2", max_seq_len=1_048_576)

sampling_params = SamplingParams(
    temperature=0.7,
    max_tokens=256,
    top_p=0.9,
)

prompt = "Explain the benefits of Index-Share in long-context LLMs."
outputs = llm.generate([prompt], sampling_params)

print(outputs[0].text)

```

### Using SGLang

```python
import sglang as sgl

# Initialize GLM-5.2 with maximum context window

model = sgl.LLM("zai-org/GLM-5.2", max_seq_len=1_048_576)

prompt = "What is the FLOP reduction of Index-Share compared to vanilla sparse attention?"
gen = model.generate(prompt, max_new_tokens=256, temperature=0.7)

print(gen.text)

```

Both configurations set `max_seq_len` to 1,048,576 tokens, triggering the IndexShare schedule inside the model's attention implementation. During inference, the system builds an index once per four layers, achieving the reported 2.9× FLOP reduction compared to per-layer index construction.

## Summary

- **IndexShare reuses one indexer across four sparse-attention layers**, eliminating three expensive index-building operations per layer group.
- **Per-token FLOPs remain constant** across shared layers because the same top-k token index is reused, preventing computational growth with depth.
- **GLM-5.2 achieves 2.9× FLOP reduction** at 1M-token contexts, as documented in the repository's [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) and confirmed by arXiv 2603.12201.
- **Deployment is supported** through inference engines like vLLM and SGLang, with specific guidance available in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) for Ascend NPU hardware.

## Frequently Asked Questions

### What exactly is an "indexer" in GLM-5.2's IndexShare architecture?

An **indexer** is the data structure constructed during sparse attention that identifies which tokens in the sequence should be attended to. In GLM-5.2, building this indexer requires sorting and bucket-creation operations over the full sequence length, making it the most computationally expensive part of sparse attention. IndexShare treats this indexer as a shared resource across four consecutive layers to avoid recomputing it.

### How does IndexShare compare to standard sparse attention in terms of FLOPs?

Standard sparse attention rebuilds the indexer at every layer, resulting in linear growth of index-building operations with model depth. **IndexShare reduces these operations by 75%** in four-layer blocks (building one index instead of four), which translates to a **2.9× reduction in total per-token FLOPs** at 1M-token contexts. The savings become more pronounced as sequence length increases because index construction complexity scales with sequence size.

### Can I use GLM-5.2's IndexShare with standard transformers inference libraries?

Yes. The IndexShare mechanism is implemented internally in the GLM-5.2 model weights and architecture, so it works automatically with compatible inference engines. The repository provides specific examples for **vLLM** and **SGLang**, as well as deployment guidance for Ascend NPUs in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md). As long as the engine supports the GLM-5.2 architecture, the FLOP-reducing IndexShare schedule activates when processing long contexts.

### Where is the IndexShare mechanism documented in the GLM-5 repository?

The primary documentation appears in the repository's **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)** and **[`README_zh.md`](https://github.com/zai-org/GLM-5/blob/main/README_zh.md)** files, which specify the 2.9× FLOP reduction and four-layer sharing pattern. Technical context and paper references are located in **[`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md)**, while deployment-specific instructions for hardware like Ascend NPUs are detailed in **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)**. The original research paper (arXiv 2603.12201) provides the complete theoretical framework.