How GLM-5's IndexShare Architecture Reduces FLOPs by 2.9× at 1M Token Context

GLM-5's IndexShare architecture reduces FLOPs by reusing a single indexer across every four sparse-attention layers, eliminating redundant index construction passes and cutting per-token computational costs by approximately 2.9× when processing 1 million tokens.

The zai-org/GLM-5 repository introduces IndexShare, a sparse-attention mechanism designed to dramatically lower the computational overhead of long-context language modeling. By sharing indexer state across multiple transformer layers, this architecture avoids the expensive sorting and bucket-creation operations that typically dominate sparse attention costs. According to the source code and documentation in zai-org/GLM-5, this approach enables efficient processing of 1 million token contexts on single GPU or NPU hardware without exhausting the FLOP budget.

The IndexShare Design Pattern

IndexShare implements a sparse-attention scheme that reuses computational artifacts across depth rather than recomputing them at every layer. In traditional sparse attention, each layer constructs its own index to identify which tokens to attend to, requiring expensive operations over the full sequence length. IndexShare modifies this by building one index and sharing it across four consecutive sparse-attention layers, effectively amortizing the construction cost over multiple processing steps.

This design is particularly effective for the GLM-S series (including GLM-5.2), where the model must handle extremely long contexts up to 1 million tokens. The architecture is documented in the repository's README.md and README_zh.md files, with theoretical justification provided in the research paper arXiv 2603.12201.

How IndexShare Reduces Computational Costs

The FLOP reduction stems from three specific technical optimizations that address the most expensive components of sparse attention:

1. Eliminating Redundant Index Construction

Index construction is the most expensive part of a sparse-attention block, requiring sorting and bucket-creation operations over the full sequence length. Instead of building a fresh index for every layer, IndexShare constructs one index and shares it across the next three layers. This eliminates three index-building passes per four-layer block, which dominate the computational cost when processing very long contexts.

2. Maintaining Constant Per-Token Interactions

Each sparse-attention layer attends to only a small subset of the full sequence (the "top-k" tokens identified by the shared index). Because the same index is used repeatedly across four layers, the number of new token-to-token interactions does not increase with depth. The per-token FLOPs remain roughly constant across the four-layer group rather than multiplying with each additional layer.

3. Quantified FLOP Savings at Scale

At a 1 million token context length, the authors measured total operations and found a 2.9-fold FLOP reduction compared with a baseline that builds a fresh index for every layer. The savings scale with sequence length because index construction complexity grows with the number of tokens, making the optimization increasingly critical for long-context applications.

Deployment in Inference Engines

The FLOP reduction manifests when running on hardware-accelerated inference engines such as vLLM, SGLang, or KTransformers. These engines instantiate the sparse-attention kernels that follow the IndexShare schedule, ensuring runtime costs scale linearly with the number of index constructions rather than with the number of layers. This linear scaling makes 1 million token contexts feasible on single GPU or Ascend NPU deployments, as documented in example/ascend.md.

Code Examples: Running GLM-5.2 with IndexShare

To leverage the IndexShare optimization, load the GLM-5.2 model with a 1 million token context window using popular inference backends. The following examples demonstrate the configuration required to activate the IndexShare schedule:

Using vLLM:

from vllm import LLM, SamplingParams

# Load the GLM‑5.2 model (the model already implements Index‑Share internally)

llm = LLM(model="zai-org/GLM-5.2", max_seq_len=1_048_576)   # 1 M tokens

sampling_params = SamplingParams(
    temperature=0.7,
    max_tokens=256,
    top_p=0.9,
)

prompt = "Explain the benefits of Index‑Share in long‑context LLMs."
outputs = llm.generate([prompt], sampling_params)

print(outputs[0].text)

Using SGLang:

import sglang as sgl

# Initialize the GLM‑5.2 model with a huge context window

model = sgl.LLM("zai-org/GLM-5.2", max_seq_len=1_048_576)

prompt = "What is the FLOP reduction of Index‑Share compared to vanilla sparse attention?"
gen = model.generate(prompt, max_new_tokens=256, temperature=0.7)

print(gen.text)

Both configurations set max_seq_len to 1,048,576 tokens, triggering the IndexShare schedule inside the model's attention implementation. During processing, the model builds an index once per four layers, yielding the reported 2.9× FLOP reduction.

Key Files and References

The GLM-5 repository contains several critical files documenting the IndexShare architecture:

  • README.md – High-level description of IndexShare and the 2.9× FLOP savings at 1M token context
  • README_zh.md – Chinese version of the technical explanation
  • example/ascend.md – Deployment instructions for running GLM-5.2 with IndexShare on Ascend NPUs
  • skills/glm-master-skill/SKILL.md – Lists the GLM-5 family and references the technical report introducing IndexShare

These files, combined with the research paper arXiv 2603.12201, provide the complete theoretical and implementation framework for the architecture.

Summary

  • IndexShare reuses a single indexer across every four sparse-attention layers, eliminating redundant index construction that typically dominates sparse attention costs.
  • The architecture achieves a 2.9× reduction in per-token FLOPs when processing 1 million token contexts by avoiding three index-building passes per four-layer block.
  • Inference engines like vLLM and SGLang implement the IndexShare schedule, enabling linear scaling of computational costs with respect to index constructions rather than layer depth.
  • The implementation is documented in zai-org/GLM-5 files including README.md, README_zh.md, and example/ascend.md, with support for both GPU and Ascend NPU deployment.

Frequently Asked Questions

What exactly is IndexShare in GLM-5?

IndexShare is a sparse-attention architecture that constructs a single index for identifying relevant tokens and shares that index across four consecutive attention layers. Instead of recomputing the expensive sorting and bucket-creation operations at every layer, the model reuses the existing index, which dramatically reduces the computational overhead required for long-context processing.

How much does IndexShare reduce FLOPs compared to standard sparse attention?

At a 1 million token context length, IndexShare reduces per-token FLOPs by approximately 2.9× compared to a baseline that builds a fresh index for every layer. This saving comes from eliminating three index construction passes per four-layer block, which are the most computationally expensive operations in sparse attention when handling very long sequences.

Where is the IndexShare implementation documented in the GLM-5 repository?

The architecture is described in the repository's README.md and README_zh.md files, which explain the FLOP reduction mechanism and performance metrics. Additional deployment context appears in example/ascend.md for Ascend NPU configurations, while skills/glm-master-skill/SKILL.md provides ecosystem context and references the original research paper.

Can I use IndexShare with standard inference frameworks like vLLM?

Yes, IndexShare is fully compatible with standard inference engines including vLLM, SGLang, and KTransformers. When loading zai-org/GLM-5.2 with a 1 million token context window, these engines automatically utilize the sparse-attention kernels that implement the IndexShare schedule, ensuring you receive the full 2.9× FLOP reduction benefit during inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →