How GLM-5.2's IndexShare Architecture Reduces FLOPs by 2.9× at 1M Context

GLM-5.2's IndexShare architecture reduces FLOPs by reusing a single sparse-attention indexer across every four consecutive layers, eliminating three expensive index-building passes per layer group and cutting per-token computational costs by approximately 2.9× when processing 1 million tokens.

The zai-org/GLM-5 repository introduces IndexShare as a core optimization in the GLM-S series, specifically engineered to handle ultra-long context windows without linear increases in computational overhead. By treating index construction as a shared resource rather than a per-layer operation, this sparse-attention scheme fundamentally changes how transformer models scale to million-token sequences.

What Is IndexShare?

IndexShare is a sparse-attention mechanism that decouples the expensive index-building process from individual attention layers. In traditional sparse attention, each layer constructs its own index to identify which tokens to attend to—a process involving sorting and bucket-creation across the full sequence. IndexShare instead builds one indexer and reuses it across multiple consecutive layers, amortizing the computational cost.

According to the repository's README.md, this architecture "reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M-token context length." The design is detailed in the research paper (arXiv 2603.12201) referenced in skills/glm-master-skill/SKILL.md, which provides the theoretical justification for the FLOP savings.

How IndexShare Reduces Computational Cost

The FLOP reduction in GLM-5.2 stems from eliminating redundant index construction while maintaining attention quality. The mechanism operates through three key principles:

Sharing the Indexer Across Layer Groups

Instead of building a fresh index for every sparse-attention layer, the model constructs one indexer and shares it across the next three layers. This means in every four-layer block, the expensive sorting and bucket-creation operations happen only once rather than four times.

The indexer construction dominates the computational cost of sparse-attention blocks because it requires processing the entire sequence. By sharing this resource, GLM-5.2 avoids doing this work three extra times per four-layer group, directly reducing the total operation count.

Maintaining Constant Per-Token Operations

Each sparse-attention layer still attends to only a small subset of the full sequence—the "top-k" tokens identified by the shared index. Because the same index is used repeatedly across the four-layer group, the number of new token-to-token interactions does not increase with depth.

This design keeps per-token FLOPs roughly constant across the shared layers, preventing the linear growth in computational cost that typically accompanies increased model depth in standard sparse attention architectures.

Measured FLOP Reduction at 1M Context

At 1 million tokens, the savings compound significantly. The authors measured the total operations for a 1M-token window and documented a 2.9-fold reduction in per-token FLOPs compared to a baseline that builds a fresh index for every layer.

The saved FLOPs primarily come from eliminating three index-building passes per four-layer block, which dominate the cost when the context is very long. This makes it feasible to deploy 1M-token contexts on single GPU or NPU hardware without exhausting the FLOP budget.

Implementation References in the GLM-5 Repository

The IndexShare architecture is documented across several key files in the zai-org/GLM-5 repository:

  • README.md – Contains the high-level description of the 2.9× FLOP reduction and the four-layer sharing pattern
  • README_zh.md – Provides the Chinese-language technical specifications confirming the same measurements
  • example/ascend.md – Demonstrates how IndexShare enables 1M-token deployment on Ascend NPUs by reducing computational overhead
  • skills/glm-master-skill/SKILL.md – Lists the GLM-5 family and points to the technical report where IndexShare is introduced

These files collectively establish that IndexShare is not merely a theoretical optimization but a implemented feature in the GLM-5.2 production model.

Running GLM-5.2 with IndexShare

When loading GLM-5.2 through hardware-accelerated inference engines like vLLM or SGLang, the IndexShare schedule activates automatically for long-context windows. The following examples demonstrate how to instantiate the model with 1M-token contexts where the FLOP reduction occurs.

Using vLLM

from vllm import LLM, SamplingParams

# Load GLM-5.2 with IndexShare optimization enabled for 1M tokens

llm = LLM(model="zai-org/GLM-5.2", max_seq_len=1_048_576)

sampling_params = SamplingParams(
    temperature=0.7,
    max_tokens=256,
    top_p=0.9,
)

prompt = "Explain the benefits of Index-Share in long-context LLMs."
outputs = llm.generate([prompt], sampling_params)

print(outputs[0].text)

Using SGLang

import sglang as sgl

# Initialize GLM-5.2 with maximum context window

model = sgl.LLM("zai-org/GLM-5.2", max_seq_len=1_048_576)

prompt = "What is the FLOP reduction of Index-Share compared to vanilla sparse attention?"
gen = model.generate(prompt, max_new_tokens=256, temperature=0.7)

print(gen.text)

Both configurations set max_seq_len to 1,048,576 tokens, triggering the IndexShare schedule inside the model's attention implementation. During inference, the system builds an index once per four layers, achieving the reported 2.9× FLOP reduction compared to per-layer index construction.

Summary

  • IndexShare reuses one indexer across four sparse-attention layers, eliminating three expensive index-building operations per layer group.
  • Per-token FLOPs remain constant across shared layers because the same top-k token index is reused, preventing computational growth with depth.
  • GLM-5.2 achieves 2.9× FLOP reduction at 1M-token contexts, as documented in the repository's README.md and confirmed by arXiv 2603.12201.
  • Deployment is supported through inference engines like vLLM and SGLang, with specific guidance available in example/ascend.md for Ascend NPU hardware.

Frequently Asked Questions

What exactly is an "indexer" in GLM-5.2's IndexShare architecture?

An indexer is the data structure constructed during sparse attention that identifies which tokens in the sequence should be attended to. In GLM-5.2, building this indexer requires sorting and bucket-creation operations over the full sequence length, making it the most computationally expensive part of sparse attention. IndexShare treats this indexer as a shared resource across four consecutive layers to avoid recomputing it.

How does IndexShare compare to standard sparse attention in terms of FLOPs?

Standard sparse attention rebuilds the indexer at every layer, resulting in linear growth of index-building operations with model depth. IndexShare reduces these operations by 75% in four-layer blocks (building one index instead of four), which translates to a 2.9× reduction in total per-token FLOPs at 1M-token contexts. The savings become more pronounced as sequence length increases because index construction complexity scales with sequence size.

Can I use GLM-5.2's IndexShare with standard transformers inference libraries?

Yes. The IndexShare mechanism is implemented internally in the GLM-5.2 model weights and architecture, so it works automatically with compatible inference engines. The repository provides specific examples for vLLM and SGLang, as well as deployment guidance for Ascend NPUs in example/ascend.md. As long as the engine supports the GLM-5.2 architecture, the FLOP-reducing IndexShare schedule activates when processing long contexts.

Where is the IndexShare mechanism documented in the GLM-5 repository?

The primary documentation appears in the repository's README.md and README_zh.md files, which specify the 2.9× FLOP reduction and four-layer sharing pattern. Technical context and paper references are located in skills/glm-master-skill/SKILL.md, while deployment-specific instructions for hardware like Ascend NPUs are detailed in example/ascend.md. The original research paper (arXiv 2603.12201) provides the complete theoretical framework.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →