How to Cache Expert Paths with IndexCache for GLM-5.2 Inference Optimization

IndexCache reduces GLM-5.2 inference latency by storing frequently used expert routing decisions and static routing tables, eliminating redundant MoE computations for up to 1 million token contexts.

The zai-org/GLM-5 repository implements IndexCache as a lightweight optimization layer for the GLM-5.2 model's Mixture-of-Experts (MoE) architecture. During long-context generation, the model recomputes expert routing paths for every token, creating significant memory traffic and computational overhead. By caching these expert paths and sparse attention indices, IndexCache delivers 2–3× speed-ups on Ascend NPU and comparable hardware back-ends.

Why IndexCache Matters for GLM-5.2 MoE Inference

GLM-5.2 uses a Mixture-of-Experts architecture where each token is dynamically routed to a small subset of expert sub-networks. This routing decision—known as the expert path—requires a fresh softmax and top-k selection for every token during inference.

For very long contexts (up to 1 million tokens), this repeated computation becomes a bottleneck. The example/ascend.md file identifies IndexCache as item 5 in the Ascend deployment optimization checklist, noting that without caching, the routing computation dominates latency in long-context scenarios.

How IndexCache Optimizes Expert Routing

IndexCache employs four complementary mechanisms to minimize computation and memory bandwidth during inference:

High-Frequency Expert Path Caching

When a token's routing pattern has been seen before, IndexCache retrieves cached expert IDs and weights directly from on-device memory. This bypasses the costly MoE softmax calculation and top-k selection for the majority of tokens, while preserving correctness for novel routing patterns.

Static Routing Table Caching

The sparse-attention indexer—the mapping structure that connects tokens to KV-pairs—is cached alongside expert paths. As documented in the repository's README.md, this enables faster lookup during the long-context attention pass without rebuilding the attention matrix for every generation step.

Chunked Prefill Strategy

Prefill (the initial context processing phase) is split into fixed-size chunks. Each chunk reuses the same cached index, dramatically reducing the per-token cost of building the attention matrix. This approach is particularly effective for the 1M token contexts supported by GLM-5.2.

Sparse Index Retrieval

IndexCache fetches only non-zero entries of the attention index from cache. This sparse retrieval pattern cuts memory bandwidth requirements and improves throughput on NPU and GPU devices where memory bandwidth is often the limiting factor.

Enabling IndexCache in Production

The cache is automatically managed by the runtime when enabled through inference configuration flags. Below are implementation examples for three major frameworks supporting GLM-5.2.

vLLM-Ascend Configuration

In vllm, enable IndexCache via the enable_index_cache parameter:

from vllm import LLM, SamplingParams

llm = LLM(
    model="zai-org/GLM-5.2",
    tensor_parallel_size=1,
    dtype="bfloat16",
    enable_index_cache=True,      # Activates IndexCache

    max_seq_len=1_000_000,
)

sampling_params = SamplingParams(
    temperature=0.7,
    top_p=0.9,
    max_tokens=256,
)

prompt = "Write a Python function that computes the nth Fibonacci number."
output = llm.generate(prompts=[prompt], sampling_params=sampling_params)
print(output[0].text)

Source: The Ascend deployment guide in example/ascend.md (lines 13-16) documents this configuration option.

SGLang Integration

For SGLang deployments, set index_cache=True when loading the model:

import sglang as sgl

model = sgl.from_pretrained(
    "zai-org/GLM-5.2",
    device="ascend",          # or "cuda"

    index_cache=True,         # Enables IndexCache

    max_context_len=1_000_000,
)

response = model.chat(
    "Explain the concept of attention mechanisms in transformer models.",
    max_new_tokens=200,
)
print(response)

Source: Referenced in example/ascend.md (lines 17-20) as part of the SGLang Ascend example.

xLLM Setup

xLLM uses the use_index_cache flag for GLM-5.2 optimization:

from xllm import XLLMModel

model = XLLMModel.from_pretrained(
    "zai-org/GLM-5.2",
    backend="ascend",
    use_index_cache=True,      # Enables the cache

    max_sequence_len=1_000_000,
)

output = model.generate(
    "Summarize the recent advances in large language model scaling.",
    max_new_tokens=150,
)
print(output)

Source: Documented in example/ascend.md (lines 21-24) within the xLLM quick-start guide.

Performance Impact and Hardware Considerations

IndexCache achieves 2–3× speed-up for long-context generation on Ascend NPU, with comparable gains available on CUDA and other back-ends. The optimization is particularly effective for:

  • Long-context scenarios (100K+ tokens) where routing computation overhead compounds
  • Batch inference where similar routing patterns recur across sequences
  • NPU deployments where memory bandwidth optimization is critical

The skills/glm-master-skill/SKILL.md file references these caching mechanisms in the context of the OpenAI-compatible API, confirming that IndexCache integrates cleanly with standard inference endpoints.

Summary

  • IndexCache eliminates redundant MoE routing computations by caching expert paths and static routing tables in zai-org/GLM-5.
  • Four mechanisms drive the optimization: high-frequency path caching, static table caching, chunked prefill, and sparse index retrieval.
  • Configuration requires a single flag: enable_index_cache=True (vLLM), index_cache=True (SGLang), or use_index_cache=True (xLLM).
  • Performance gains reach 2–3× for 1M token contexts on Ascend NPU without sacrificing model accuracy.
  • Source references include example/ascend.md (deployment guide), README.md (architecture rationale), and skills/glm-master-skill/SKILL.md (API integration).

Frequently Asked Questions

What is IndexCache in GLM-5.2?

IndexCache is an on-device optimization layer that stores frequently used expert routing results and sparse attention indices. It prevents the model from recomputing softmax-based routing decisions for tokens with previously seen patterns, significantly reducing inference latency in the MoE architecture.

How do I enable IndexCache for GLM-5.2 inference?

Enable IndexCache by setting the appropriate configuration flag in your inference framework: use enable_index_cache=True for vLLM-Ascend, index_cache=True for SGLang, or use_index_cache=True for xLLM. The cache activates automatically once the flag is set, requiring no changes to model weights or input formatting.

Does IndexCache affect model accuracy?

No. IndexCache preserves full model accuracy by computing fresh routing decisions for novel patterns while only caching verified results. The system maintains correctness for infrequent or new routing patterns that fall outside the cached set, ensuring output quality remains identical to uncached inference.

Which hardware platforms support IndexCache?

IndexCache is optimized for Ascend NPU where it provides 2–3× speed-ups, but the mechanism works across GPU (CUDA) and other NPU back-ends. The sparse index retrieval and chunked prefill strategies provide bandwidth and latency benefits on any hardware where memory traffic is a bottleneck.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →