# How the MTP Layer Improves Speculative Decoding in GLM-5

> Discover how GLM-5's MTP layer boosts speculative decoding, extending acceptance by 20% and cutting FLOPs by 2.9× with efficient sparse-attention. Learn more!

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: internals
- Published: 2026-06-19

---

**The Multi-Token Prediction (MTP) layer in GLM-5 enables the model to predict and verify multiple tokens in a single forward pass, extending speculative acceptance length by approximately 20% and reducing per-token FLOPs by roughly 2.9× through efficient sparse-attention reuse.**

The GLM-5 2 release from the zai-org/GLM-5 repository introduces a sophisticated Multi-Token Prediction (MTP) layer that fundamentally changes how speculative decoding operates. By moving from single-token verification to batch verification of draft sequences, the MTP layer reduces latency and increases throughput for long-context generation tasks.

## Understanding the MTP Layer Architecture

### Multi-Token Prediction Mechanism

The MTP layer deviates from traditional autoregressive generation by predicting **multiple tokens in a single forward pass**. In the GLM-5 2 architecture, instead of generating one token at a time and checking its validity, the MTP layer emits a short sequence of future tokens simultaneously. This capability is documented in the repository's [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) at line 25, which details how the layer integrates with the model's sparse attention mechanisms.

### Integration with Speculative Decoding

During speculative decoding, a fast draft model first generates candidate tokens. The MTP layer then **verifies those tokens in a batch** rather than individually. According to the source code in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) at line 7, this "Multi-Token Prediction Optimization" serves as a core improvement for single-step generation, allowing the main model to accept longer streaks of draft tokens before requiring regeneration.

## Performance Improvements from MTP-Enhanced Speculative Decoding

### Extended Acceptance Length

The MTP layer increases the **speculative acceptance length** by up to approximately 20% compared to standard speculative decoding. This means the model can validate longer sequences of draft tokens in a single verification step. As noted in the [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md) architecture description, this extended acceptance reduces the number of costly verification passes required during generation.

### FLOP Efficiency via IndexShare

The MTP layer achieves significant computational savings through the **IndexShare mechanism**, which reuses the same sparse-attention indexer across several layers. At a 1 million-token context length, this reuse cuts per-token FLOPs by approximately **2.9×**. The efficiency gains come from avoiding redundant attention computations during the multi-token verification process.

### Latency and Throughput Benefits

By verifying tokens in batches, the MTP layer delivers three primary performance advantages:

- **Higher throughput** – Fewer forward passes are required per generated token.
- **Reduced latency** – The acceptance length of speculative chunks grows, allowing the model to retain draft outputs longer before regeneration.
- **Better FLOP efficiency** – The IndexShare mechanism minimizes redundant computation across the context window.

## Implementation Details in the GLM-5 Source Code

The MTP layer's implementation spans several key files in the zai-org/GLM-5 repository:

- **[`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md)** (line 25): Documents the MTP layer's role in speculative decoding and the ~20% acceptance-length gain.
- **[`README_zh.md`](https://github.com/zai-org/GLM-5/blob/main/README_zh.md)** (line 25): Provides the Chinese language version of the architectural specifications.
- **[`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md)** (line 7): Lists "Multi-Token Prediction Optimization" as a core improvement for single-step generation and attention preprocessing.
- **[`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md)**: Defines the skill configuration used by the GLM-5 API, where the MTP layer is integrated into the inference pipeline exposed to end-users.

## Enabling MTP-Based Speculative Decoding: Code Examples

You can activate the MTP layer through the Hugging Face Transformers API or vLLM. In both cases, the MTP layer engages automatically when speculative decoding is enabled, handling multi-token predictions and extending the accepted chunk length.

### Hugging Face Transformers Implementation

```python

# Example: using GLM‑5.2 with speculative decoding in the 🤗‑Transformers API

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "zai-org/GLM-5.2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    # Enable speculative decoding; the library internally activates the MTP layer

    trust_remote_code=True,
    device_map="auto",
    # The length of the speculative chunk the MTP layer will try to accept

    # (default is 4; increasing it may raise acceptance up to ~20 %)

    speculative_accept_len=8,
)

prompt = "Explain why the sky is blue."
input_ids = tokenizer(prompt, return_tensors="pt").input_ids.to(model.device)

# Generate with speculative decoding (MTP enabled)

output_ids = model.generate(
    input_ids,
    max_new_tokens=128,
    do_sample=False,
    # Turn on the speculative pipeline

    use_speculative=True,
)

print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

```

### vLLM Server Implementation

```python

# Example: vLLM inference with speculative decoding (MTP automatically used)

from vllm import LLM, SamplingParams

llm = LLM(
    model="zai-org/GLM-5.2",
    enable_speculative=True,          # activates the MTP layer

    speculative_accept_len=6,        # longer acceptance → higher speed

)

sampling_params = SamplingParams(
    max_tokens=150,
    temperature=0.0,
)

prompt = "Write a short Python function that computes the Fibonacci sequence."
outputs = llm.generate([prompt], sampling_params)

print(outputs[0].text)

```

## Summary

- The **MTP layer** in GLM-5 predicts multiple tokens in a single forward pass, enabling batch verification during speculative decoding.
- Acceptance length increases by approximately **20%**, reducing the number of verification steps required.
- The **IndexShare mechanism** reuses sparse-attention indexers across layers, cutting per-token FLOPs by roughly **2.9×** at 1 million-token contexts.
- Implementation spans [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md), [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md), and [`skills/glm-master-skill/SKILL.md`](https://github.com/zai-org/GLM-5/blob/main/skills/glm-master-skill/SKILL.md) in the zai-org/GLM-5 repository.
- Both Hugging Face Transformers and vLLM support MTP-based speculative decoding via `use_speculative=True` and `enable_speculative=True` parameters.

## Frequently Asked Questions

### What is the MTP layer in GLM-5?

The **Multi-Token Prediction (MTP)** layer is an architectural component in GLM-5 2 that predicts multiple future tokens simultaneously rather than generating them one at a time. According to the source code in [`README.md`](https://github.com/zai-org/GLM-5/blob/main/README.md), this layer enables the model to verify longer sequences of draft tokens during speculative decoding, which significantly reduces generation latency.

### How does the MTP layer improve speculative decoding speed?

The MTP layer improves speed by extending the **speculative acceptance length** by approximately 20%, allowing the model to accept longer chunks of draft tokens in a single verification pass. As implemented in the [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) documentation, this batch verification approach reduces the number of forward passes required per generated token, directly translating to higher throughput and lower latency.

### What is the IndexShare mechanism and how does it reduce FLOPs?

The **IndexShare mechanism** is an optimization within the MTP layer that reuses the same sparse-attention indexer across multiple layers during multi-token prediction. This reuse eliminates redundant attention computations, reducing per-token FLOPs by roughly **2.9×** when processing contexts up to 1 million tokens, as detailed in the GLM-5 architecture documentation.

### How do I enable MTP-based speculative decoding in my inference pipeline?

To enable the MTP layer, set `use_speculative=True` in the Hugging Face Transformers `generate()` method or `enable_speculative=True` when initializing the vLLM engine. The `speculative_accept_len` parameter controls the maximum chunk size the MTP layer will attempt to verify, with higher values potentially increasing acceptance rates up to the ~20% limit mentioned in the source code.