# MTPLX Common Errors and Solutions: Troubleshooting Guide for Multi-Token Prediction Runtime

> Fix MTPLX common errors like MTP head issues or scheduler mismatches. Our troubleshooting guide helps you resolve runtime problems with MTPLX and get your multi-token predictions working smoothly.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-11

---

**MTPLX errors typically stem from missing MTP heads, scheduler mode mismatches, or Metal memory pressure, and can be resolved by verifying model compatibility with `mtplx inspect`, aligning numerics profiles with scheduler configurations, and adjusting context windows to fit unified memory limits.**

MTPLX is a native macOS LLM runtime that accelerates inference using **multi-token prediction (MTP)**. Because it tightly couples model architecture, Metal kernels, and a sophisticated scheduler, several runtime-level errors occur when the configuration, model, or hardware do not match the expectations of the MTP pipeline. This guide explains the most frequent MTPLX common errors and solutions, referencing specific source files from the `youssofal/MTPLX` repository.

## RuntimeError: MTP is not enabled for this runtime

This is the most frequently encountered error when working with MTPLX. The error originates in [`mtplx/runtime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py) at lines 339-342, where a guard checks `self.mtp_enabled` before any draft-MTP operation.

### Why This Happens

The runtime raises this **RuntimeError** whenever asked to perform MTP-draft operations while the loaded model lacks an MTP head or the head failed to load. This guard prevents silent fallbacks to autoregressive (AR) decoding, ensuring exactness in the prediction pipeline.

### Solutions

- **Verify model MTP support**: Use `mtplx inspect <model>` to confirm the model includes an MTP head. Look for `MTP: enabled ✔` in the output.
- **Select correct models**: Choose models explicitly built for MTPLX, such as `mlx-community/Qwen3.8-Optimized-Speed`.
- **Use AR-only mode**: If you deliberately want AR-only behavior, launch with `--no-mtp` or set `scheduler_mode=ar_batch`.

## Scheduler and Numerics Compatibility Errors

### RuntimeError: balanced requires scheduler_mode=mtp_batch

In [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (lines 42-45), validation logic enforces that the *balanced* numerics profile only works with the **mtp_batch** scheduler. This profile relies on a batch-wide verification kernel that exists only in that scheduler lane.

**Fix**: Switch to `--scheduler-mode mtp_batch` when using `--mtp-batch-numerics balanced`. Alternatively, use `b1-exact` which works with any scheduler mode.

## Metal Memory Allocation Failures

When MTPLX encounters `[metal::malloc] Unable to allocate …` errors, the issue originates in [`mtplx/server.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server.py) (lines 30-33), where `_stream_error_kind` maps Metal allocation errors to **memory_refusal** HTTP 507 responses.

### Why Memory Errors Occur

Metal cannot allocate the requested buffer size, usually because the model's KV cache exceeds available unified memory on Apple Silicon.

### Solutions

- Reduce the **context window** using `--context-window`
- Lower **max-active-requests** to decrease concurrent memory pressure
- Use smaller models or lower-precision quantization (4-bit)
- On memory-constrained Macs, enable **hyper** mode to keep width = 1

## Session and Tool-Calling Errors

### RuntimeError: rich tool schemas are unsupported

As implemented in [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py) (lines 4191-4192), the OpenAI-compatible endpoint supports only minimal tool schemas. Richer JSON-Schema definitions trigger this guard.

**Fix**: Stick to minimal tool schemas (name, description, parameters with simple types) as defined in the OpenAI spec.

### RuntimeError: cache filtering failed after partial mutation

This error appears in [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py) (lines 1537-1540) during cache integrity checks. It occurs when a session-bank operation attempts to filter the KV cache but is interrupted, leaving the cache inconsistent.

**Fix**: Avoid mutating a session while generation is in-flight. Use `await session.flush()` before editing. Upgrade to the latest MTPLX version which patches this race condition.

## Model Checkpoint Compatibility Errors

### AttributeError: 'DecoderLayer' object has no attribute 'input_layernorm'

Documented in [`CHANGELOG.md`](https://github.com/youssofal/MTPLX/blob/main/CHANGELOG.md) (lines 148-149), this error indicates the model checkpoint is missing required layers, often caused by loading a pre-MTP checkpoint instead of an MTPLX-ready one.

**Fix**: Pull the latest model release using `mtplx pull …` and ensure you use an **MTPLX-ready** checkpoint rather than a vanilla MLX checkpoint.

## Practical Code Examples

### Detecting Missing MTP Heads Programmatically

```python
import mtplx
from mtplx.runtime import RuntimeError

try:
    # Load a model that *should* have an MTP head

    rt = mtplx.Runtime(model="mlx-community/Qwen3.8-Optimized-Speed")
    # Attempt a draft call – will raise if MTP is unavailable

    rt.draft_mtp(tokens=[0, 1, 2])
except RuntimeError as e:
    if "MTP is not enabled" in str(e):
        raise RuntimeError(
            "Selected model lacks an MTP head. "
            "Pick a model from the MTPLX catalog that includes MTP, "
            "or launch with '--no-mtp' for AR‑only decoding."
        )

```

*Source*: Guard in [`mtplx/runtime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py) lines 339-342.

### Launching with Correct Scheduler and Numerics

```bash

# Correct launch for a balanced numerics profile

mtplx serve \
  --model mlx-community/Qwen3.8-Optimized-Speed \
  --scheduler-mode mtp_batch \
  --mtp-batch-numerics balanced

```

If you omit `--scheduler-mode mtp_batch`, the server aborts with `RuntimeError: balanced requires scheduler_mode=mtp_batch`.

*Source*: Validation in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) lines 42-45.

### Handling Memory Allocation Errors

```python
def start_server():
    try:
        mtplx.serve(port=8000)
    except RuntimeError as exc:
        if "[metal::malloc]" in str(exc):
            print(
                "⚠️  Not enough unified memory for the current context window. "
                "Consider reducing '--context-window' or using a smaller model."
            )
            raise

```

*Source*: Error mapping in [`mtplx/server.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server.py) lines 30-33.

### Verifying MTP Availability via CLI

```bash
$ mtplx inspect mlx-community/Qwen3.8-Optimized-Speed

# Output includes:

#   MTP: enabled ✔

```

If the output shows `MTP: disabled ✘`, switch to a different model from the MTPLX catalog.

## Summary

- **MTP head availability** determines whether the runtime can perform multi-token prediction. Verify with `mtplx inspect` before loading models.
- **Scheduler-numerics coupling** is strict: the `balanced` profile requires `mtp_batch` mode, enforced in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py).
- **Memory pressure** manifests as `[metal::malloc]` errors; resolve by reducing context windows or using lower quantization.
- **Session integrity** requires avoiding mutations during active generation to prevent cache filtering failures.
- **Model compatibility** requires MTPLX-ready checkpoints, not vanilla MLX models, to avoid attribute errors in decoder layers.

## Frequently Asked Questions

### Why does MTPLX say "MTP is not enabled" when I know the model supports it?

This occurs when the model checkpoint lacks the specific MTP head weights required by the MTPLX runtime. According to the source code in [`mtplx/runtime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py), the runtime checks `self.mtp_enabled` before any draft operation. Even if the model architecture supports MTP, the weights must be present in the checkpoint. Use `mtplx inspect <model>` to verify the weights are included, or ensure you are not launching with `--no-mtp` which explicitly disables the feature.

### Can I use the balanced numerics profile with any scheduler mode?

No. The **balanced** numerics profile strictly requires `scheduler_mode=mtp_batch` because it relies on a batch-wide verification kernel only available in that mode. As implemented in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) lines 42-45, the server validates this compatibility at startup and raises `RuntimeError: balanced requires scheduler_mode=mtp_batch` if the configuration is incorrect. Use `b1-exact` if you need a numerics profile compatible with all scheduler modes.

### How do I fix "[metal::malloc] Unable to allocate" errors on my Mac?

This error indicates Metal cannot allocate sufficient unified memory for the KV cache. According to [`mtplx/server.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server.py), these errors map to HTTP 507 responses. Reduce memory pressure by lowering `--context-window` or `--max-active-requests`, switch to a smaller model or 4-bit quantization, or enable **hyper** mode on memory-constrained Apple Silicon devices. The error typically occurs when the requested buffer size exceeds available RAM shared between CPU and GPU.

### What causes cache filtering failures in MTPLX?

The `RuntimeError: cache filtering failed after partial mutation` occurs when a session-bank operation interrupts cache filtering, leaving the KV cache inconsistent. As noted in [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py), this happens when mutating a session while generation is in-flight. Always `await session.flush()` before editing session parameters, and upgrade to the latest MTPLX version which patches this race condition.