# Can You Attach a Separately Supplied MTP Sidecar to an MLX Trunk in MTPLX?

> Learn why you cannot attach a separately supplied MTP sidecar to an MLX trunk in MTPLX. Discover the architectural reasons for this restriction in the youssofal/MTPLX repository.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-04

---

**No, MTPLX explicitly prohibits grafting a user-supplied MTP sidecar onto an arbitrary MLX trunk.** The runtime architecture requires that the multi-token prediction head be built, trained, and cryptographically verified together with the specific trunk weights it will accelerate.

The MTPLX inference engine (youssofal/MTPLX) enforces strict coupling between transformer trunk weights and their corresponding MTP (Multi-Token Prediction) sidecars. This design constraint exists to preserve the mathematical exactness guarantees of speculative sampling (Leviathan & Chen rejection sampling), ensuring that draft tokens generated by the sidecar are valid for the exact trunk weights loaded into memory.

## Why MTPLX Blocks External MTP Sidecar Attachment

### The Exactness Guarantee Requirement

MTPLX relies on mathematically exact speculative sampling to maintain correctness during accelerated decoding. Because the MTP head must generate draft tokens that are valid for the specific probability distribution of the trunk, it **must be trained on the exact same trunk weights** to preserve this guarantee. 

According to the source documentation in [`README.md`](https://github.com/youssofal/MTPLX/blob/main/README.md) (lines 71‑73), matching architecture fields, tensor shapes, or provenance labels cannot prove that an external head was trained against those exact trunk weights. A mismatched sidecar would break the exactness of the draft-verification loop, potentially causing silent errors or degraded output quality.

### The Verification Pipeline

When you attempt to load a model, MTPLX runs a strict inspection protocol defined in [`mtplx/engine_session.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py) (lines 180‑185). This verification checks that the sidecar’s tensor shapes, projection matrices, and provenance tags **match** the trunk. If any discrepancy is detected, the runtime raises a `ValueError` before any weights are loaded, preventing runtime corruption.

The sidecar is expected to reside at `mtp/weights.safetensors` relative to the model root, as defined in [`mtplx/artifacts.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/artifacts.py) (lines 387‑390). The runtime hardcodes this path and expects the file to be bundled during the Forge packaging process, not supplied separately at load time.

## How to Properly Load or Create MTP-Enabled Models

Since you cannot attach a separately supplied MTP sidecar to an MLX trunk after the fact, you must use one of two supported workflows:

### Loading a Complete Pre-Bundled Artifact

Use a model that already contains its compatible MTP tensors packaged at build time. The MTPLX CLI handles the download and automatic sidecar loading:

```bash
mtplx start --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed

```

This command loads the trunk and the bundled `mtp/weights.safetensors` sidecar, immediately enabling multi-token prediction without manual grafting.

### Verifying Sidecar Presence Before Loading

Inspect any MTPLX model to confirm it contains a verified sidecar before attempting to start the engine:

```bash
mtplx inspect Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed --json

```

If the JSON output contains `"mtp_sidecar": true`, the artifact passed the verification pipeline and the sidecar is proven compatible with the trunk.

### Building a Verified Sidecar with Forge

If you have a raw Hugging Face checkpoint without an MTPLX sidecar, you must use **Forge** to generate a proper artifact. The Forge command (implemented in [`mtplx/commands/forge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py), lines 1524‑1624) automates the conversion:

```bash
mtplx forge build \
    --repo huggingface/your-model \
    --output ./my_mtplx_model \
    --quant int8 \
    --depth 1

```

Forge performs four critical steps:
1. Converts the model to MLX format.
2. Trains (or extracts) the MTP adapter on the **same trunk**.
3. Verifies that the sidecar actually speeds up decoding on sample data.
4. Packages the sidecar alongside the trunk in `mtp/weights.safetensors`.

Because the sidecar is generated **in-process** with the exact trunk, no external grafting is needed nor supported.

## Attempting Manual Sidecar Grafting (Will Fail)

If you attempt to programmatically load a detached sidecar using the Python API, MTPLX will reject the operation. The following code demonstrates the unsupported pattern:

```python
import mlx.core as mx
from mtplx import MTPLXRuntime, MTPContract

# Load trunk only

runtime = MTPLXRuntime(model_path="path/to/trunk")

# Attempt to manually attach a sidecar – this is NOT supported

runtime.load_mtp_sidecar("path/to/your_sidecar.safetensors")  # ❌

```

The runtime will raise an error similar to:

```

ValueError: MTP sidecar payload does not match the loaded trunk weights.

```

This enforcement occurs in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) (lines 3531‑3538) during the session initialization, ensuring that only bundled, verified sidecars may be attached to the engine.

## Summary

- **MTPLX does not support** attaching a separately supplied MTP sidecar to an arbitrary MLX trunk.
- **Exactness guarantee** requires the sidecar to be trained on and cryptographically bound to the specific trunk weights.
- **Use complete models** that already include `mtp/weights.safetensors` for immediate multi-token prediction.
- **Use Forge** (`mtplx forge build`) to generate a valid MTPLX artifact from a Hugging Face checkpoint, which automatically creates and verifies the sidecar.
- **Manual grafting attempts** via the Python API fail with a `ValueError` during the verification pipeline defined in [`mtplx/engine_session.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py).

## Frequently Asked Questions

### Can I mix an MTP sidecar from one model with a different trunk if the architectures match?

No. Even if architecture fields and tensor shapes align, MTPLX rejects the load because matching metadata cannot prove the head was trained against those exact trunk weights. The runtime enforces this in [`mtplx/engine_session.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py) to prevent breaking the exactness of the speculative sampling loop. You must use a sidecar generated by Forge from the same source checkpoint as the trunk.

### How do I verify if an MTPLX model has a valid MTP sidecar?

Run the `mtplx inspect` command with the `--json` flag. Validated models return `"mtp_sidecar": true` in their metadata, indicating the sidecar passed shape, projection matrix, and provenance checks. This inspection occurs before the engine attempts to load weights, preventing runtime errors.

### What is the Forge command to create a compatible sidecar?

Execute `mtplx forge build --repo <huggingface-id> --output <path>`. Forge automatically trains the MTP adapter on the trunk, verifies it improves inference speed, and packages it at `mtp/weights.safetensors`. This is the only supported method for creating a sidecar that the runtime will accept.

### Does MTPLX support third-party or community-created MTP sidecars?

No. The runtime only accepts sidecars that were generated during the Forge build process for the specific trunk in question. Community sidecars created separately cannot satisfy the verification checks in [`mtplx/artifacts.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/artifacts.py) because they lack the in-process provenance tags that prove co-training with the trunk weights.