Can You Attach a Separately Supplied MTP Sidecar to an MLX Trunk in MTPLX?
No, MTPLX explicitly prohibits grafting a user-supplied MTP sidecar onto an arbitrary MLX trunk. The runtime architecture requires that the multi-token prediction head be built, trained, and cryptographically verified together with the specific trunk weights it will accelerate.
The MTPLX inference engine (youssofal/MTPLX) enforces strict coupling between transformer trunk weights and their corresponding MTP (Multi-Token Prediction) sidecars. This design constraint exists to preserve the mathematical exactness guarantees of speculative sampling (Leviathan & Chen rejection sampling), ensuring that draft tokens generated by the sidecar are valid for the exact trunk weights loaded into memory.
Why MTPLX Blocks External MTP Sidecar Attachment
The Exactness Guarantee Requirement
MTPLX relies on mathematically exact speculative sampling to maintain correctness during accelerated decoding. Because the MTP head must generate draft tokens that are valid for the specific probability distribution of the trunk, it must be trained on the exact same trunk weights to preserve this guarantee.
According to the source documentation in README.md (lines 71‑73), matching architecture fields, tensor shapes, or provenance labels cannot prove that an external head was trained against those exact trunk weights. A mismatched sidecar would break the exactness of the draft-verification loop, potentially causing silent errors or degraded output quality.
The Verification Pipeline
When you attempt to load a model, MTPLX runs a strict inspection protocol defined in mtplx/engine_session.py (lines 180‑185). This verification checks that the sidecar’s tensor shapes, projection matrices, and provenance tags match the trunk. If any discrepancy is detected, the runtime raises a ValueError before any weights are loaded, preventing runtime corruption.
The sidecar is expected to reside at mtp/weights.safetensors relative to the model root, as defined in mtplx/artifacts.py (lines 387‑390). The runtime hardcodes this path and expects the file to be bundled during the Forge packaging process, not supplied separately at load time.
How to Properly Load or Create MTP-Enabled Models
Since you cannot attach a separately supplied MTP sidecar to an MLX trunk after the fact, you must use one of two supported workflows:
Loading a Complete Pre-Bundled Artifact
Use a model that already contains its compatible MTP tensors packaged at build time. The MTPLX CLI handles the download and automatic sidecar loading:
mtplx start --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed
This command loads the trunk and the bundled mtp/weights.safetensors sidecar, immediately enabling multi-token prediction without manual grafting.
Verifying Sidecar Presence Before Loading
Inspect any MTPLX model to confirm it contains a verified sidecar before attempting to start the engine:
mtplx inspect Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed --json
If the JSON output contains "mtp_sidecar": true, the artifact passed the verification pipeline and the sidecar is proven compatible with the trunk.
Building a Verified Sidecar with Forge
If you have a raw Hugging Face checkpoint without an MTPLX sidecar, you must use Forge to generate a proper artifact. The Forge command (implemented in mtplx/commands/forge.py, lines 1524‑1624) automates the conversion:
mtplx forge build \
--repo huggingface/your-model \
--output ./my_mtplx_model \
--quant int8 \
--depth 1
Forge performs four critical steps:
- Converts the model to MLX format.
- Trains (or extracts) the MTP adapter on the same trunk.
- Verifies that the sidecar actually speeds up decoding on sample data.
- Packages the sidecar alongside the trunk in
mtp/weights.safetensors.
Because the sidecar is generated in-process with the exact trunk, no external grafting is needed nor supported.
Attempting Manual Sidecar Grafting (Will Fail)
If you attempt to programmatically load a detached sidecar using the Python API, MTPLX will reject the operation. The following code demonstrates the unsupported pattern:
import mlx.core as mx
from mtplx import MTPLXRuntime, MTPContract
# Load trunk only
runtime = MTPLXRuntime(model_path="path/to/trunk")
# Attempt to manually attach a sidecar – this is NOT supported
runtime.load_mtp_sidecar("path/to/your_sidecar.safetensors") # ❌
The runtime will raise an error similar to:
ValueError: MTP sidecar payload does not match the loaded trunk weights.
This enforcement occurs in mtplx/cli.py (lines 3531‑3538) during the session initialization, ensuring that only bundled, verified sidecars may be attached to the engine.
Summary
- MTPLX does not support attaching a separately supplied MTP sidecar to an arbitrary MLX trunk.
- Exactness guarantee requires the sidecar to be trained on and cryptographically bound to the specific trunk weights.
- Use complete models that already include
mtp/weights.safetensorsfor immediate multi-token prediction. - Use Forge (
mtplx forge build) to generate a valid MTPLX artifact from a Hugging Face checkpoint, which automatically creates and verifies the sidecar. - Manual grafting attempts via the Python API fail with a
ValueErrorduring the verification pipeline defined inmtplx/engine_session.py.
Frequently Asked Questions
Can I mix an MTP sidecar from one model with a different trunk if the architectures match?
No. Even if architecture fields and tensor shapes align, MTPLX rejects the load because matching metadata cannot prove the head was trained against those exact trunk weights. The runtime enforces this in mtplx/engine_session.py to prevent breaking the exactness of the speculative sampling loop. You must use a sidecar generated by Forge from the same source checkpoint as the trunk.
How do I verify if an MTPLX model has a valid MTP sidecar?
Run the mtplx inspect command with the --json flag. Validated models return "mtp_sidecar": true in their metadata, indicating the sidecar passed shape, projection matrix, and provenance checks. This inspection occurs before the engine attempts to load weights, preventing runtime errors.
What is the Forge command to create a compatible sidecar?
Execute mtplx forge build --repo <huggingface-id> --output <path>. Forge automatically trains the MTP adapter on the trunk, verifies it improves inference speed, and packages it at mtp/weights.safetensors. This is the only supported method for creating a sidecar that the runtime will accept.
Does MTPLX support third-party or community-created MTP sidecars?
No. The runtime only accepts sidecars that were generated during the Forge build process for the specific trunk in question. Community sidecars created separately cannot satisfy the verification checks in mtplx/artifacts.py because they lack the in-process provenance tags that prove co-training with the trunk weights.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →