How to Attach an MTP Sidecar to an MLX Trunk with Forge
Yes—Forge can attach a separately supplied MTP sidecar to any MLX trunk, provided the trunk is wrapped by the MTPLX runtime and the sidecar follows the expected layout conventions.
The MTPLX framework (github.com/youssofal/MTPLX) extends MLX with a modular sidecar architecture that allows quantized MTP (Model Transmission Protocol) weights to be injected into existing models at packaging time. Forge, the built-in conversion tool, automates this integration by discovering sidecar files and embedding them into the model’s runtime contract.
How Forge Discovers and Loads MTP Sidecars
Forge implements sidecar detection through a filesystem scan that looks for mtp/weights.safetensors inside the model directory. When found, the file is registered with the MTPContract, a configuration object that governs quantization settings and runtime behavior.
In mtplx/hf_loader.py, the _mtp_sidecar_candidates function (lines 473-527) validates candidate files and returns paths to eligible sidecars. Forge then copies the discovered weights into the final artifact and updates the runtime manifest. The contract itself is defined in mtplx/mtp_patch.py (lines 373-379), where sidecar detection and repair hooks ensure the sidecar metadata aligns with the trunk’s architecture.
The Forge driver in mtplx/commands/forge.py constructs the MTPContract during the conversion pipeline. If the sidecar is present but fails validation—for example, due to mismatched tensor names or missing metadata—Forge raises a ForgeError and aborts the build.
Prerequisites for Attaching an External Sidecar
Before attaching a sidecar, verify three specific conditions:
- MTPLX Runtime Wrapper: The trunk must be instantiated through
MTPLXRuntime, not a rawmlx.nn.Module. Pure MLX trunks lack theMTPContractinterface required for sidecar injection. - Correct File Location: The sidecar must be a valid Safetensors file placed at
mtp/weights.safetensorsrelative to the model root directory. - Matching Quantization Parameters: The sidecar’s quantization bits, group size, and mode must exactly match the parameters declared in the
MTPContract. Mismatches trigger a verification error during the Forge step.
Step-by-Step: Attaching a Separate MTP Sidecar
You can attach a sidecar programmatically via the Python API or through the MTPLX CLI.
Python API Approach
Use the forge function from mtplx.commands.forge after configuring the runtime contract:
from mtplx.commands.forge import forge
from mtplx.mtp_patch import MTPContract
from mtplx.runtime import MTPLXRuntime
import shutil
import pathlib
# 1. Initialize the runtime with MTP enabled
runtime = MTPLXRuntime(
contract=MTPContract(
mtp_enabled=True,
mtp_quant_bits=4,
mtp_quant_group_size=32
),
)
# 2. Stage the sidecar in the expected location
model_dir = pathlib.Path("/path/to/my_model")
sidecar_src = pathlib.Path("/path/to/my_mtp_sidecar.safetensors")
(model_dir / "mtp").mkdir(exist_ok=True)
shutil.copy(sidecar_src, model_dir / "mtp" / "weights.safetensors")
# 3. Execute Forge to package the model with the sidecar
forge(
model_dir=model_dir,
runtime=runtime,
output_dir=pathlib.Path("/tmp/forge_artifact"),
)
During execution, Forge copies mtp/weights.safetensors into the artifact and writes an mtplx_runtime.json containing the contract. At inference time, mtplx/runtime.py checks MTPContract.mtp_enabled and loads the sidecar weights instead of the raw FP16/FP32 weights.
CLI Approach
The MTPLX command-line interface provides equivalent functionality:
mtplx forge \
--model-dir /path/to/my_model \
--enable-mtp \
--mtp-bits 4 \
--mtp-group-size 32 \
--output /tmp/forge_artifact
Ensure the sidecar exists at /path/to/my_model/mtp/weights.safetensors before running the command. The CLI automatically validates the file and injects it into the build.
When Sidecar Attachment Does Not Work
Attachment fails in two primary scenarios:
Pure MLX Trunks Without MTPLX Wrapping
If the model is a standard MLX module never passed through MTPLXRuntime, Forge cannot inject the sidecar because the MTPContract object does not exist. You must first wrap the trunk using the MTPLX runtime before invoking Forge.
Incompatible Sidecar Formats
Sidecars with incorrect tensor schemas, missing required metadata, or quantization parameters that diverge from the contract are rejected during the verification phase. Forge raises ForgeError: MTP did not accelerate this model when validation fails, preventing runtime crashes from mismatched weight shapes.
Summary
- Forge automatically attaches MTP sidecars when it finds
mtp/weights.safetensorsin the model directory. - The attachment requires an MTPContract configured via
mtplx/mtp_patch.pyand processed bymtplx/commands/forge.py. - The trunk must be wrapped by MTPLXRuntime; raw MLX modules are not compatible.
- Quantization parameters in the sidecar must match the contract exactly, verified by logic in
mtplx/hf_loader.py. - Runtime loading occurs in
mtplx/runtime.py, which swaps in sidecar weights whenmtp_enabledis true.
Frequently Asked Questions
Can I attach an MTP sidecar to any Hugging Face model converted to MLX?
No—you can only attach sidecars to models that have been wrapped by the MTPLX runtime. Standard MLX conversions lack the MTPContract interface that Forge requires to register and validate sidecar weights. You must first load the model through MTPLXRuntime before running Forge.
What file format must the MTP sidecar use?
The sidecar must be a Safetensors file named weights.safetensors and placed in a subdirectory named mtp/ inside the model folder. Forge specifically searches for this path pattern in mtplx/hf_loader.py (lines 473-527) and rejects other formats or locations.
Does the sidecar量化 [quantization] settings need to match the base model?
Yes—the quantization bits, group size, and mode declared in the MTPContract must exactly match the sidecar’s internal metadata. Forge validates these parameters during the conversion pipeline in mtplx/commands/forge.py. A mismatch triggers a ForgeError and stops the build process.
How does the runtime know to use the sidecar instead of base weights?
At inference time, mtplx/runtime.py checks the mtp_enabled field in the loaded MTPContract. If true, the runtime loads weights from the sidecar path (packaged within the Forge artifact) rather than the original FP16 or FP32 parameters. This switching logic is transparent to the downstream inference code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →