How to Build an MTPLX-Ready MTP Model Using Forge: Complete CLI Guide

Forge converts standard Hugging Face checkpoints into MTPLX-ready artifacts through a three-phase pipeline of probing, building, and verification, enabling Multi-Task-Prompt Learning deployment.

The youssofal/MTPLX repository provides Forge as its command-line entry point to build an MTPLX-ready MTP model using Forge. This tool automates the transformation of existing models into the MLX-affine format required for MTPLX inference, handling everything from format detection to MTP contract calibration.

Architecture and Core Components

At the heart of the system lies mtplx/commands/forge.py, which implements the Forge CLI dispatcher. The architecture delegates heavy computation to external subprocesses—primarily mlx_lm—ensuring the main process remains isolated from memory spikes during conversion. Forge operates through three deterministic phases: Probe, Build, and Verify, with an optional Publish stage for distribution.

Key supporting modules include mtplx/commands/forge_mixed_convert.py for per-module quantization overrides and mtplx/gemma4_pair.py for detecting specialized assistant-pair bundles. Progress reporting writes JSON status files (download.json, convert.json, verify.json) to enable real-time UI monitoring.

Phase 1: Probing Source Compatibility

Before conversion, Forge inspects the source to determine forgeability and detect the underlying format. The probe_source() function in mtplx/commands/forge.py analyzes the checkpoint and returns metadata via _source_format_from_config, classifying inputs as BF16, MLX-affine, AutoAWQ, compressed-tensors (AWQ/NVFP4), or Gemma-4 assistant-pair bundles.

To probe a Hugging Face repository:

mtplx forge \
  --forge-action probe \
  --source EleutherAI/gpt-neox-20b

This command validates whether the model supports MTPLX conversion and identifies the specific source format without downloading weights, enabling you to verify prerequisites before committing resources.

Phase 2: Building the MTPLX Artifact

The build phase converts the probed model into an MTPLX-ready artifact. This process executes _cmd_build() in mtplx/commands/forge.py, which orchestrates several internal steps:

  1. _prepare_source() downloads the repository or validates the local path
  2. _convert_with_mlx_lm() (or format-specific variants) converts weights to MLX-affine format
  3. _calibrate_mtp_contract() discovers feasible multi-task prompt learning depths
  4. _stamp_runtime_metadata() writes the mtplx_runtime.json contract file

The MTP contract—captured via _runtime_or_default_mtp_contract()—defines the depth-wise hidden-variant layout that serves as the inference "contract" for downstream tasks.

To build with custom quantization:

mtplx forge \
  --forge-action build \
  --repo EleutherAI/gpt-neox-20b \
  --branded-name neox-20b-mtplx \
  --out ./outputs \
  --run-id build-001 \
  --recipe '{"body_dtype":"bf16","body_bits":4}'

When the recipe contains module_overrides, Forge delegates to mtplx/commands/forge_mixed_convert.py to apply per-module precision predicates, enabling selective quantization of specific layers while maintaining FP16/BF16 elsewhere.

Phase 3: Verifying the MTP Contract

Verification ensures the built artifact satisfies the MTP contract through empirical testing. The _run_verify() function launches mtplx tune (or the _run_verify_family_serve() fallback) to generate Autoregressive (AR) and MTP depth rows, confirming the calibrated depths are functional.

Progress streams to verify.json in the output directory, allowing monitoring tools to track verification status in real time.

To verify a built model:

mtplx forge \
  --forge-action verify \
  --path ./outputs/neox-20b-mtplx \
  --out ./verify \
  --run-id verify-001 \
  --max \
  --json

This produces a verification document containing performance rows for each MTP depth and confirms the integrity of the mtplx_runtime.json metadata stamped during the build phase.

Publishing to Hugging Face (Optional)

After successful verification, distribute the artifact using the publish action. This creates a new Hugging Face repository and uploads the MTPLX-ready files, recording upload manifests in publish.json.

To publish the verified model:

mtplx forge \
  --forge-action publish \
  --source ./outputs/neox-20b-mtplx \
  --repo myorg/neox-20b-mtplx \
  --out ./publish \
  --run-id publish-001

Summary

  • Probe your source model using probe_source() in mtplx/commands/forge.py to validate forgeability and detect formats like AWQ or compressed-tensors before committing resources.
  • Build the artifact via _cmd_build(), which downloads, converts to MLX-affine format, and calibrates the MTP contract through _calibrate_mtp_contract(), outputting mtplx_runtime.json.
  • Verify the contract using _run_verify() to generate AR/MTP depth rows and ensure the model meets runtime requirements.
  • Publish the final artifact to Hugging Face to make the MTPLX-ready model available for mtplx serve or downstream workflows.
  • Apply mixed-precision overrides through forge_mixed_convert.py when using recipes containing module_overrides for per-layer quantization control.

Frequently Asked Questions

What is the difference between probing and building in Forge?

Probing inspects the model metadata without downloading weights, using probe_source() to determine if the checkpoint is compatible and which converter (AWQ, compressed-tensors, or standard) is required. Building actually downloads the weights, converts them to MLX-affine format, and calibrates the MTP contract, producing the mtplx_runtime.json file required for inference.

Where is the MTP contract stored and how is it calibrated?

The MTP contract is stored in mtplx_runtime.json within the output directory. Forge calibrates this contract during the build phase via _calibrate_mtp_contract() in mtplx/commands/forge.py, which runs a quick tuning process to discover feasible depths and hidden-variant layouts specific to the model architecture.

Can I use custom quantization recipes during the build process?

Yes. Pass a JSON recipe to the --recipe argument containing body_dtype, body_bits, or module_overrides. When module_overrides are present, Forge invokes mtplx/commands/forge_mixed_convert.py to apply per-module quantization predicates, allowing fine-grained control over which layers use 4-bit, 8-bit, or FP16 precision.

How does Forge handle different model formats like AutoAWQ or Gemma-4?

Forge automatically detects the source format through _source_format_from_config() during the probe phase. For AutoAWQ and compressed-tensors formats, it routes to specialized converters like _convert_compressed_tensors_awq(). For Gemma-4 assistant-pair bundles, it uses helpers from mtplx/gemma4_pair.py to handle the unique weight bundling before standard MLX conversion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →