# How to Build an MTPLX-Ready MTP Model Using Forge: Complete CLI Guide

> Build an MTPLX-ready MTP model using Forge CLI. This guide covers the three-phase pipeline probing, building, and verification to enable Multi-Task-Prompt Learning deployment.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-02

---

**Forge converts standard Hugging Face checkpoints into MTPLX-ready artifacts through a three-phase pipeline of probing, building, and verification, enabling Multi-Task-Prompt Learning deployment.**

The `youssofal/MTPLX` repository provides **Forge** as its command-line entry point to build an MTPLX-ready MTP model using Forge. This tool automates the transformation of existing models into the MLX-affine format required for MTPLX inference, handling everything from format detection to MTP contract calibration.

## Architecture and Core Components

At the heart of the system lies [`mtplx/commands/forge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py), which implements the `Forge` CLI dispatcher. The architecture delegates heavy computation to external subprocesses—primarily `mlx_lm`—ensuring the main process remains isolated from memory spikes during conversion. Forge operates through three deterministic phases: **Probe**, **Build**, and **Verify**, with an optional **Publish** stage for distribution.

Key supporting modules include [`mtplx/commands/forge_mixed_convert.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge_mixed_convert.py) for per-module quantization overrides and [`mtplx/gemma4_pair.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/gemma4_pair.py) for detecting specialized assistant-pair bundles. Progress reporting writes JSON status files ([`download.json`](https://github.com/youssofal/MTPLX/blob/main/download.json), [`convert.json`](https://github.com/youssofal/MTPLX/blob/main/convert.json), [`verify.json`](https://github.com/youssofal/MTPLX/blob/main/verify.json)) to enable real-time UI monitoring.

## Phase 1: Probing Source Compatibility

Before conversion, Forge inspects the source to determine forgeability and detect the underlying format. The `probe_source()` function in [`mtplx/commands/forge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py) analyzes the checkpoint and returns metadata via `_source_format_from_config`, classifying inputs as BF16, MLX-affine, AutoAWQ, compressed-tensors (AWQ/NVFP4), or Gemma-4 assistant-pair bundles.

**To probe a Hugging Face repository:**

```bash
mtplx forge \
  --forge-action probe \
  --source EleutherAI/gpt-neox-20b

```

This command validates whether the model supports MTPLX conversion and identifies the specific source format without downloading weights, enabling you to verify prerequisites before committing resources.

## Phase 2: Building the MTPLX Artifact

The build phase converts the probed model into an MTPLX-ready artifact. This process executes `_cmd_build()` in [`mtplx/commands/forge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py), which orchestrates several internal steps:

1. `_prepare_source()` downloads the repository or validates the local path
2. `_convert_with_mlx_lm()` (or format-specific variants) converts weights to MLX-affine format
3. `_calibrate_mtp_contract()` discovers feasible multi-task prompt learning depths
4. `_stamp_runtime_metadata()` writes the [`mtplx_runtime.json`](https://github.com/youssofal/MTPLX/blob/main/mtplx_runtime.json) contract file

**The MTP contract**—captured via `_runtime_or_default_mtp_contract()`—defines the depth-wise hidden-variant layout that serves as the inference "contract" for downstream tasks.

**To build with custom quantization:**

```bash
mtplx forge \
  --forge-action build \
  --repo EleutherAI/gpt-neox-20b \
  --branded-name neox-20b-mtplx \
  --out ./outputs \
  --run-id build-001 \
  --recipe '{"body_dtype":"bf16","body_bits":4}'

```

When the recipe contains `module_overrides`, Forge delegates to [`mtplx/commands/forge_mixed_convert.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge_mixed_convert.py) to apply per-module precision predicates, enabling selective quantization of specific layers while maintaining FP16/BF16 elsewhere.

## Phase 3: Verifying the MTP Contract

Verification ensures the built artifact satisfies the MTP contract through empirical testing. The `_run_verify()` function launches `mtplx tune` (or the `_run_verify_family_serve()` fallback) to generate Autoregressive (AR) and MTP depth rows, confirming the calibrated depths are functional.

Progress streams to [`verify.json`](https://github.com/youssofal/MTPLX/blob/main/verify.json) in the output directory, allowing monitoring tools to track verification status in real time.

**To verify a built model:**

```bash
mtplx forge \
  --forge-action verify \
  --path ./outputs/neox-20b-mtplx \
  --out ./verify \
  --run-id verify-001 \
  --max \
  --json

```

This produces a verification document containing performance rows for each MTP depth and confirms the integrity of the [`mtplx_runtime.json`](https://github.com/youssofal/MTPLX/blob/main/mtplx_runtime.json) metadata stamped during the build phase.

## Publishing to Hugging Face (Optional)

After successful verification, distribute the artifact using the publish action. This creates a new Hugging Face repository and uploads the MTPLX-ready files, recording upload manifests in [`publish.json`](https://github.com/youssofal/MTPLX/blob/main/publish.json).

**To publish the verified model:**

```bash
mtplx forge \
  --forge-action publish \
  --source ./outputs/neox-20b-mtplx \
  --repo myorg/neox-20b-mtplx \
  --out ./publish \
  --run-id publish-001

```

## Summary

- **Probe** your source model using `probe_source()` in [`mtplx/commands/forge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py) to validate forgeability and detect formats like AWQ or compressed-tensors before committing resources.
- **Build** the artifact via `_cmd_build()`, which downloads, converts to MLX-affine format, and calibrates the MTP contract through `_calibrate_mtp_contract()`, outputting [`mtplx_runtime.json`](https://github.com/youssofal/MTPLX/blob/main/mtplx_runtime.json).
- **Verify** the contract using `_run_verify()` to generate AR/MTP depth rows and ensure the model meets runtime requirements.
- **Publish** the final artifact to Hugging Face to make the MTPLX-ready model available for `mtplx serve` or downstream workflows.
- **Apply mixed-precision overrides** through [`forge_mixed_convert.py`](https://github.com/youssofal/MTPLX/blob/main/forge_mixed_convert.py) when using recipes containing `module_overrides` for per-layer quantization control.

## Frequently Asked Questions

### What is the difference between probing and building in Forge?

**Probing** inspects the model metadata without downloading weights, using `probe_source()` to determine if the checkpoint is compatible and which converter (AWQ, compressed-tensors, or standard) is required. **Building** actually downloads the weights, converts them to MLX-affine format, and calibrates the MTP contract, producing the [`mtplx_runtime.json`](https://github.com/youssofal/MTPLX/blob/main/mtplx_runtime.json) file required for inference.

### Where is the MTP contract stored and how is it calibrated?

The MTP contract is stored in [`mtplx_runtime.json`](https://github.com/youssofal/MTPLX/blob/main/mtplx_runtime.json) within the output directory. Forge calibrates this contract during the build phase via `_calibrate_mtp_contract()` in [`mtplx/commands/forge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py), which runs a quick tuning process to discover feasible depths and hidden-variant layouts specific to the model architecture.

### Can I use custom quantization recipes during the build process?

Yes. Pass a JSON recipe to the `--recipe` argument containing `body_dtype`, `body_bits`, or `module_overrides`. When `module_overrides` are present, Forge invokes [`mtplx/commands/forge_mixed_convert.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge_mixed_convert.py) to apply per-module quantization predicates, allowing fine-grained control over which layers use 4-bit, 8-bit, or FP16 precision.

### How does Forge handle different model formats like AutoAWQ or Gemma-4?

Forge automatically detects the source format through `_source_format_from_config()` during the probe phase. For AutoAWQ and compressed-tensors formats, it routes to specialized converters like `_convert_compressed_tensors_awq()`. For Gemma-4 assistant-pair bundles, it uses helpers from [`mtplx/gemma4_pair.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/gemma4_pair.py) to handle the unique weight bundling before standard MLX conversion.