How to Use Forge to Build and Verify Custom MTP Models

Forge is the MTPLX command-line front-end that converts standard Hugging Face models or local checkpoints into MTP-enabled artifacts and validates them against the MTPLX runtime contract.

The youssofal/MTPLX repository provides Forge as the primary tool to use Forge to build and verify custom MTP models. It automates the conversion of compatible transformers into the MLX format while calibrating Multi-Token Prediction (MTP) contracts. This guide walks through the three-phase workflow—Probe, Build, and Verify—using actual commands and implementation details from the source code.

The Three-Phase Forge Workflow

Forge operates through three distinct phases, each implemented as a subcommand in [mtplx/commands/forge.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py).

Phase 1: Probe the Source

Before building, determine if your model is compatible. The probe command calls probe_source() at line 80 to detect the source format and existing metadata.

mtplx forge probe --source <HF-repo-id|/path/to/checkpoint> [--json]

The function returns a JSON verdict indicating whether the source is "forgeable":

{
  "verdict": "forgeable",
  "source_format": "mlx_affine",
  "has_mtp_weights": true,
  "message": "MLX affine source detected; Forge will package and verify MTP metadata."
}

If the verdict shows "forgeable", proceed to the build phase.

Phase 2: Build a Custom MTP Model

The build command executes _cmd_build() starting at line 38, orchestrating weight conversion and contract calibration.

mtplx forge build \
    --repo <HF-repo-id|/local/path> \
    --branded-name MyModel-MTPLX \
    --out ./out \
    --run-id build-001 \
    --dtype fp16

Internal execution flow in [mtplx/commands/forge.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py):

  • _body_dtype() (lines 84-90): Validates the recipe and dtype parameters
  • _prepare_source() (lines 94-121): Downloads or locates the source checkpoint
  • _convert_with_mlx_lm() (lines 124-148): Converts weights to MLX format via _mlx_lm_convert_command()
  • _calibrate_mtp_contract() (lines 141-150): Runs the MTP calibration side-car
  • _stamp_runtime_metadata() (lines 171-186): Writes the mtplx_runtime.json contract

The output directory contains the converted checkpoint and mtplx_runtime.json describing the MTP depth and scaling factors.

Phase 3: Verify the Built Model

Verification ensures the model meets runtime performance guarantees. The verify command triggers _cmd_verify() → _run_verify() at line 86.

mtplx forge verify \
    --path ./out/MyModel-MTPLX \
    --out ./verify \
    --run-id verify-001 \
    --stamp

During verification:

  1. probe_source() reloads existing metadata
  2. _run_verify() executes an AR + MTP benchmark using the validator in [mtplx/benchmarks/validators/basic.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/validators/basic.py)
  3. Results are saved to verify/verify.json
  4. With --stamp, the verified contract updates mtplx_runtime.json in-place

Essential Forge Flags

Flag Purpose
--max Enables ThermalForge fan-boost during benchmarking (Apple Silicon optimization)
--allow-degraded-mtp Bypasses the requantize guard (REQUIREMENT_REFUSAL) for edge cases
--json Outputs structured JSON for CI/CD pipelines
--stamp Writes verified metadata back to the model directory

Core Implementation Files

Forge's functionality spans several critical modules:

Summary

Frequently Asked Questions

What input formats does Forge support?

Forge accepts Hugging Face repository IDs, local checkpoint paths, and MLX affine formats. The probe_source() function in [mtplx/commands/forge.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py) automatically detects the format and determines the appropriate conversion lane, including special handling for compressed tensors via _convert_compressed_tensors_awq().

How does Forge calibrate MTP contracts?

During the build phase, _calibrate_mtp_contract() analyzes the converted weights to determine optimal MTP depth and scaling factors. This process writes the mtplx_runtime.json file that the MTPLX runtime uses to configure multi-token prediction behavior. The verification phase later benchmarks these contracts to ensure they meet performance targets.

Can I verify a model without stamping the results?

Yes. Omit the --stamp flag during verification to run benchmarks without modifying the model's mtplx_runtime.json. This is useful for testing experimental builds or comparing performance across different calibration settings without affecting the production artifact.

What happens if the source model is not forgeable?

The probe command returns a verdict of "unforgeable" with a descriptive message indicating the incompatibility reason, such as unsupported quantization schemes or missing architecture metadata. Forge refuses to proceed with the build phase until the source meets the requirements detected by probe_source().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →