How to Verify a Model with MTPLX Forge: Step-by-Step CLI and API Guide

MTPLX Forge provides a dedicated verify subcommand that measures arithmetic-reuse (AR) and Multi-Tier-Precision (MTP) performance by executing either the standard tune lane or the specialized family-serve lane for native backend models.

Verifying converted models in the youssofal/MTPLX repository ensures that arithmetic-reuse optimizations and MTP contracts perform correctly under load. To verify a model with MTPLX Forge, you invoke the forge verify command, which orchestrates subprocesses, merges partial results in real-time, and emits structured performance data to verify.json.

Verification Pipeline Overview

The verification flow implemented in mtplx/commands/forge.py proceeds through eight distinct stages:

  1. Argument Dispatch: The CLI entry point cmd_forge_public routes to _cmd_verify (line 84) when forge verify is invoked.
  2. Output Setup: The helper _run_dir creates a timestamped folder under outputs/forge-verify unless --out and --run-id specify a custom path.
  3. Metadata Loading: _read_runtime ingests existing mtplx_runtime.json to reuse previously calibrated MTP contracts.
  4. Engine Selection: _run_verify (line 2504) branches into either the tune lane or the family-serve lane for models requiring native backends like Qwen 4-exp.
  5. Tune Lane Execution: For standard models, Forge spawns mtplx tune as a subprocess with --max-tokens, the selected prompt suite, and MTP contract options (--base-hidden-variant, --mtp-hidden-variant, --concat-order). During execution, _merge_candidate_rows continuously aggregates partial results into a live verify.json.
  6. Family-Serve Lane: For native backend models, _run_verify_family_serve (line 2617) boots a temporary mtplx serve instance on a random port, optionally wrapping requests with SmartFanController when --max is set, then issues first-load requests to gather AR and family-default rows.
  7. Post-Processing: _annotate_verify_rows enriches raw data before _stamp_runtime_metadata writes results back to mtplx_runtime.json if --stamp is provided.
  8. Output Emission: By default, pretty-printed JSON goes to stdout; --json emits machine-readable JSON. Exit codes are 0 (success), 1 (no rows gathered), or the subprocess error code.

Essential CLI Flags

Flag Description
--path <dir> Path to the model directory containing mtplx_runtime.json or raw model files.
--out <dir> Destination directory for verification artifacts.
--run-id <id> Identifier grouping a verification run for CI pipelines.
--max Enables MAX-fan mode via SmartFanController, engaging the ThermalForge daemon.
--max-tokens <N> Upper bound on tokens per request (default: 2048).
--suite <name> Prompt suite driving verification (default: long-code-uncapped).
--stamp Updates mtplx_runtime.json to mark the model as verified.
--json Outputs pure JSON to stdout for scripting.

The Two Verification Lanes

MTPLX Forge utilizes two distinct verification paths depending on model requirements.

The Tune Lane

Most models traverse the tune lane, which leverages existing MTPLX tuning infrastructure. This lane executes mtplx tune as a subprocess, passing MTP contract parameters such as --base-hidden-variant and --concat-order. While the subprocess runs, Forge continuously merges partial rows and writes incremental updates to verify.json, enabling downstream tools to monitor progress without waiting for completion.

The Family-Serve Lane

Models requiring native backends (specifically Qwen 4-exp) are routed through the family-serve lane. This path invokes _run_verify_family_serve to boot a temporary mtplx serve instance on a random port. When --max is specified, the system initializes SmartFanController from mtplx/thermal.py to communicate with the privileged ThermalForge daemon, ramping fans to test MAX performance mode. The lane issues first-load requests to collect AR metrics and family-default rows without requiring a full tuning cycle.

Step-by-Step Usage Examples

Basic Model Verification

Verify a local model directory with MAX-fan enabled and machine-readable output:

mtplx forge verify \
    --path ./my-model \
    --out ./verify-run \
    --run-id nightly-2026-09-02 \
    --max \
    --suite long-code-uncapped \
    --json > verify-run/result.json
  • --path targets the model directory.
  • --out and --run-id create ./verify-run/nightly-2026-09-02 for artifacts.
  • --max forces ThermalForge fan-boost mode.
  • --json formats output for programmatic consumption.

Programmatic Verification via Python API

For automated pipelines, invoke _cmd_verify directly:

from pathlib import Path
from mtplx.commands.forge import _cmd_verify, ForgeError

args = type("Args", (), {
    "path": Path("./my-model"),
    "out": Path("./verify-run"),
    "run_id": "py-run-001",
    "max": True,
    "suite": "long-code-uncapped",
    "json": True,
    "stamp": False,
})

try:
    exit_code = _cmd_verify(args)
    print(f"Verification finished with exit code {exit_code}")
except ForgeError as e:
    print(f"Forge failed: {e}")

The function returns 0 on success, 1 if no verification rows were produced, or a specific error code on failure.

Analyzing Verification Results

Inspect the generated verify.json to extract AR and MTP metrics:

import json
from pathlib import Path

with Path("./verify-run/nightly-2026-09-02/verify.json").open() as f:
    payload = json.load(f)

for row in payload["rows"]:
    print(f"Depth {row['depth']}: AR={row['ar']:.2f}%, "
          f"MTP={row.get('mtp', 'n/a')}")

Typical output shows decreasing AR percentages at deeper token depths, with MTP values indicating Multi-Tier-Precision efficiency:


Depth 0: AR=99.8%, MTP=n/a
Depth 1: AR=98.3%, MTP=95.2%
Depth 2: AR=96.7%, MTP=92.1%

Core Implementation Files

The verification logic is distributed across these key modules:

  • mtplx/commands/forge.py: Contains _cmd_verify, _run_verify, and _run_verify_family_serve, implementing the dual-lane orchestration and live row merging.
  • mtplx/thermal.py: Implements SmartFanController and ThermalForge daemon integration for --max mode.
  • mtplx/artifacts.py: Handles model introspection and mtplx_runtime.json I/O via _read_runtime and inspection utilities.
  • mtplx/mtp_patch.py: Defines MTP contract schemas and calibration logic via _calibrate_mtp_contract and _runtime_or_default_mtp_contract.

Summary

  • MTPLX Forge provides the verify subcommand to measure AR and MTP performance in converted models.
  • The system selects between a tune lane (standard models) and a family-serve lane (native backend models like Qwen 4-exp).
  • Live progress reporting merges partial rows into verify.json during execution, eliminating wait times for final results.
  • MTP contract reuse avoids redundant calibration by loading existing contracts from mtplx_runtime.json via _runtime_or_default_mtp_contract.
  • ThermalForge integration via --max tests thermal boundaries using SmartFanController from mtplx/thermal.py.
  • Exit codes indicate success (0), empty results (1), or subprocess failures.

Frequently Asked Questions

What exit codes does forge verify return?

The command returns 0 when verification succeeds and at least one row is gathered. It returns 1 if the engine completes but produces no measurable rows. Any non-zero exit codes from underlying subprocesses (such as mtplx tune) propagate directly to the caller.

When should I use the --max flag?

Use --max when you need to verify performance under maximum thermal conditions. This flag engages the SmartFanController class from mtplx/thermal.py, which communicates with the privileged ThermalForge daemon to ramp cooling fans to maximum speed during the verification run.

How does MTP contract reuse work during verification?

Before running the verification engine, _cmd_verify calls _read_runtime to load any existing mtplx_runtime.json. If a valid MTP contract exists, _runtime_or_default_mtp_contract returns it immediately; otherwise, the system invokes _calibrate_mtp_contract from mtplx/mtp_patch.py to generate a new contract. This prevents unnecessary re-calibration of previously optimized models.

Can I verify models that require native backends like Qwen 4-exp?

Yes. Forge detects these requirements and automatically routes them through the family-serve lane managed by _run_verify_family_serve (line 2617). This lane boots a temporary mtplx serve instance to exercise the contract-specific sampler, ensuring native backend models receive proper verification without the standard tuning subprocess.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →