How to Verify a Model with MTPLX Forge: Step-by-Step CLI and API Guide
MTPLX Forge provides a dedicated verify subcommand that measures arithmetic-reuse (AR) and Multi-Tier-Precision (MTP) performance by executing either the standard tune lane or the specialized family-serve lane for native backend models.
Verifying converted models in the youssofal/MTPLX repository ensures that arithmetic-reuse optimizations and MTP contracts perform correctly under load. To verify a model with MTPLX Forge, you invoke the forge verify command, which orchestrates subprocesses, merges partial results in real-time, and emits structured performance data to verify.json.
Verification Pipeline Overview
The verification flow implemented in mtplx/commands/forge.py proceeds through eight distinct stages:
- Argument Dispatch: The CLI entry point
cmd_forge_publicroutes to_cmd_verify(line 84) whenforge verifyis invoked. - Output Setup: The helper
_run_dircreates a timestamped folder underoutputs/forge-verifyunless--outand--run-idspecify a custom path. - Metadata Loading:
_read_runtimeingests existingmtplx_runtime.jsonto reuse previously calibrated MTP contracts. - Engine Selection:
_run_verify(line 2504) branches into either thetunelane or thefamily-servelane for models requiring native backends like Qwen 4-exp. - Tune Lane Execution: For standard models, Forge spawns
mtplx tuneas a subprocess with--max-tokens, the selected prompt suite, and MTP contract options (--base-hidden-variant,--mtp-hidden-variant,--concat-order). During execution,_merge_candidate_rowscontinuously aggregates partial results into a liveverify.json. - Family-Serve Lane: For native backend models,
_run_verify_family_serve(line 2617) boots a temporarymtplx serveinstance on a random port, optionally wrapping requests withSmartFanControllerwhen--maxis set, then issues first-load requests to gather AR and family-default rows. - Post-Processing:
_annotate_verify_rowsenriches raw data before_stamp_runtime_metadatawrites results back tomtplx_runtime.jsonif--stampis provided. - Output Emission: By default, pretty-printed JSON goes to stdout;
--jsonemits machine-readable JSON. Exit codes are 0 (success), 1 (no rows gathered), or the subprocess error code.
Essential CLI Flags
| Flag | Description |
|---|---|
--path <dir> |
Path to the model directory containing mtplx_runtime.json or raw model files. |
--out <dir> |
Destination directory for verification artifacts. |
--run-id <id> |
Identifier grouping a verification run for CI pipelines. |
--max |
Enables MAX-fan mode via SmartFanController, engaging the ThermalForge daemon. |
--max-tokens <N> |
Upper bound on tokens per request (default: 2048). |
--suite <name> |
Prompt suite driving verification (default: long-code-uncapped). |
--stamp |
Updates mtplx_runtime.json to mark the model as verified. |
--json |
Outputs pure JSON to stdout for scripting. |
The Two Verification Lanes
MTPLX Forge utilizes two distinct verification paths depending on model requirements.
The Tune Lane
Most models traverse the tune lane, which leverages existing MTPLX tuning infrastructure. This lane executes mtplx tune as a subprocess, passing MTP contract parameters such as --base-hidden-variant and --concat-order. While the subprocess runs, Forge continuously merges partial rows and writes incremental updates to verify.json, enabling downstream tools to monitor progress without waiting for completion.
The Family-Serve Lane
Models requiring native backends (specifically Qwen 4-exp) are routed through the family-serve lane. This path invokes _run_verify_family_serve to boot a temporary mtplx serve instance on a random port. When --max is specified, the system initializes SmartFanController from mtplx/thermal.py to communicate with the privileged ThermalForge daemon, ramping fans to test MAX performance mode. The lane issues first-load requests to collect AR metrics and family-default rows without requiring a full tuning cycle.
Step-by-Step Usage Examples
Basic Model Verification
Verify a local model directory with MAX-fan enabled and machine-readable output:
mtplx forge verify \
--path ./my-model \
--out ./verify-run \
--run-id nightly-2026-09-02 \
--max \
--suite long-code-uncapped \
--json > verify-run/result.json
--pathtargets the model directory.--outand--run-idcreate./verify-run/nightly-2026-09-02for artifacts.--maxforces ThermalForge fan-boost mode.--jsonformats output for programmatic consumption.
Programmatic Verification via Python API
For automated pipelines, invoke _cmd_verify directly:
from pathlib import Path
from mtplx.commands.forge import _cmd_verify, ForgeError
args = type("Args", (), {
"path": Path("./my-model"),
"out": Path("./verify-run"),
"run_id": "py-run-001",
"max": True,
"suite": "long-code-uncapped",
"json": True,
"stamp": False,
})
try:
exit_code = _cmd_verify(args)
print(f"Verification finished with exit code {exit_code}")
except ForgeError as e:
print(f"Forge failed: {e}")
The function returns 0 on success, 1 if no verification rows were produced, or a specific error code on failure.
Analyzing Verification Results
Inspect the generated verify.json to extract AR and MTP metrics:
import json
from pathlib import Path
with Path("./verify-run/nightly-2026-09-02/verify.json").open() as f:
payload = json.load(f)
for row in payload["rows"]:
print(f"Depth {row['depth']}: AR={row['ar']:.2f}%, "
f"MTP={row.get('mtp', 'n/a')}")
Typical output shows decreasing AR percentages at deeper token depths, with MTP values indicating Multi-Tier-Precision efficiency:
Depth 0: AR=99.8%, MTP=n/a
Depth 1: AR=98.3%, MTP=95.2%
Depth 2: AR=96.7%, MTP=92.1%
Core Implementation Files
The verification logic is distributed across these key modules:
mtplx/commands/forge.py: Contains_cmd_verify,_run_verify, and_run_verify_family_serve, implementing the dual-lane orchestration and live row merging.mtplx/thermal.py: ImplementsSmartFanControllerand ThermalForge daemon integration for--maxmode.mtplx/artifacts.py: Handles model introspection andmtplx_runtime.jsonI/O via_read_runtimeand inspection utilities.mtplx/mtp_patch.py: Defines MTP contract schemas and calibration logic via_calibrate_mtp_contractand_runtime_or_default_mtp_contract.
Summary
- MTPLX Forge provides the
verifysubcommand to measure AR and MTP performance in converted models. - The system selects between a tune lane (standard models) and a family-serve lane (native backend models like Qwen 4-exp).
- Live progress reporting merges partial rows into
verify.jsonduring execution, eliminating wait times for final results. - MTP contract reuse avoids redundant calibration by loading existing contracts from
mtplx_runtime.jsonvia_runtime_or_default_mtp_contract. - ThermalForge integration via
--maxtests thermal boundaries usingSmartFanControllerfrommtplx/thermal.py. - Exit codes indicate success (
0), empty results (1), or subprocess failures.
Frequently Asked Questions
What exit codes does forge verify return?
The command returns 0 when verification succeeds and at least one row is gathered. It returns 1 if the engine completes but produces no measurable rows. Any non-zero exit codes from underlying subprocesses (such as mtplx tune) propagate directly to the caller.
When should I use the --max flag?
Use --max when you need to verify performance under maximum thermal conditions. This flag engages the SmartFanController class from mtplx/thermal.py, which communicates with the privileged ThermalForge daemon to ramp cooling fans to maximum speed during the verification run.
How does MTP contract reuse work during verification?
Before running the verification engine, _cmd_verify calls _read_runtime to load any existing mtplx_runtime.json. If a valid MTP contract exists, _runtime_or_default_mtp_contract returns it immediately; otherwise, the system invokes _calibrate_mtp_contract from mtplx/mtp_patch.py to generate a new contract. This prevents unnecessary re-calibration of previously optimized models.
Can I verify models that require native backends like Qwen 4-exp?
Yes. Forge detects these requirements and automatically routes them through the family-serve lane managed by _run_verify_family_serve (line 2617). This lane boots a temporary mtplx serve instance to exercise the contract-specific sampler, ensuring native backend models receive proper verification without the standard tuning subprocess.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →