# Does Forge Validate That the Converted Model Is Actually Faster?

> Forge validates model speedup by benchmarking every converted model and raising an error if the MTL-P version isn't faster than the original autoregressive implementation.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: best-practices
- Published: 2026-09-08

---

**Yes, Forge explicitly benchmarks every converted model and raises a `ForgeError` if the MTL-P version does not demonstrate a measurable speedup over the original autoregressive implementation.**

The **MTPLX** toolchain includes a **Forge** subsystem that converts models to use Multi-Token Prediction (MTP) through MTL-P. Unlike standard conversion tools that accept any output, Forge enforces a strict performance contract by validating that the converted artifact actually accelerates inference before marking the conversion as successful.

## How Forge Enforces Performance Validation

### Benchmarking Against the Reference Implementation

Inside [`mtplx/commands/forge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py) (approximately lines 3580–3600), Forge executes a comparative benchmark between the original autoregressive model (**AR**) and the newly converted **MTP** implementation. The system measures wall-clock inference times for both variants using identical inputs to calculate an objective speedup metric.

The core validation logic evaluates an outcome object produced by the benchmark:

```python

# Simplified excerpt from mtplx/commands/forge.py

outcome = benchmark.compare(ar_time, mtp_time)
if outcome.speedup <= 0:
    raise ForgeError("MTP did not accelerate this model; AR was faster.")

```

### The Speedup Threshold Check

Forge treats any speedup value of zero or below as a failed conversion. The code explicitly checks `if outcome.speedup <= 0:` immediately after the benchmark completes. This threshold is absolute—there is no tolerance for marginal or negative performance changes.

## Handling Validation Failures

### Raising ForgeError for Slow Conversions

When validation fails, Forge aborts the workflow and raises a `ForgeError` with specific messaging. The test suite in [`tests/test_forge_cli.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_forge_cli.py) (line 976) asserts that Forge emits explicit error strings such as **"MTP did not accelerate this model; AR was faster"** or **"MTP did not accelerate this model; draft acceptance collapsed"** when the benchmark indicates no performance improvement.

You can observe this behavior programmatically:

```python
from mtplx.commands.forge import ForgeError, run_forge

try:
    result = run_forge(
        source="path/to/model",
        recipe={"quant": "bf16", "target": "mlx"},
        run_id="my-run",
    )
except ForgeError as exc:
    # Conversion rejected: MTP was not faster

    print(f"Validation failed: {exc}")
else:
    # Successful conversion with proven speedup

    print("Conversion validated: MTP is faster than AR")

```

### Preventing Artifact Generation

If the speedup validation fails, Forge prevents the creation of the final stamped artifact. This ensures that only demonstrably faster models receive the [`mtplx_runtime.json`](https://github.com/youssofal/MTPLX/blob/main/mtplx_runtime.json) contract and are eligible for deployment, protecting downstream pipelines from performance regressions.

## Successful Validation and Provenance

### Stamping the Runtime Contract

Upon passing the performance check, Forge invokes the logic in [`mtplx/metadata_scrub.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/metadata_scrub.py) (line 3) to stamp the [`mtplx_runtime.json`](https://github.com/youssofal/MTPLX/blob/main/mtplx_runtime.json) file with provenance data. This establishes an auditable record that the artifact has been verified as faster than its AR counterpart, including the input paths and conversion parameters used during the Forge process.

### Command-Line Behavior

When invoking Forge from the CLI, a failed validation exits with a non-zero status:

```bash
mtplx forge \
    --source path/to/model \
    --recipe '{"quant":"bf16","target":"mlx"}' \
    --run-id my-run

```

If the conversion does not yield a speedup, the command prints:

```

ForgeError: MTP did not accelerate this model; AR was faster.

```

## Summary

- Forge benchmarks every conversion against the original AR implementation before accepting it.
- A `speedup` value of zero or below triggers an immediate `ForgeError` and aborts the process.
- Error messages explicitly distinguish between general slowness and draft acceptance collapse.
- Only validated, faster models receive the [`mtplx_runtime.json`](https://github.com/youssofal/MTPLX/blob/main/mtplx_runtime.json) provenance stamp.
- The validation is enforced in [`mtplx/commands/forge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py) and verified by [`tests/test_forge_cli.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_forge_cli.py).

## Frequently Asked Questions

### What threshold does Forge use to determine if a model is faster?

Forge requires a strictly positive speedup (speedup > 0). If the benchmark returns zero or a negative value, the conversion fails validation and Forge raises an error.

### Can I bypass the performance validation in Forge?

No. The speedup check is hardcoded in the Forge workflow within [`mtplx/commands/forge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/forge.py). There is no configuration flag to disable this validation, ensuring all produced artifacts meet the performance contract.

### Where does Forge store the benchmark results?

Upon successful validation, Forge records provenance data in [`mtplx_runtime.json`](https://github.com/youssofal/MTPLX/blob/main/mtplx_runtime.json) via [`mtplx/metadata_scrub.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/metadata_scrub.py), documenting the conversion parameters and verification status for the accelerated artifact.

### What error will I see if the MTP model is slower?

You will receive a `ForgeError` with the message **"MTP did not accelerate this model; AR was faster"** or **"MTP did not accelerate this model; draft acceptance collapsed"**, depending on the specific failure mode detected during benchmarking.