Experimental AR-Only Paths for MLX LM in MTPLX: Complete Model Guide

MTPLX provides three experimental AR-only model variants—lfm2-moe-ar, iquestcoder-ar, and llama-ar—that run on the MLX LM stack without native Multi-Token-Probability (MTP) heads, registered in mtplx/backends/registry.py with the support level experimental-mlx-lm-ar-only.

The MTPLX repository (youssofal/MTPLX) extends the MLX LM ecosystem with dedicated experimental AR-only paths for architectures lacking verified Multi-Token-Probability implementations. These paths enable high-performance autoregressive inference on specific model variants using the mlx_lm_ar backend while bypassing MTP head requirements entirely.

What Are Experimental AR-Only MLX LM Paths?

Experimental AR-only paths in MTPLX refer to model configurations that utilize the MLX LM backend for pure autoregressive (AR) generation without relying on Multi-Token-Probability (MTP) heads. According to the source code in mtplx/backends/registry.py, these entries are explicitly marked with the support level experimental-mlx-lm-ar-only because the MLX LM backend does not yet provide verified MTP implementations for these specific architectures. Despite their experimental status, these models support native-ar-only runtime compatibility and are fully functional for causal language modeling tasks.

The Three Experimental AR-Only Models

MTPLX currently ships with three distinct experimental AR-only model identifiers, each targeting different architectural requirements while sharing the mlx_lm_ar backend.

LiquidAI LFM2.5 MoE (MLX)

The lfm2-moe-ar identifier provides access to the LiquidAI LFM2.5 MoE architecture through MLX LM. Defined in mtplx/backends/registry.py at lines 141-150, this variant implements a hybrid ShortConv + GQA MoE structure without an MTP head. The model loads via the bundled mlx-lm lfm2_moe module, making it suitable for efficient mixture-of-experts inference on Apple Silicon hardware.

IQuest Coder V1 (MLX)

Registered at lines 163-172 of mtplx/backends/registry.py, the iquestcoder-ar model offers a Llama-architecture coder specifically remapped for AR-only execution. This variant strips the MTP head typically associated with IQuest models, serving target-only autoregressive generation through the MLX LM stack.

Llama-Architecture AR (MLX)

The llama-ar entry (lines 183-193 in mtplx/backends/registry.py) supports plain Llama-style checkpoints without MTP heads, including compatible variants like G9v3 and MiniCPM5-1B. This serves as the general-purpose experimental AR-only path for standard Llama architectures that require MLX LM compatibility without multi-token prediction capabilities.

How to Use Experimental AR-Only Models

MTPLX provides multiple interfaces for interacting with these experimental paths, from command-line discovery to programmatic loading.

Listing Available Models via CLI

You can filter the model registry to display only experimental AR-only MLX LM variants using the support-level filter:

mtplx model list --support-level experimental-mlx-lm-ar-only

This command queries mtplx/commands/public.py to return the three identifiers: lfm2-moe-ar, iquestcoder-ar, and llama-ar.

Loading Models Programmatically

To load an experimental AR-only model in Python, import the Model class from mtplx and specify the appropriate identifier with MTP disabled:

from mtplx import Model

# Select one of the experimental AR-only variants

model_id = "lfm2-moe-ar"  # or "iquestcoder-ar", "llama-ar"

# Initialize with AR-only configuration

model = Model.from_pretrained(
    model_id,
    mlp=False,      # Disable MTP head (default for AR-only models)

    device="gpu"    # Use "cpu" for CPU-only inference

)

# Execute generation

output = model.generate(
    "Explain quantum entanglement in one sentence.", 
    max_new_tokens=20
)
print(output)

The mlp=False parameter ensures the model initializes without expecting an MTP head, matching the experimental path configuration stored in the registry.

Serving Models with AR-Only Mode

When deploying models via the MTPLX server, explicitly disable MTP to align with the experimental AR-only path:

mtplx serve --model llama-ar --mtp-off

The --mtp-off flag is processed by mtplx/cli.py to force AR-only execution, ensuring compatibility with the native-ar-only runtime specification defined in the registry entries.

Technical Implementation Details

The experimental AR-only infrastructure relies on coordinated implementation across several core modules. The mtplx/backends/registry.py file defines the architecture catalog and assigns the experimental-mlx-lm-ar-only support level to the three variants, while mtplx/models/ (including files like deepseek_v4.py and llama.py) provides the backend implementations accessed via the mlx_lm_ar backend. The CLI entry point in mtplx/cli.py parses the --mtp-off flag to select appropriate runtime paths, and mtplx/commands/public.py handles registry filtering for the model list command.

Summary

  • Three experimental variants: MTPLX supports lfm2-moe-ar, iquestcoder-ar, and llama-ar as AR-only MLX LM models.
  • Registry location: All three are defined in mtplx/backends/registry.py with support level experimental-mlx-lm-ar-only.
  • Backend compatibility: These models utilize the mlx_lm_ar backend with native-ar-only runtime compatibility.
  • CLI integration: Use --support-level experimental-mlx-lm-ar-only to list models and --mtp-off when serving.

Frequently Asked Questions

What does "AR-only" mean in MTPLX?

AR-only refers to pure autoregressive generation mode where the model predicts one token at a time sequentially without using Multi-Token-Probability (MTP) heads. In MTPLX, these paths disable MTP functionality to support architectures that either lack MTP implementations or run on backends without verified MTP support.

Why are these MLX LM paths marked as experimental?

These paths carry the experimental-mlx-lm-ar-only designation because the MLX LM backend does not yet provide verified MTP implementations for the specific architectures involved (LiquidAI LFM2.5 MoE, IQuest Coder, and certain Llama variants). The AR-only implementation is stable and functional, but the experimental flag indicates that MTP capabilities may be added in future releases.

Can I use experimental AR-only models for production inference?

While the native-ar-only runtime compatibility ensures these models run correctly on MLX LM, the experimental designation suggests they should be deployed with appropriate testing. The underlying inference engines in mtplx/models/ are production-ready, but the lack of MTP heads means you cannot leverage multi-token prediction optimizations available in non-experimental variants.

How do I disable MTP when loading an experimental model?

To disable MTP, set mlp=False when calling Model.from_pretrained() in Python, or use the --mtp-off flag when serving via the CLI. This configuration aligns with the experimental AR-only path requirements defined in mtplx/backends/registry.py and prevents errors from missing MTP head weights.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →