How to Run AR-only Models with MTPLX Using the `--no-mtp` Flag

Add the --no-mtp flag to any MTPLX command to disable Mixed-Token Prediction and enable pure autoregressive (AR) decoding for models like Laguna-S-2.1 that lack an MTP head.

MTPLX is designed to serve both Mixed-Token-Prediction (MTP) models and pure autoregressive models through a unified interface. When working with AR-only architectures, you must explicitly tell the runtime to skip MTP initialization to prevent errors about missing prediction heads. This guide explains exactly how the --no-mtp flag works across the MTPLX codebase and how to use it correctly with the youssofal/MTPLX repository.

Why AR-only Models Require the --no-mtp Flag

Unlike standard MTP models, AR-only models such as Laguna-S-2.1 do not ship with a native MTP head in their architecture. According to the source code in mtplx/model_catalog.py (line 51), these target-only AR models will raise runtime errors if the system attempts to initialize MTP components. The --no-mtp flag instructs MTPLX to bypass all MTP-related processing and use pure AR decoding throughout the generation pipeline.

Without this flag, the runtime expects to find MTP weights and configuration that simply do not exist in AR-only checkpoints, leading to initialization failures before inference can begin.

How the --no-mtp Flag Works in the MTPLX Architecture

The flag propagates through three critical layers of the MTPLX system, enforcing AR-only mode at every stage of the execution pipeline.

CLI Argument Parsing in mtplx/cli.py

The entry point for the flag is defined in mtplx/cli.py at lines 304–313 and 696. When present in the command line, the argument parser stores args.no_mtp = True and passes this boolean downstream to all subcommands. This top-level integration means the flag is available globally across the CLI interface, including the serve, run, and start commands.

Runtime Model Loading in mtplx/runtime.py

At line 635 in mtplx/runtime.py, the system checks the flag value and conditionally loads the model with mtp=False. This forces the generation pipeline to instantiate pure AR decoding modules rather than attempting to initialize the multi-token prediction heads. The runtime comment at this location explicitly documents this behavior as the mechanism for forcing AR-only execution.

Generation Mode Selection in mtplx/commands/public.py

Helper functions in mtplx/commands/public.py (spanning lines 12168–12685) inspect the flag to determine which generation mode constant to use. When --no-mtp is detected, the system selects GENERATION_MODE_AR instead of GENERATION_MODE_MTP, ensuring that sampling loops, token selection, and caching strategies all respect the autoregressive-only constraint. The flag also appends a "–no-mtp" suffix to help messages throughout this file to improve CLI discoverability.

Practical Examples for Running AR-only Models

Use the --no-mtp flag with any MTPLX command to serve or query AR-only models. The following examples demonstrate the correct syntax for common operations:


# Start an interactive CLI session with an AR-only model

mtplx start cli --no-mtp

# Serve an AR-only model as an API endpoint

mtplx serve --model path/to/laguna-s-2.1 --no-mtp

# Run a one-shot generation command in AR-only mode

mtplx run "Explain quantum tunneling" --no-mtp

# Use the quick-start example with the flag

mtplx quickstart --no-mtp

These patterns are documented in docs/quickstart.md (line 32) and README.md (line 156), confirming that the flag integrates seamlessly with the primary user-facing workflows.

Key Files and Their Roles

Understanding where the --no-mtp logic lives helps with debugging and extending MTPLX for custom AR-only deployments:

  • mtplx/cli.py (lines 304–313, 696) – Defines the argument parser entry point that captures --no-mtp and stores it as args.no_mtp
  • mtplx/runtime.py (line 635) – Enforces the flag by loading models with mtp=False when AR-only mode is requested
  • mtplx/commands/public.py (lines 12168–12685) – Implements generation-mode switching between GENERATION_MODE_AR and GENERATION_MODE_MTP
  • mtplx/model_catalog.py (line 51) – Documents which target-only AR models require the flag, serving as the source of truth for model compatibility
  • docs/api.md (line 63) – Provides API reference documentation for the flag
  • docs/quickstart.md (line 32) – Shows practical usage examples for new users

Summary

  • AR-only models like Laguna-S-2.1 lack MTP heads and will fail without the --no-mtp flag.
  • The flag is defined in mtplx/cli.py and stored as args.no_mtp = True for downstream consumption.
  • At runtime, mtplx/runtime.py uses the flag to load models with mtp=False, forcing pure autoregressive decoding.
  • Generation logic in mtplx/commands/public.py switches between GENERATION_MODE_AR and GENERATION_MODE_MTP based on the flag state.
  • The flag works globally across all MTPLX commands including serve, run, start, and quickstart.

Frequently Asked Questions

Which models require the --no-mtp flag when using MTPLX?

AR-only models such as Laguna-S-2.1 require the --no-mtp flag because they do not contain a native MTP head in their architecture. According to mtplx/model_catalog.py (line 51), these models are explicitly cataloged as target-only AR models that must be served with this flag to avoid initialization errors regarding missing prediction heads.

How does MTPLX internally process the --no-mtp flag?

The flag triggers a three-stage enforcement mechanism: the CLI parser in mtplx/cli.py (lines 304–313) captures the argument as args.no_mtp, the runtime loader in mtplx/runtime.py (line 635) initializes the model with mtp=False, and the command layer in mtplx/commands/public.py selects GENERATION_MODE_AR instead of the default MTP mode for the generation loop.

Can I use the --no-mtp flag with any MTPLX subcommand?

Yes. Because the flag is attached to the top-level argument parser in mtplx/cli.py, it is available across all subcommands including mtplx serve, mtplx run, mtplx start cli, and mtplx quickstart. The documentation in docs/quickstart.md (line 32) and README.md (line 156) confirms this global availability for both interactive and API serving scenarios.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →