What Is the `mtplx_runtime.json` Contract File? Purpose and Implementation Guide
The mtplx_runtime.json file is a runtime contract that declares a model's inference-time capabilities, feature flags, and environment configuration, enabling the MTPLX framework to safely configure and serve Mixture-of-Token-Parallel-Layers models.
The mtplx_runtime.json contract file sits at the intersection of model artifacts and runtime execution in the youssofal/MTPLX repository. This JSON configuration bridges static model checkpoints with dynamic inference requirements, allowing the framework to verify compatibility, enable optional features like Multi-Token Prediction (MTP) layers, and inject environment-specific optimizations before serving begins.
Core Purpose of the Runtime Contract
The contract serves as a canonical declaration of what a given model can do at inference time. Located alongside model checkpoints—typically at /models/{model_name}/mtplx_runtime.json—this file is parsed immediately when the MTPLX server loads a model.
According to the source implementation, the contract accomplishes three critical functions:
- Capability Advertisement: Signals whether advanced features like MTP layers, draft samplers, or specific LM-head configurations are available.
- Environment Configuration: Supplies per-model overrides for system environment variables that optimize inference performance.
- Integrity Verification: Provides a
verifiedflag used by release pipelines to block unverified models from public deployment.
How the Contract Is Loaded
The primary entry point for reading the contract is the load_runtime_contract function in mtplx/backends/registry.py. This function reads the JSON file and returns a dictionary containing the model's runtime parameters.
When a user requests model serving, the framework executes the following sequence:
- Discovery: The server searches for
mtplx_runtime.jsonin the model's directory. - Parsing:
load_runtime_contractdeserializes the JSON into a Python dictionary. - Validation: The UI layer in
mtplx/ui/onboarding.pyvalidates the contract structure between lines 447-537, failing the launch if the file is missing or malformed. - Propagation: Valid contracts are passed to
mtplx/server/openai.pyto configure the inference engine.
Key Configuration Fields
Feature Gating and MTP Layer Configuration
The contract controls Multi-Token Prediction (MTP) through the mtp_depth_max field. When this value exceeds zero, the runtime enables MTP layers for speculative decoding.
In mtplx/backends/registry.py, the loader checks this field to conditionally initialize parallel token processing:
from mtplx.backends.registry import load_runtime_contract
contract, error = load_runtime_contract("/models/champion/mtplx_runtime.json")
if error:
raise RuntimeError(f"Contract loading failed: {error}")
# Enable MTP layers only if advertised
if contract.get("mtp_depth_max", 0) > 0:
model.enable_mtp_layers(contract["mtp_depth_max"])
Draft Sampling and LM Head Specifications
The contract defines parameters for draft sampling strategies used in speculative execution. Functions such as draft_sampler_spec_from_runtime_contract in mtplx/draft_sampling.py and draft_lm_head_spec_from_runtime_contract in mtplx/draft_lm_head.py extract configuration details to build appropriate sampler objects.
These specifications determine how the model generates draft tokens and manages LM-head placement during the "draft" inference path, directly impacting latency and throughput.
Runtime Environment Overrides
The runtime_env_overrides field allows models to specify system environment variables that optimize inference. Common overrides include thread-count variables like OMP_NUM_THREADS or CUDA-specific configurations.
As implemented in the loading sequence, these overrides are injected into the process environment before the model begins execution:
import os
from mtplx.backends.registry import load_runtime_contract
# Example: Applying environment overrides from contract
contract, _ = load_runtime_contract(model_path)
overrides = contract.get("runtime_env_overrides", {})
for key, value in overrides.items():
os.environ[key] = value
Verification and Release Status
The verified boolean flag acts as a gatekeeper for public deployment. Release pipelines check this field to ensure only validated models enter production. Tests in tests/test_public_cli.py explicitly block serving when verified is missing or false, preventing experimental checkpoints from being served as stable releases.
Error Handling and Missing Contracts
When mtplx_runtime.json is absent, the framework treats the model as runtime-contract-less. This state severely restricts available features and triggers warnings in the onboarding UI. According to validation logic in mtplx/ui/onboarding.py, missing contracts cause the launch sequence to fail or operate in a degraded mode, depending on the deployment context.
Summary
- The
mtplx_runtime.jsonfile is a mandatory runtime contract that describes model capabilities, configuration, and verification status in the youssofal/MTPLX ecosystem. - Loaded via
load_runtime_contractinmtplx/backends/registry.py, the file enables feature gating for MTP layers and draft samplers through fields likemtp_depth_max. - Environment optimization is achieved through
runtime_env_overrides, which injects per-model system variables before inference begins. - Safety mechanisms rely on the
verifiedflag to prevent unverified models from reaching public release, with validation occurring inmtplx/ui/onboarding.pyandmtplx/server/openai.py.
Frequently Asked Questions
What happens if mtplx_runtime.json is missing from the model directory?
The framework classifies the model as runtime-contract-less, which disables advanced features like MTP layers and draft sampling. The onboarding UI in mtplx/ui/onboarding.py will either block the launch or display warnings indicating limited functionality, and tests in tests/test_tail_profile_truth.py verify that such models cannot pass public release gates.
How does the contract affect model serving performance?
The contract directly impacts performance through speculative decoding configuration. By setting mtp_depth_max greater than zero, the contract enables Multi-Token Prediction layers that reduce latency. Additionally, runtime_env_overrides allows per-model tuning of thread counts and GPU settings, optimizing throughput for specific hardware configurations.
Where is the contract validated in the user interface?
Validation occurs in mtplx/ui/onboarding.py between lines 447-537, where the system loads and inspects the JSON structure. This code checks for required fields, validates the verified status, and displays the contract contents to users before allowing model serving to commence.
What fields are required for a valid mtplx_runtime.json?
While the schema is extensible, critical fields include mtp_depth_max (integer controlling parallel token layers), runtime_env_overrides (dictionary of environment variables), and verified (boolean indicating release pipeline approval). The draft_sampler_spec_from_runtime_contract and related functions expect specific sub-fields for draft sampling configuration, though these may be optional depending on the serving mode.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →