MTPLX Runtime Contract System Tiers: Storage Layers and Contract Policies Explained
The MTPLX runtime contract system uses two distinct tier categories: three storage tiers (Hot, Warm, Cold) for state management and four contract-activation tiers (Tool, No-Tool, Read-Only-Force-Answer, PI-Convergence) for policy enforcement.
MTPLX implements a tiered architecture that governs how model state, cache entries, and contract metadata are stored and enforced at runtime. The youssofal/MTPLX repository organizes these tiers into separate concerns: storage durability and contract policy selection. This guide examines both category types, their implementation in the source code, and practical usage patterns.
Storage Tiers in MTPLX
The storage tier system manages where model state lives across the memory hierarchy. Each tier represents a trade-off between access speed and durability.
Hot Tier: In-Memory Generation State
The Hot tier maintains pure in-memory state that the generator reads and writes on every token. This tier lives entirely within Python objects inside the current SessionBank instance.
- Zero disk I/O overhead
- Lost on process termination
- Ideal for per-token generation with no durability requirements
Warm Tier: Hybrid Memory-Disk Cache
The Warm tier provides a hybrid layer that holds recently-used data in memory while spilling to disk when memory pressure grows. The SessionBank holds a reference to a WarmTier object that lazily materializes rows on demand.
This tier balances performance against memory constraints, keeping recent context fast without requiring full disk I/O for larger histories.
Cold Tier: Persistent Disk Storage
The Cold tier implements fully persisted, disk-backed storage for historic rows, checkpoints, and contract files. This tier is defined in mtplx/cache_bank/cold_tier.py by the SessionBankColdTier class.
The Cold tier provides:
- Durability across sessions
- Checkpoint-restore capability
- Storage for the JSON runtime contract that drives feature-gating
Tests across the repository demonstrate Cold tier usage, including tests/test_ssd_spill.py and tests/test_cold_tier_write_budget.py.
from mtplx.cache_bank import SessionBankColdTier
# Create disk-backed tier storing rows under "/tmp/mtplx_cold"
cold = SessionBankColdTier(base_dir="/tmp/mtplx_cold", mode="on")
# Writes survive process restarts
cold.put_entry(row_id=42, data=b"...")
Contract-Activation Tiers in MTPLX
Beyond storage, the MTPLX runtime contract system defines contract-activation tiers that determine which policy contract attaches to a generation request. These tiers are inspected by server logic in mtplx/server/openai.py (lines 1364-1365) and by the onboarding UI reading mtplx_runtime.json.
Tool Contract Tier
The Tool contract provides a full schema-free description of the tool-call protocol. It defines rules such as "Read a file in ONE call" and "Never print file contents."
This tier activates when tool_contract_active evaluates to True, typically when the request contains tool calls.
No-Tool Contract Tier
The No-Tool contract supplies a minimal contract that explicitly disables tool usage. Injected when no_tools_contract_active is True.
Read-Only-Force-Answer Contract Tier
The Read-Only-Force-Answer contract limits the model to returning only the requested number of list items with markdown-list formatting rules. Activated via read_only_force_answer_contract_active.
PI-Convergence Contract Tier
The PI-Convergence contract appends a tiny suffix that nudges the model toward "plan-in-answer" style responses. Controlled by pi_convergence_contract_active.
import mtplx.server.openai as openai
# Activate tool contract (normally automatic with tool calls)
openai._tool_contract_active_for_mode = lambda _: True
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "List files in current directory."}
]
# Contract injected as first system message
with_contract = openai._with_mtplx_tool_contract(messages)
print(with_contract[0]["content"]) # "MTPLX tool contract: ..."
Runtime Contract File and UI Integration
The MTPLX runtime contract system reads its configuration from mtplx_runtime.json in the model directory. The onboarding UI in mtplx/ui/onboarding.py (lines 445-447) loads this file to expose contract status.
from mtplx.ui.onboarding import RuntimeContractStub
from pathlib import Path
import json
stub = RuntimeContractStub()
model_dir = Path("/path/to/model_dir")
contract_path = model_dir / "mtplx_runtime.json"
if contract_path.is_file():
stub.runtime_contract_data = json.loads(contract_path.read_text())
else:
stub.runtime_contract_error = "Missing contract file"
Key Implementation Files
| File | Purpose |
|---|---|
mtplx/cache_bank/cold_tier.py |
Cold tier implementation with SessionBankColdTier class |
mtplx/ui/onboarding.py |
Runtime contract JSON loading for UI |
mtplx/server/openai.py |
Contract tier selection logic |
mtplx/kpi/reference_vllm.py |
KPI counters tracking active contract tiers |
mtplx/kpi/runtime_kpis.py |
Runtime KPI definitions for tier observation |
Summary
- The MTPLX runtime contract system separates storage concerns from policy enforcement through distinct tier categories.
- Storage tiers range from ephemeral Hot (memory-only) through hybrid Warm to persistent Cold (
SessionBankColdTierinmtplx/cache_bank/cold_tier.py). - Contract tiers (Tool, No-Tool, Read-Only-Force-Answer, PI-Convergence) are boolean flags evaluated by server logic to inject appropriate policy contracts.
- The runtime contract file
mtplx_runtime.jsonprovides durable configuration read by both server and UI components. - KPI tracking in
mtplx/kpi/enables observability of which tiers are active during generation.
Frequently Asked Questions
How does the Warm tier differ from the Cold tier in MTPLX?
The Warm tier maintains a memory-first cache with lazy disk spilling under memory pressure, while the Cold tier immediately persists all data to disk. The Warm tier is accessed through a WarmTier object referenced by SessionBank, whereas Cold tier operations use SessionBankColdTier with explicit base_dir configuration.
When does MTPLX activate the Tool contract tier?
MTPLX activates the Tool contract tier when tool_contract_active evaluates to True, which occurs automatically when a generation request contains tool calls. The server logic in mtplx/server/openai.py then injects the schema-free tool protocol description as the first system message.
What happens if mtplx_runtime.json is missing?
The onboarding UI in mtplx/ui/onboarding.py detects missing contract files through RuntimeContractStub, setting runtime_contract_error to "Missing contract file" rather than loading configuration data. The server may fall back to default contract behaviors depending on implementation context.
Can multiple contract tiers be active simultaneously?
The source code in mtplx/server/openai.py evaluates contract tiers through independent boolean flags, suggesting they can coexist. However, the Tool and No-Tool contracts are mutually exclusive by design, while PI-Convergence and Read-Only-Force-Answer may combine with either tool state depending on request configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →