How MTPLX's Memory Plan Decode System Manages VRAM to Prevent Metal OOM Errors

MTPLX’s memory plan decode system partitions GPU memory into non-reclaimable commitments (model weights, KV cache, and a 3 GiB runtime transient reserve) and reclaimable caches (session bank and MLX allocator pool), then enforces a dynamic ceiling that automatically shrinks the cache budget as live KV grows to prevent exceeding Metal’s memory envelope.

The MTPLX inference engine (youssofal/MTPLX) implements a deterministic VRAM orchestration layer in mtplx/memory_plan.py that eliminates out-of-memory crashes during large language model inference. Unlike heuristic allocators, this system pre-computes a complete memory plan before model loading, separating fixed commitments from elastic caches and applying runtime ceilings that spill to SSD rather than violating the Metal GPU memory cap.

Core VRAM Budgeting Pipeline

Detecting Physical Memory Limits

The system begins by establishing the hardware baseline. The detect_total_ram_bytes() function (lines 39‑71) queries the host’s physical RAM via sysctl on macOS or os.sysconf on Linux, returning the total bytes available for engine budgeting.

Enforcing the Usable Engine Envelope

Once physical RAM is known, usable_engine_bytes() (lines 73‑81) applies the 75 % envelope rule (ENGINE_RAM_FRACTION), clamped to a minimum of 8 GiB (ENGINE_RAM_FLOOR_BYTES) and a maximum of 192 GiB (ENGINE_RAM_CAP_BYTES). This computed envelope matches Metal’s internal allocator limits and represents the absolute ceiling for all GPU allocations.

Applying User Overrides

The plan_memory() entry point (lines 51‑88) accepts optional constraints:

  • memory_budget_bytes: Simulates a smaller GPU (e.g., requesting a 48 GiB seat on a 128 GiB machine)
  • usable_bytes_override: Explicit Metal limit from environment flags

The final usable value and its source (usable_source) are recorded in the plan, ensuring downstream components respect the hard boundary.

Commitments vs. Caches Architecture

Non-Reclaimable Commitments

Commitments are memory regions that cannot be reclaimed without aborting active requests. The system reserves three fixed commitments:

  • Model weights (model_weights_bytes): Pinned for the duration of inference
  • Full-context KV cache: Pre-allocated for the worst-case context window, sized by kv_bytes_per_token_effective (dense KV size × KV_QUANT_BYTE_FACTOR for quantization)
  • Runtime transients (RUNTIME_TRANSIENTS_BYTES): A 3 GiB static buffer (lines 54‑57) for temporary operation spikes

Reclaimable Cache Pools

All remaining VRAM funds two elastic pools that the system can evict to host memory or SSD on demand:

  • Session bank: Stores reusable activation tensors and intermediate results
  • MLX allocator pool: Dynamic scratch space for Metal kernels

This split allows MTPLX to maintain a "warm" cache during idle periods while guaranteeing that commitments never exceed the usable envelope.

Calculating Safe Context Windows

Per-Token Cost Estimation

Before determining the maximum context length, the plan computes per_token cost by aggregating:

  • Quantized KV size (controlled by KV_QUANT_BYTE_FACTOR at lines 24‑27)
  • Family-specific auxiliary bytes (aux_bytes_per_token)
  • Pre-fill transient overhead (prefill_transient_bytes_per_token)

Context Fit Calculation

The KV budget is derived as:

kv_budget = usable - model_weights_bytes - RUNTIME_TRANSIENTS_BYTES - BANK_FLOOR_BYTES

Dividing this budget by per_token and aligning down to the 4096-token block size (CONTEXT_ALIGN_TOKENS) yields fit_raw (lines 66‑71). The final resolved_context is the minimum of this machine fit, the model’s architectural limit (model_max_context), and any user-requested override (requested_context).

Runtime Memory Protection Mechanisms

The Dynamic Ceiling

During serving, bank_dynamic_ceiling() (lines 84‑92) recalculates the session bank limit based on working_set_bytes (currently live KV cache). As the working set grows, the ceiling lowers proportionally, forcing the cache to evict older entries to SSD before the Metal allocator hits its hard cap.

Transient Spike Protection

To handle pathological allocation spikes during complex turns, transient_reserve_bytes() (lines 150‑166) reserves headroom capped at 16 GiB (TRANSIENT_RESERVE_CAP_BYTES) or half of the remaining "play" bytes (usable minus commitments). This ensures a single turn cannot starve the session bank permanently.

N-Gram Table Streaming

The system optionally treats the n-gram lookup table as a streamed resource rather than a pinned commitment. ngram_table_resident_policy (lines 91‑107) checks the MTPLX_NGRAM_RESIDENT environment variable; when disabled, the table pages in on demand, freeing VRAM for model weights and KV cache.

Practical Implementation Examples

Building a Memory Plan

from mtplx.memory_plan import plan_memory, describe_plan

# 48 GiB Mac with a 10 GiB model and default KV settings

plan = plan_memory(
    total_ram_bytes=48 * 1024**3,
    model_weights_bytes=10 * 1024**3,
)

print(describe_plan(plan))

# Output: "48G Mac: engine budget 36.0G, weights 10.0G, context 8192 (machine-bound), session bank up to 24.0G"

Applying a Memory Budget Constraint


# Simulate a 48 GiB seat on a 128 GiB machine

plan = plan_memory(
    total_ram_bytes=128 * 1024**3,
    model_weights_bytes=30 * 1024**3,
    memory_budget_bytes=48 * 1024**3,
)

assert plan.usable_bytes == 48 * 1024**3  # Budget override applied

Adjusting the Dynamic Ceiling at Runtime

from mtplx.memory_plan import bank_dynamic_ceiling

# As KV cache grows to 8 GiB, the bank ceiling shrinks

working_kv = 8 * 1024**3
new_ceiling = bank_dynamic_ceiling(plan, working_set_bytes=working_kv)
print(f"Reduced bank limit: {new_ceiling / 1024**3:.1f} GiB")

Reserving for Transient Spikes

from mtplx.memory_plan import transient_reserve_bytes

reserve = transient_reserve_bytes(
    peak_bytes=12 * 1024**3,
    active_bytes=3 * 1024**3,
    play_bytes=plan.usable_bytes - plan.model_weights_bytes,
)
print(f"Reserved {reserve / 1024**3:.1f} GiB for spikes")  # Capped at 16 GiB

Summary

  • Fixed commitments (weights, KV cache, 3 GiB runtime transients) are subtracted from the usable envelope first, ensuring they never trigger OOM.
  • Reclaimable caches (session bank, MLX pool) consume remaining VRAM and spill to SSD when the dynamic ceiling lowers.
  • Dynamic ceilings automatically shrink the cache budget as live KV grows, preventing Metal allocator failures during long conversations.
  • Quantization awareness via KV_QUANT_BYTE_FACTOR allows the plan to accurately budget for Q8 or Q4 KV caches without changing allocation logic.
  • Transient protection caps spike reserves at 16 GiB to balance safety against cache warmth.

Frequently Asked Questions

How does MTPLX prevent Metal out-of-memory errors while maximizing context length?

The system pre-calculates a hard usable envelope (75 % of physical RAM, capped at 192 GiB) and subtracts non-negotiable commitments (weights, KV cache, runtime transients) before allocating any cache memory. During inference, bank_dynamic_ceiling() continuously shrinks the cache limit as the KV working set grows, forcing eviction to SSD rather than allowing allocations to breach the Metal limit.

What is the difference between commitments and caches in the memory plan decode system?

Commitments are memory regions that cannot be reclaimed without aborting the request, specifically model_weights_bytes, the full-context KV reserve, and the static 3 GiB RUNTIME_TRANSIENTS_BYTES buffer. Caches comprise the session bank and MLX allocator pool; these are treated as elastic resources that can be reclaimed instantly when the dynamic ceiling drops, allowing the system to trade latency for stability.

How does KV quantization affect VRAM budgeting?

The plan applies KV_QUANT_BYTE_FACTOR (lines 24‑27) to the dense KV size when computing kv_bytes_per_token_effective. Setting this factor to 0.5 for Q4 quantization or 1.0 for full precision directly reduces the per-token commitment, increasing the KV budget available for longer context windows without changing the commitment-vs-cache architecture.

Can I manually limit MTPLX VRAM usage below the default 75 % envelope?

Yes. Pass memory_budget_bytes to plan_memory() to simulate a smaller GPU. For example, providing 48 * 1024**3 on a 128 GiB machine forces the plan to treat 48 GiB as the usable ceiling, overriding the 75 % calculation and preventing the engine from claiming the full hardware envelope.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →