# How MTPLX's Memory Plan Decode System Manages VRAM to Prevent Metal OOM Errors

> MTPLX's memory plan decode system prevents Metal OOM errors by dynamically managing VRAM, partitioning memory, and shrinking caches as KV grows.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: internals
- Published: 2026-09-13

---

**MTPLX’s memory plan decode system partitions GPU memory into non-reclaimable commitments (model weights, KV cache, and a 3 GiB runtime transient reserve) and reclaimable caches (session bank and MLX allocator pool), then enforces a dynamic ceiling that automatically shrinks the cache budget as live KV grows to prevent exceeding Metal’s memory envelope.**

The MTPLX inference engine (youssofal/MTPLX) implements a deterministic VRAM orchestration layer in **[`mtplx/memory_plan.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/memory_plan.py)** that eliminates out-of-memory crashes during large language model inference. Unlike heuristic allocators, this system pre-computes a complete memory plan before model loading, separating fixed commitments from elastic caches and applying runtime ceilings that spill to SSD rather than violating the Metal GPU memory cap.

## Core VRAM Budgeting Pipeline

### Detecting Physical Memory Limits

The system begins by establishing the hardware baseline. The **`detect_total_ram_bytes()`** function (lines 39‑71) queries the host’s physical RAM via `sysctl` on macOS or `os.sysconf` on Linux, returning the total bytes available for engine budgeting.

### Enforcing the Usable Engine Envelope

Once physical RAM is known, **`usable_engine_bytes()`** (lines 73‑81) applies the **75 % envelope rule** (`ENGINE_RAM_FRACTION`), clamped to a minimum of **8 GiB** (`ENGINE_RAM_FLOOR_BYTES`) and a maximum of **192 GiB** (`ENGINE_RAM_CAP_BYTES`). This computed envelope matches Metal’s internal allocator limits and represents the absolute ceiling for all GPU allocations.

### Applying User Overrides

The **`plan_memory()`** entry point (lines 51‑88) accepts optional constraints:
- **`memory_budget_bytes`**: Simulates a smaller GPU (e.g., requesting a 48 GiB seat on a 128 GiB machine)
- **`usable_bytes_override`**: Explicit Metal limit from environment flags

The final `usable` value and its source (`usable_source`) are recorded in the plan, ensuring downstream components respect the hard boundary.

## Commitments vs. Caches Architecture

### Non-Reclaimable Commitments

Commitments are memory regions that cannot be reclaimed without aborting active requests. The system reserves three fixed commitments:

- **Model weights** (`model_weights_bytes`): Pinned for the duration of inference
- **Full-context KV cache**: Pre-allocated for the worst-case context window, sized by `kv_bytes_per_token_effective` (dense KV size × `KV_QUANT_BYTE_FACTOR` for quantization)
- **Runtime transients** (`RUNTIME_TRANSIENTS_BYTES`): A **3 GiB** static buffer (lines 54‑57) for temporary operation spikes

### Reclaimable Cache Pools

All remaining VRAM funds two elastic pools that the system can evict to host memory or SSD on demand:
- **Session bank**: Stores reusable activation tensors and intermediate results
- **MLX allocator pool**: Dynamic scratch space for Metal kernels

This split allows MTPLX to maintain a "warm" cache during idle periods while guaranteeing that commitments never exceed the usable envelope.

## Calculating Safe Context Windows

### Per-Token Cost Estimation

Before determining the maximum context length, the plan computes **`per_token`** cost by aggregating:
- Quantized KV size (controlled by `KV_QUANT_BYTE_FACTOR` at lines 24‑27)
- Family-specific auxiliary bytes (`aux_bytes_per_token`)
- Pre-fill transient overhead (`prefill_transient_bytes_per_token`)

### Context Fit Calculation

The **KV budget** is derived as:

```text
kv_budget = usable - model_weights_bytes - RUNTIME_TRANSIENTS_BYTES - BANK_FLOOR_BYTES

```

Dividing this budget by `per_token` and aligning down to the **4096-token block size** (`CONTEXT_ALIGN_TOKENS`) yields **`fit_raw`** (lines 66‑71). The final `resolved_context` is the minimum of this machine fit, the model’s architectural limit (`model_max_context`), and any user-requested override (`requested_context`).

## Runtime Memory Protection Mechanisms

### The Dynamic Ceiling

During serving, **`bank_dynamic_ceiling()`** (lines 84‑92) recalculates the session bank limit based on **`working_set_bytes`** (currently live KV cache). As the working set grows, the ceiling lowers proportionally, forcing the cache to evict older entries to SSD before the Metal allocator hits its hard cap.

### Transient Spike Protection

To handle pathological allocation spikes during complex turns, **`transient_reserve_bytes()`** (lines 150‑166) reserves headroom capped at **16 GiB** (`TRANSIENT_RESERVE_CAP_BYTES`) or half of the remaining "play" bytes (usable minus commitments). This ensures a single turn cannot starve the session bank permanently.

### N-Gram Table Streaming

The system optionally treats the n-gram lookup table as a streamed resource rather than a pinned commitment. **`ngram_table_resident_policy`** (lines 91‑107) checks the `MTPLX_NGRAM_RESIDENT` environment variable; when disabled, the table pages in on demand, freeing VRAM for model weights and KV cache.

## Practical Implementation Examples

### Building a Memory Plan

```python
from mtplx.memory_plan import plan_memory, describe_plan

# 48 GiB Mac with a 10 GiB model and default KV settings

plan = plan_memory(
    total_ram_bytes=48 * 1024**3,
    model_weights_bytes=10 * 1024**3,
)

print(describe_plan(plan))

# Output: "48G Mac: engine budget 36.0G, weights 10.0G, context 8192 (machine-bound), session bank up to 24.0G"

```

### Applying a Memory Budget Constraint

```python

# Simulate a 48 GiB seat on a 128 GiB machine

plan = plan_memory(
    total_ram_bytes=128 * 1024**3,
    model_weights_bytes=30 * 1024**3,
    memory_budget_bytes=48 * 1024**3,
)

assert plan.usable_bytes == 48 * 1024**3  # Budget override applied

```

### Adjusting the Dynamic Ceiling at Runtime

```python
from mtplx.memory_plan import bank_dynamic_ceiling

# As KV cache grows to 8 GiB, the bank ceiling shrinks

working_kv = 8 * 1024**3
new_ceiling = bank_dynamic_ceiling(plan, working_set_bytes=working_kv)
print(f"Reduced bank limit: {new_ceiling / 1024**3:.1f} GiB")

```

### Reserving for Transient Spikes

```python
from mtplx.memory_plan import transient_reserve_bytes

reserve = transient_reserve_bytes(
    peak_bytes=12 * 1024**3,
    active_bytes=3 * 1024**3,
    play_bytes=plan.usable_bytes - plan.model_weights_bytes,
)
print(f"Reserved {reserve / 1024**3:.1f} GiB for spikes")  # Capped at 16 GiB

```

## Summary

- **Fixed commitments** (weights, KV cache, 3 GiB runtime transients) are subtracted from the usable envelope first, ensuring they never trigger OOM.
- **Reclaimable caches** (session bank, MLX pool) consume remaining VRAM and spill to SSD when the dynamic ceiling lowers.
- **Dynamic ceilings** automatically shrink the cache budget as live KV grows, preventing Metal allocator failures during long conversations.
- **Quantization awareness** via `KV_QUANT_BYTE_FACTOR` allows the plan to accurately budget for Q8 or Q4 KV caches without changing allocation logic.
- **Transient protection** caps spike reserves at 16 GiB to balance safety against cache warmth.

## Frequently Asked Questions

### How does MTPLX prevent Metal out-of-memory errors while maximizing context length?

The system pre-calculates a **hard usable envelope** (75 % of physical RAM, capped at 192 GiB) and subtracts non-negotiable commitments (weights, KV cache, runtime transients) before allocating any cache memory. During inference, **`bank_dynamic_ceiling()`** continuously shrinks the cache limit as the KV working set grows, forcing eviction to SSD rather than allowing allocations to breach the Metal limit.

### What is the difference between commitments and caches in the memory plan decode system?

**Commitments** are memory regions that cannot be reclaimed without aborting the request, specifically `model_weights_bytes`, the full-context KV reserve, and the static 3 GiB `RUNTIME_TRANSIENTS_BYTES` buffer. **Caches** comprise the session bank and MLX allocator pool; these are treated as elastic resources that can be reclaimed instantly when the dynamic ceiling drops, allowing the system to trade latency for stability.

### How does KV quantization affect VRAM budgeting?

The plan applies **`KV_QUANT_BYTE_FACTOR`** (lines 24‑27) to the dense KV size when computing `kv_bytes_per_token_effective`. Setting this factor to `0.5` for Q4 quantization or `1.0` for full precision directly reduces the per-token commitment, increasing the KV budget available for longer context windows without changing the commitment-vs-cache architecture.

### Can I manually limit MTPLX VRAM usage below the default 75 % envelope?

Yes. Pass **`memory_budget_bytes`** to `plan_memory()` to simulate a smaller GPU. For example, providing `48 * 1024**3` on a 128 GiB machine forces the plan to treat 48 GiB as the usable ceiling, overriding the 75 % calculation and preventing the engine from claiming the full hardware envelope.