# MTPLX Admission Policies for Managing Memory Pressure: 5 Layers of Defense

> Discover MTPLX admission policies managing memory pressure with 5 layers of defense. Learn how shrink_to_bytes and shrink_for_admission prevent out-of-memory crashes and protect active sessions.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-05

---

**MTPLX implements a deterministic, multi-stage admission control system—centered in `SessionBank` and the OpenAI-compatible server—that uses `shrink_to_bytes()` for global pressure relief and `shrink_for_admission()` for request-time eviction to prevent out-of-memory crashes while protecting active sessions.**

The youssofal/MTPLX repository provides a model-serving infrastructure designed to handle high-concurrency inference workloads on memory-constrained systems. Managing memory pressure in this context requires granular admission policies that decide when to evict cached sessions, throttle incoming requests, or prioritize active computations over new ones. The implementation spans [`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py) for core eviction logic and [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) for admission flags and scheduling constraints.

## Core Admission Policies in MTPLX

MTPLX operates five distinct policies that form a layered defense against memory exhaustion. Each policy targets a specific stage of the request lifecycle, from system-wide pressure detection to per-request prefill validation.

### Memory-Pressure Guard (System-Wide Eviction)

The **memory-pressure guard** responds to macOS system-wide pressure reports, such as when the operating system begins swapping memory. Located at lines **2786–2809** in [`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py), this policy invokes the `shrink_to_bytes()` method to perform a global eviction pass.

The method accepts a `target_bytes` parameter and an optional `protect_active` flag. When `protect_active=True`, the guard preserves entries belonging to active sessions, evicting only the least-recently-used (LRU) idle entries until the memory budget is satisfied. This prevents mid-run re-prefill stalls while respecting system-level memory constraints.

### Prefill Admission Shed (Targeted Eviction)

Before executing a memory-heavy prefill operation, the **prefill admission shed** validates whether the projected request exceeds the sustained-pressure limit. Implemented at lines **2840–2865** in [`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py), this policy triggers `shrink_for_admission()` to perform escalating eviction.

The procedure first removes non-terminal entries, then falls back to a generic LRU pass if the deficit remains. Critical to this logic is the `protect_tokens` parameter, which accepts a list of token IDs that must remain intact during the shed. This ensures that tokens required for the incoming prompt are not evicted during the admission check.

### Prefill Admission Chain Shed (Per-Request Ceiling)

For deployments requiring stricter isolation, the **prefill admission chain shed** enforces a per-request memory ceiling. This policy is controlled via the server flag `_prefill_admission_chain_shed_enabled`, defined around lines **17432–17450** in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py).

When enabled, the server calls `shrink_for_admission()` with a larger target derived from the *admission chain* budget. This creates a secondary validation layer that runs in coordination with the standard prefill shed, providing defense-in-depth for multi-tenant or resource-constrained environments.

### Hyper-Mode Admission Cap (Hard Throttling)

When the scheduler operates in **hyper mode** (`--scheduler-mode hyper`), MTPLX enforces a hard admission cap of **exactly one** concurrent request. This simplification is implemented at lines **2752–2768** in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py).

The server validates launch flags and rejects configurations that would permit multi-admission semantics. By limiting active sessions to one, the memory-budget logic avoids complex multi-tenant arbitration, making eviction decisions deterministic and lightweight.

### Dynamic-Ceiling Policy (Active Session Protection)

The **dynamic-ceiling policy** allows the memory-pressure guard to tolerate temporary overages when evicting entries would harm active computations. Located within the `protect_active=True` branch of `shrink_to_bytes()` at lines **2790–2804**, this logic permits the bank to remain above `target_bytes` if the only remaining entries belong to the active session.

This policy trades strict adherence to the memory budget for inference latency stability, ensuring that active sessions complete without expensive re-prefill operations.

## Implementation Workflow

The admission policies execute in a deterministic cascade to handle memory pressure gracefully:

1. **System pressure detection** triggers `shrink_to_bytes()` for global LRU eviction of idle sessions.
2. **Prefill projection** invokes `shrink_for_admission()` to enforce targeted eviction of non-terminal entries before admitting new requests.
3. **Admission chain validation** optionally applies a larger eviction target when `_prefill_admission_chain_shed_enabled` is active.
4. **Hyper-mode enforcement** overrides all previous logic with a hard cap of one admitted request, simplifying the eviction calculus.

## Practical Code Examples

The following examples demonstrate how to interact with the admission policies programmatically:

```python

# Example: manually invoking the memory-pressure guard

from mtplx.session_bank import SessionBank

bank = SessionBank(...)

# Target to keep the bank under 2 GiB

evicted = bank.shrink_to_bytes(target_bytes=2 * 1024**3, protect_active=True)
print(f"Evicted {evicted} entries to respect memory pressure")

```

```python

# Example: admission-shed before a large prefill

# Assume the upcoming request needs 1 GiB more than the current budget

non_term, term = bank.shrink_for_admission(
    target_bytes=bank.total_nbytes + 1 * 1024**3,
    protect_tokens=[123, 456, 789],   # tokens that must stay intact

    reason="prefill_admission_chain",
)
print(f"Admission shed removed {non_term} non-terminal and {term} terminal entries")

```

## Summary

- **Memory-pressure guard** (`shrink_to_bytes`): Handles macOS system-wide pressure via global LRU eviction with active-session protection.
- **Prefill admission shed** (`shrink_for_admission`): Performs escalating eviction (non-terminal first) before admitting memory-heavy prefills.
- **Admission chain shed**: Enforces per-request ceilings via server flags in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py).
- **Hyper-mode cap**: Restricts concurrency to one request, eliminating multi-tenant memory arbitration complexity.
- **Dynamic-ceiling policy**: Allows temporary overages to protect active sessions from costly re-prefill operations.

## Frequently Asked Questions

### What triggers the memory-pressure guard in MTPLX?

The memory-pressure guard triggers when macOS reports system-wide memory pressure, typically when the OS begins swapping. According to the source code in [`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py), this invokes `shrink_to_bytes()` to evict LRU entries until the `target_bytes` budget is met, optionally protecting active sessions via the `protect_active` parameter.

### How does `shrink_for_admission` differ from `shrink_to_bytes`?

While `shrink_to_bytes` performs a global LRU eviction suitable for system-pressure events, `shrink_for_admission` implements an escalating strategy specifically for prefill-time decisions. As implemented at lines 2840–2865, it first attempts to remove non-terminal entries before falling back to LRU eviction, and it accepts a `protect_tokens` list to preserve specific tokens required for the incoming request.

### What is the hyper-mode admission cap and when should I use it?

The hyper-mode admission cap enforces a maximum of one concurrent request when the server runs with `--scheduler-mode hyper`. This mode, enforced at lines 2752–2768 in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), simplifies memory management by eliminating multi-session arbitration overhead. Use this mode when serving large models on single-GPU systems where memory fragmentation risks outweigh concurrency benefits.

### How can I enable the admission chain shed for stricter memory isolation?

Enable the admission chain shed by setting the `_prefill_admission_chain_shed_enabled` flag in the server configuration, located around lines 17432–17450 in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py). When active, this policy invokes `shrink_for_admission()` with an admission-chain-derived budget target, adding a secondary validation layer before accepting memory-intensive requests.