MTPLX Admission Policies for Managing Memory Pressure: 5 Layers of Defense
MTPLX implements a deterministic, multi-stage admission control system—centered in SessionBank and the OpenAI-compatible server—that uses shrink_to_bytes() for global pressure relief and shrink_for_admission() for request-time eviction to prevent out-of-memory crashes while protecting active sessions.
The youssofal/MTPLX repository provides a model-serving infrastructure designed to handle high-concurrency inference workloads on memory-constrained systems. Managing memory pressure in this context requires granular admission policies that decide when to evict cached sessions, throttle incoming requests, or prioritize active computations over new ones. The implementation spans mtplx/session_bank.py for core eviction logic and mtplx/server/openai.py for admission flags and scheduling constraints.
Core Admission Policies in MTPLX
MTPLX operates five distinct policies that form a layered defense against memory exhaustion. Each policy targets a specific stage of the request lifecycle, from system-wide pressure detection to per-request prefill validation.
Memory-Pressure Guard (System-Wide Eviction)
The memory-pressure guard responds to macOS system-wide pressure reports, such as when the operating system begins swapping memory. Located at lines 2786–2809 in mtplx/session_bank.py, this policy invokes the shrink_to_bytes() method to perform a global eviction pass.
The method accepts a target_bytes parameter and an optional protect_active flag. When protect_active=True, the guard preserves entries belonging to active sessions, evicting only the least-recently-used (LRU) idle entries until the memory budget is satisfied. This prevents mid-run re-prefill stalls while respecting system-level memory constraints.
Prefill Admission Shed (Targeted Eviction)
Before executing a memory-heavy prefill operation, the prefill admission shed validates whether the projected request exceeds the sustained-pressure limit. Implemented at lines 2840–2865 in mtplx/session_bank.py, this policy triggers shrink_for_admission() to perform escalating eviction.
The procedure first removes non-terminal entries, then falls back to a generic LRU pass if the deficit remains. Critical to this logic is the protect_tokens parameter, which accepts a list of token IDs that must remain intact during the shed. This ensures that tokens required for the incoming prompt are not evicted during the admission check.
Prefill Admission Chain Shed (Per-Request Ceiling)
For deployments requiring stricter isolation, the prefill admission chain shed enforces a per-request memory ceiling. This policy is controlled via the server flag _prefill_admission_chain_shed_enabled, defined around lines 17432–17450 in mtplx/server/openai.py.
When enabled, the server calls shrink_for_admission() with a larger target derived from the admission chain budget. This creates a secondary validation layer that runs in coordination with the standard prefill shed, providing defense-in-depth for multi-tenant or resource-constrained environments.
Hyper-Mode Admission Cap (Hard Throttling)
When the scheduler operates in hyper mode (--scheduler-mode hyper), MTPLX enforces a hard admission cap of exactly one concurrent request. This simplification is implemented at lines 2752–2768 in mtplx/server/openai.py.
The server validates launch flags and rejects configurations that would permit multi-admission semantics. By limiting active sessions to one, the memory-budget logic avoids complex multi-tenant arbitration, making eviction decisions deterministic and lightweight.
Dynamic-Ceiling Policy (Active Session Protection)
The dynamic-ceiling policy allows the memory-pressure guard to tolerate temporary overages when evicting entries would harm active computations. Located within the protect_active=True branch of shrink_to_bytes() at lines 2790–2804, this logic permits the bank to remain above target_bytes if the only remaining entries belong to the active session.
This policy trades strict adherence to the memory budget for inference latency stability, ensuring that active sessions complete without expensive re-prefill operations.
Implementation Workflow
The admission policies execute in a deterministic cascade to handle memory pressure gracefully:
- System pressure detection triggers
shrink_to_bytes()for global LRU eviction of idle sessions. - Prefill projection invokes
shrink_for_admission()to enforce targeted eviction of non-terminal entries before admitting new requests. - Admission chain validation optionally applies a larger eviction target when
_prefill_admission_chain_shed_enabledis active. - Hyper-mode enforcement overrides all previous logic with a hard cap of one admitted request, simplifying the eviction calculus.
Practical Code Examples
The following examples demonstrate how to interact with the admission policies programmatically:
# Example: manually invoking the memory-pressure guard
from mtplx.session_bank import SessionBank
bank = SessionBank(...)
# Target to keep the bank under 2 GiB
evicted = bank.shrink_to_bytes(target_bytes=2 * 1024**3, protect_active=True)
print(f"Evicted {evicted} entries to respect memory pressure")
# Example: admission-shed before a large prefill
# Assume the upcoming request needs 1 GiB more than the current budget
non_term, term = bank.shrink_for_admission(
target_bytes=bank.total_nbytes + 1 * 1024**3,
protect_tokens=[123, 456, 789], # tokens that must stay intact
reason="prefill_admission_chain",
)
print(f"Admission shed removed {non_term} non-terminal and {term} terminal entries")
Summary
- Memory-pressure guard (
shrink_to_bytes): Handles macOS system-wide pressure via global LRU eviction with active-session protection. - Prefill admission shed (
shrink_for_admission): Performs escalating eviction (non-terminal first) before admitting memory-heavy prefills. - Admission chain shed: Enforces per-request ceilings via server flags in
mtplx/server/openai.py. - Hyper-mode cap: Restricts concurrency to one request, eliminating multi-tenant memory arbitration complexity.
- Dynamic-ceiling policy: Allows temporary overages to protect active sessions from costly re-prefill operations.
Frequently Asked Questions
What triggers the memory-pressure guard in MTPLX?
The memory-pressure guard triggers when macOS reports system-wide memory pressure, typically when the OS begins swapping. According to the source code in mtplx/session_bank.py, this invokes shrink_to_bytes() to evict LRU entries until the target_bytes budget is met, optionally protecting active sessions via the protect_active parameter.
How does shrink_for_admission differ from shrink_to_bytes?
While shrink_to_bytes performs a global LRU eviction suitable for system-pressure events, shrink_for_admission implements an escalating strategy specifically for prefill-time decisions. As implemented at lines 2840–2865, it first attempts to remove non-terminal entries before falling back to LRU eviction, and it accepts a protect_tokens list to preserve specific tokens required for the incoming request.
What is the hyper-mode admission cap and when should I use it?
The hyper-mode admission cap enforces a maximum of one concurrent request when the server runs with --scheduler-mode hyper. This mode, enforced at lines 2752–2768 in mtplx/server/openai.py, simplifies memory management by eliminating multi-session arbitration overhead. Use this mode when serving large models on single-GPU systems where memory fragmentation risks outweigh concurrency benefits.
How can I enable the admission chain shed for stricter memory isolation?
Enable the admission chain shed by setting the _prefill_admission_chain_shed_enabled flag in the server configuration, located around lines 17432–17450 in mtplx/server/openai.py. When active, this policy invokes shrink_for_admission() with an admission-chain-derived budget target, adding a secondary validation layer before accepting memory-intensive requests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →