How to Optimize Memory Usage for Large Language Models in MLX-Omni-Server

MLX-Omni-Server implements a multi-layered caching architecture combining LRU + TTL model wrapper caching, token-level prompt caching, and configurable KV-cache constraints to minimize VRAM consumption when serving LLMs on Apple Silicon.

MLX-Omni-Server provides a unified API for running large language models like Qwen-3 and LLaMA-3 on Apple Silicon via the mlx-lm library. Because these models can occupy several gigabytes of VRAM, the server employs sophisticated memory optimization techniques that prevent redundant model loading and eliminate repeated prompt tokenization. Understanding how to configure these caching layers is essential to optimize memory usage for large language models while maintaining low latency and high throughput.

Implement LRU and TTL Model Wrapper Caching

The primary defense against memory bloat is the wrapper cache, a singleton instance shared across all API endpoints. Located in src/mlx_omni_server/chat/mlx/wrapper_cache.py, this cache stores fully-initialized ChatGenerator instances (model, tokenizer, and optional adapters) to avoid reloading the same model multiple times.

The implementation combines two eviction strategies to keep VRAM usage bounded:

  • LRU Eviction: When the cache reaches max_size (default 3), the least-recently accessed model is removed (lines 92-106 of wrapper_cache.py).
  • TTL Eviction: A background daemon thread automatically clears items older than ttl_seconds (default 300 seconds) to prevent stale entries from consuming VRAM (lines 70-87).

All cache mutations are protected by threading.Lock, ensuring thread-safe access in multi-client environments.

Adjusting Cache Limits

Modify cache boundaries at runtime to balance memory capacity against model diversity:

from mlx_omni_server.chat.mlx.wrapper_cache import wrapper_cache

# Retain up to 5 models, each persisting for 10 minutes

wrapper_cache.set_max_size(5)
wrapper_cache._ttl_seconds = 600

(Source: lines 88-107, 118-124 of wrapper_cache.py)

Reuse Tokenized Prompts with Prompt Caching

The prompt cache in src/mlx_omni_server/chat/mlx/prompt_cache.py eliminates redundant tokenization and KV-cache recomputation. When enable_prompt_cache=True in a generation request, the server:

  1. Tokenizes the prompt once using tokenizer.encode(prompt).
  2. Looks up cached token sequences via PromptCache.get_prompt_cache.
  3. Feeds cached KV-cache entries to mlx_lm.generate via the prompt_cache argument, drastically reducing first-token latency.

This mechanism is integrated in ChatGenerator.generate_stream (lines 29-33 and 41-44 of chat_generator.py).

Enabling Prompt Caching

Activate prompt caching in your generation calls to skip redundant computation:

response = chat_generator.generate(
    messages=[{"role": "user", "content": "Explain quantum tunneling"}],
    enable_prompt_cache=True,
    max_tokens=256,
)

Configure Memory-Efficient Generation Parameters

Fine-tune VRAM usage per request through parameters passed via ChatGenerator._create_mlx_kwargs (lines 310-311 of chat_generator.py). These arguments translate directly into mlx-lm memory controls:

Parameter Memory Impact
max_kv_size Caps the KV-cache size in tokens; smaller values reduce memory but may truncate long contexts.
kv_bits Quantizes KV-cache entries to lower precision (e.g., 8-bit vs 16-bit), halving cache memory requirements.

Applying KV Constraints

Restrict memory footprint for specific high-throughput scenarios:

response = chat_generator.generate(
    messages=[{"role": "user", "content": "Summarize this document"}],
    max_kv_size=1024,   # Limit KV cache to 1024 tokens

    kv_bits=8,          # Use 8-bit precision for cache entries

    max_tokens=512,
)

Utilize Global Cache Access Patterns

Instead of manually instantiating generators, use ChatGenerator.get_or_create to automatically leverage the wrapper cache. According to the MLX-Omni-Server source code, this method checks for existing instances before loading new models, ensuring optimal memory reuse.

from mlx_omni_server.chat.mlx.chat_generator import ChatGenerator

# First call loads the model; subsequent calls return cached instance

generator = ChatGenerator.get_or_create("mlx-community/Qwen3-0.6B-4bit")
result = generator.generate(messages=[{"role": "user", "content": "Hello"}])

(Implementation: lines 30-38 of chat_generator.py)

Monitor Cache Metrics in Real Time

Inspect current cache utilization via wrapper_cache.get_cache_info() (lines 250-285 of wrapper_cache.py) to diagnose memory pressure:

info = wrapper_cache.get_cache_info()
print(info)

# {'cache_size': 2, 'max_size': 3, 'ttl_seconds': 300,

#  'cached_keys': [...], 'lru_order': [...], 'ttl_info': [...]}

Summary

  • Wrapper caching with LRU and TTL eviction prevents redundant model loading, keeping only hot models in VRAM (default maximum: 3).
  • Prompt caching eliminates repeated tokenization and enables KV-cache reuse between identical prompts, reducing compute and memory churn.
  • Generation parameters (max_kv_size, kv_bits) provide per-request control over KV-cache memory footprint.
  • Thread-safe design ensures cache consistency across concurrent API requests via threading.Lock protection in wrapper_cache.py.
  • Runtime monitoring via get_cache_info() enables dynamic adjustment of cache limits based on observed usage patterns.

Frequently Asked Questions

What is the default cache size limit in MLX-Omni-Server?

The wrapper cache defaults to retaining 3 model instances simultaneously. When this limit is reached, the least-recently used model is evicted from VRAM to make room for new requests, as implemented in lines 92-106 of wrapper_cache.py.

How does prompt caching reduce memory usage?

Prompt caching reduces memory pressure by avoiding redundant tokenization and enabling reuse of pre-computed KV-cache entries for identical prompts. When enabled via enable_prompt_cache=True, subsequent requests skip the expensive initial forward pass, reducing both latency and transient memory spikes during token processing.

Can I adjust cache settings without restarting the server?

Yes. The wrapper cache exposes runtime configuration methods including set_max_size() and the _ttl_seconds attribute. These modifications take effect immediately for subsequent model loading operations, allowing dynamic scaling of memory resources based on current demand without server restarts.

What happens when the KV-cache exceeds max_kv_size?

When max_kv_size is specified in chat_generator.py, the KV-cache is truncated to the specified token limit. This constraint reduces VRAM consumption but may cause the model to lose context for earlier tokens in long conversations, potentially degrading output quality for extended dialogues.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →