# How to Optimize Memory Usage for Large Language Models in MLX-Omni-Server

> Optimize memory usage for large language models with MLX-Omni-Server. Discover multi-layered caching and KV-cache constraints for efficient LLM serving on Apple Silicon.

- Repository: [madroid/mlx-omni-server](https://github.com/madroidmaq/mlx-omni-server)
- Tags: performance
- Published: 2026-03-06

---

**MLX-Omni-Server implements a multi-layered caching architecture combining LRU + TTL model wrapper caching, token-level prompt caching, and configurable KV-cache constraints to minimize VRAM consumption when serving LLMs on Apple Silicon.**

MLX-Omni-Server provides a unified API for running large language models like Qwen-3 and LLaMA-3 on Apple Silicon via the **mlx-lm** library. Because these models can occupy several gigabytes of VRAM, the server employs sophisticated memory optimization techniques that prevent redundant model loading and eliminate repeated prompt tokenization. Understanding how to configure these caching layers is essential to optimize memory usage for large language models while maintaining low latency and high throughput.

## Implement LRU and TTL Model Wrapper Caching

The primary defense against memory bloat is the **wrapper cache**, a singleton instance shared across all API endpoints. Located in [`src/mlx_omni_server/chat/mlx/wrapper_cache.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/mlx/wrapper_cache.py), this cache stores fully-initialized `ChatGenerator` instances (model, tokenizer, and optional adapters) to avoid reloading the same model multiple times.

The implementation combines two eviction strategies to keep VRAM usage bounded:

- **LRU Eviction**: When the cache reaches `max_size` (default 3), the least-recently accessed model is removed (lines 92-106 of [`wrapper_cache.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/wrapper_cache.py)).
- **TTL Eviction**: A background daemon thread automatically clears items older than `ttl_seconds` (default 300 seconds) to prevent stale entries from consuming VRAM (lines 70-87).

All cache mutations are protected by `threading.Lock`, ensuring thread-safe access in multi-client environments.

### Adjusting Cache Limits

Modify cache boundaries at runtime to balance memory capacity against model diversity:

```python
from mlx_omni_server.chat.mlx.wrapper_cache import wrapper_cache

# Retain up to 5 models, each persisting for 10 minutes

wrapper_cache.set_max_size(5)
wrapper_cache._ttl_seconds = 600

```

*(Source: lines 88-107, 118-124 of [`wrapper_cache.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/wrapper_cache.py))*

## Reuse Tokenized Prompts with Prompt Caching

The **prompt cache** in [`src/mlx_omni_server/chat/mlx/prompt_cache.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/mlx/prompt_cache.py) eliminates redundant tokenization and KV-cache recomputation. When `enable_prompt_cache=True` in a generation request, the server:

1. Tokenizes the prompt once using `tokenizer.encode(prompt)`.
2. Looks up cached token sequences via `PromptCache.get_prompt_cache`.
3. Feeds cached KV-cache entries to `mlx_lm.generate` via the `prompt_cache` argument, drastically reducing first-token latency.

This mechanism is integrated in `ChatGenerator.generate_stream` (lines 29-33 and 41-44 of [`chat_generator.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/chat_generator.py)).

### Enabling Prompt Caching

Activate prompt caching in your generation calls to skip redundant computation:

```python
response = chat_generator.generate(
    messages=[{"role": "user", "content": "Explain quantum tunneling"}],
    enable_prompt_cache=True,
    max_tokens=256,
)

```

## Configure Memory-Efficient Generation Parameters

Fine-tune VRAM usage per request through parameters passed via `ChatGenerator._create_mlx_kwargs` (lines 310-311 of [`chat_generator.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/chat_generator.py)). These arguments translate directly into **mlx-lm** memory controls:

| Parameter | Memory Impact |
|-----------|---------------|
| `max_kv_size` | Caps the KV-cache size in tokens; smaller values reduce memory but may truncate long contexts. |
| `kv_bits` | Quantizes KV-cache entries to lower precision (e.g., 8-bit vs 16-bit), halving cache memory requirements. |

### Applying KV Constraints

Restrict memory footprint for specific high-throughput scenarios:

```python
response = chat_generator.generate(
    messages=[{"role": "user", "content": "Summarize this document"}],
    max_kv_size=1024,   # Limit KV cache to 1024 tokens

    kv_bits=8,          # Use 8-bit precision for cache entries

    max_tokens=512,
)

```

## Utilize Global Cache Access Patterns

Instead of manually instantiating generators, use `ChatGenerator.get_or_create` to automatically leverage the wrapper cache. According to the MLX-Omni-Server source code, this method checks for existing instances before loading new models, ensuring optimal memory reuse.

```python
from mlx_omni_server.chat.mlx.chat_generator import ChatGenerator

# First call loads the model; subsequent calls return cached instance

generator = ChatGenerator.get_or_create("mlx-community/Qwen3-0.6B-4bit")
result = generator.generate(messages=[{"role": "user", "content": "Hello"}])

```

*(Implementation: lines 30-38 of [`chat_generator.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/chat_generator.py))*

## Monitor Cache Metrics in Real Time

Inspect current cache utilization via `wrapper_cache.get_cache_info()` (lines 250-285 of [`wrapper_cache.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/wrapper_cache.py)) to diagnose memory pressure:

```python
info = wrapper_cache.get_cache_info()
print(info)

# {'cache_size': 2, 'max_size': 3, 'ttl_seconds': 300,

#  'cached_keys': [...], 'lru_order': [...], 'ttl_info': [...]}

```

## Summary

- **Wrapper caching** with LRU and TTL eviction prevents redundant model loading, keeping only hot models in VRAM (default maximum: 3).
- **Prompt caching** eliminates repeated tokenization and enables KV-cache reuse between identical prompts, reducing compute and memory churn.
- **Generation parameters** (`max_kv_size`, `kv_bits`) provide per-request control over KV-cache memory footprint.
- **Thread-safe design** ensures cache consistency across concurrent API requests via `threading.Lock` protection in [`wrapper_cache.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/wrapper_cache.py).
- **Runtime monitoring** via `get_cache_info()` enables dynamic adjustment of cache limits based on observed usage patterns.

## Frequently Asked Questions

### What is the default cache size limit in MLX-Omni-Server?

The wrapper cache defaults to retaining **3 model instances** simultaneously. When this limit is reached, the least-recently used model is evicted from VRAM to make room for new requests, as implemented in lines 92-106 of [`wrapper_cache.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/wrapper_cache.py).

### How does prompt caching reduce memory usage?

Prompt caching reduces memory pressure by avoiding redundant tokenization and enabling reuse of pre-computed **KV-cache entries** for identical prompts. When enabled via `enable_prompt_cache=True`, subsequent requests skip the expensive initial forward pass, reducing both latency and transient memory spikes during token processing.

### Can I adjust cache settings without restarting the server?

Yes. The wrapper cache exposes runtime configuration methods including `set_max_size()` and the `_ttl_seconds` attribute. These modifications take effect immediately for subsequent model loading operations, allowing dynamic scaling of memory resources based on current demand without server restarts.

### What happens when the KV-cache exceeds max_kv_size?

When `max_kv_size` is specified in [`chat_generator.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/chat_generator.py), the KV-cache is truncated to the specified token limit. This constraint reduces VRAM consumption but may cause the model to lose context for earlier tokens in long conversations, potentially degrading output quality for extended dialogues.