How Prompt Caching Improves Generation Performance in MLX-Omni-Server

Prompt caching accelerates multi-turn chat generation by reusing the model's KV cache for unchanged prompt prefixes, eliminating redundant attention computations while maintaining identical output quality.

In the mlx-omni-server repository, prompt caching provides a critical performance optimization for conversational AI workloads. By storing intermediate transformer states across generation turns, the system avoids recomputing expensive attention layers for static portions of dialogue history, significantly reducing latency and resource consumption in long conversations.

How Prompt Caching Works in MLX-Omni-Server

KV-Cache Reuse and Common Prefix Detection

The mechanism centers on the KV-cache (key-value cache), which stores intermediate attention results as the model processes tokens. In src/mlx_omni_server/chat/mlx/prompt_cache.py, the PromptCache class manages this state by calculating the longest common prefix between the current prompt and previously cached tokens.

When a new request arrives, the system determines common_prefix_len to identify matching token sequences. If the conversation history remains static and only a new user message is appended, the model bypasses attention calculations for the entire prefix and resumes generation from the cached state.

Reduced Token Processing Overhead

During generation, the ChatGenerator.generate_stream method in src/mlx_omni_server/chat/mlx/chat_generator.py passes the tokenized prompt through self.prompt_cache.get_prompt_cache(...). This returns (processed_prompt, cached_tokens), where only the suffix (new tokens) requires actual forward passes through the model.

The cached_tokens count appears in generation statistics as cache_hit_tokens, providing direct visibility into optimization efficiency. This selective processing reduces both GPU/CPU compute cycles and memory bandwidth consumption.

Performance Benefits of Prompt Caching

Latency and Throughput Improvements

Prompt caching delivers measurable speedups specifically in conversational scenarios:

  • First-token latency remains equivalent to uncached generation initially, as the system loads the existing KV cache state
  • Subsequent-token latency drops significantly because the model skips heavy attention calculations for cached prefixes
  • Throughput increases substantially during long conversations where system prompts and prior turns remain unchanged

Automatic Cache Trimming

For models supporting the can_trim_prompt_cache capability, the system automatically trims the cache to the new prefix length rather than performing a full reset. This preserves maximum previously computed state while adapting to conversation changes, ensuring optimal memory utilization without manual intervention.

Implementing Prompt Caching in Your Applications

Enable the optimization using the enable_prompt_cache parameter:

from mlx_omni_server.chat.mlx.chat_generator import ChatGenerator

# Initialize the generator

generator = ChatGenerator.create("mlx-community/Qwen3-0.6B-4bit")

# First turn - builds the initial cache

response1 = generator.generate(
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Hello!"}
    ],
    enable_prompt_cache=True,
    max_tokens=512
)

# Second turn - automatically reuses cache for the system prompt

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "assistant", "content": response1.content.text},
    {"role": "user", "content": "What is the weather today?"}
]

response2 = generator.generate(
    messages=messages,
    enable_prompt_cache=True,
    max_tokens=512
)

Monitor cache efficiency in streaming mode:

stream = generator.generate_stream(
    messages=conversation_history,
    enable_prompt_cache=True,
    max_tokens=128
)

for result in stream:
    # Access cache hit statistics

    reused_count = result.stats.cache_hit_tokens
    print(f"Reused {reused_count} tokens from cache")

Verify caching behavior programmatically:

def verify_caching_performance():
    gen = ChatGenerator.create("mlx-community/Qwen3-0.6B-4bit")
    messages = [
        {"role": "system", "content": "You are a bot."},
        {"role": "user", "content": "Say hi"},
    ]

    # First call processes the full prompt

    first = gen.generate(messages=messages, enable_prompt_cache=True)
    assert first.stats.cache_hit_tokens == 0

    # Extend the conversation

    messages.append({"role": "assistant", "content": first.content.text})
    messages.append({"role": "user", "content": "How are you?"})

    # Second call reuses the system prompt tokens

    second = gen.generate(messages=messages, enable_prompt_cache=True)
    assert second.stats.cache_hit_tokens > 0

Summary

  • Prompt caching eliminates redundant computation by reusing KV-cache states for unchanged prompt prefixes in multi-turn conversations managed by mlx-omni-server.
  • The PromptCache class in src/mlx_omni_server/chat/mlx/prompt_cache.py detects common prefixes via common_prefix_len to minimize processing overhead.
  • Integration within ChatGenerator.generate_stream exposes cache statistics through cache_hit_tokens, enabling real-time performance monitoring.
  • Enabling enable_prompt_cache=True reduces GPU/CPU compute and memory bandwidth, particularly benefiting long conversational contexts with stable system prompts.
  • Automatic cache trimming via can_trim_prompt_cache preserves computational state without requiring manual cache management.

Frequently Asked Questions

What is the default behavior for prompt caching in MLX-Omni-Server?

By default, enable_prompt_cache is set to False, requiring the model to recompute the entire prompt for every generation turn. You must explicitly set enable_prompt_cache=True to activate the optimization and realize performance gains.

How does prompt caching affect first-token latency?

First-token latency remains approximately equivalent to uncached generation because the system must load and initialize the KV cache state. The primary performance improvements appear in subsequent token generation and overall throughput during extended multi-turn conversations.

Can I use prompt caching with any model architecture in MLX-Omni-Server?

Prompt caching works with models supported by the underlying MLX framework. However, the automatic trimming feature (can_trim_prompt_cache) depends on specific model implementations. Verify your specific model supports trimming by checking its configuration or the src/mlx_omni_server/chat/mlx/prompt_cache.py implementation details.

How can I verify that prompt caching is actually improving performance?

Inspect the cache_hit_tokens field in the generation statistics returned by generate_stream or generate. A value greater than zero indicates successful cache reuse, while the initial call in a conversation sequence should show zero cache hits. Compare processing times between cached and uncached turns to measure latency improvements.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →