How to Run the MLX-VLM FastAPI Server: A Complete Setup Guide
Start the MLX-VLM FastAPI server using the mlx_vlm.server CLI command with the --model flag to preload a vision-language model, then access OpenAI-compatible endpoints at http://localhost:8080/chat/completions.
The MLX-VLM repository by Blaizzy provides a production-ready FastAPI server that exposes Apple Silicon-optimized vision-language models (VLMs) through an OpenAI-compatible HTTP API. This server implementation in mlx_vlm/server.py wraps the core inference utilities and manages model caching, vision feature caching, and streaming generation for both text and multimodal inputs.
Installation and Prerequisites
Before you can run the MLX-VLM FastAPI server, install the package from PyPI. The server requires Python 3.9+ and runs on Apple Silicon (M1/M2/M3) or x86 Macs with MLX support.
pip install -U mlx-vlm
The installation registers the mlx_vlm.server console script entry point, which maps to the main() function in mlx_vlm/server.py (lines 1444–1464). Alternatively, invoke the module directly using python -m mlx_vlm.server.
Starting the MLX-VLM FastAPI Server
The server supports both lazy loading (model loads on first request) and preloading (model loads at startup). Preloading is recommended for production to avoid latency on the first inference request.
Basic Startup Command
mlx_vlm.server \
--model mlx-community/Qwen2-VL-2B-Instruct-4bit \
--host 0.0.0.0 \
--port 8080 \
--trust-remote-code
Key parameters include:
--model: Hugging Face model path or local directory to preload via thePRELOAD_MODELenvironment variable logic--hostand--port: Bind address (default:0.0.0.0:8080)--adapter: Optional LoRA adapter path (PRELOAD_ADAPTER)--trust-remote-code: Required for models executing custom architecture code
Environment Variable Configuration
Instead of CLI flags, you can set environment variables before starting the server:
export PRELOAD_MODEL="mlx-community/Qwen2-VL-2B-Instruct-4bit"
export PRELOAD_ADAPTER="path/to/adapter"
mlx_vlm.server --port 8080
The lifespan() async context manager in mlx_vlm/server.py (lines 88–101) handles startup model loading when these variables are present and ensures cleanup on shutdown.
Core Architecture and Implementation
Understanding the server architecture helps optimize deployment and debug issues.
FastAPI Application Structure
The server creates a FastAPI instance with custom lifespan management:
# From mlx_vlm/server.py lines 14-22
app = FastAPI(
title="MLX-VLM Server",
description="OpenAI-compatible API for vision-language models",
lifespan=lifespan
)
The lifespan() function manages the application lifecycle, loading models into the model_cache dictionary at startup and clearing the VisionFeatureCache on shutdown.
Model Caching Mechanism
The server maintains a module-level cache to avoid reloading models between requests:
model_cache: Dictionary storing the loaded model, processor, config, andVisionFeatureCacheget_cached_model(): Retrieves or initializes the model (lines 22–30)unload_model_sync(): Synchronously removes the model from memory
The VisionFeatureCache (instantiated in mlx_vlm/vision_cache.py) stores image embeddings in an LRU cache to avoid re-encoding identical images across multiple API calls.
Uvicorn Integration
The CLI entry point invokes uvicorn programmatically:
# From mlx_vlm/server.py lines 1444-1464
def main():
parser = argparse.ArgumentParser()
# ... argument parsing ...
uvicorn.run("mlx_vlm.server:app", host=args.host, port=args.port)
This allows passing configuration directly from parsed CLI arguments to the ASGI server.
API Endpoints and Usage Examples
The server exposes OpenAI-compatible endpoints defined in mlx_vlm/server.py (lines 510–545).
Chat Completions Endpoint
POST /chat/completions supports multimodal inputs (text, images, audio) and streaming via Server-Sent Events (SSE).
Non-streaming request:
curl -X POST http://localhost:8080/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mlx-community/Qwen2-VL-2B-Instruct-4bit",
"messages": [{"role": "user", "content": "Describe this image"}],
"max_tokens": 200
}'
Streaming request with image:
curl -X POST http://localhost:8080/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mlx-community/Qwen2-VL-2B-Instruct-4bit",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "input_image", "image_url": "https://example.com/photo.jpg"}
]
}],
"stream": true,
"max_tokens": 150
}'
Streaming responses use the ChatStreamChunk model, returning SSE data lines containing JSON-encoded tokens.
Management Endpoints
GET /health: Returns server status and currently loaded modelPOST /unload: Unloads the model frommodel_cache, freeing GPU/CPU memoryGET /models: Lists available locally cached Hugging Face models
Health check example:
curl http://localhost:8080/health
# Response: {"status":"healthy","loaded_model":"mlx-community/Qwen2-VL-2B-Instruct-4bit","loaded_adapter":null}
Unload model:
curl -X POST http://localhost:8080/unload
Python Client Example
Use the OpenAI SDK to interact with the server:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="any")
response = client.chat.completions.create(
model="mlx-community/Qwen2-VL-2B-Instruct-4bit",
messages=[{"role": "user", "content": "Explain quantum computing"}],
max_tokens=200,
stream=False
)
print(response.choices[0].message.content)
Advanced Configuration Options
The server supports performance tuning through KV-cache quantization and memory management flags.
KV-Cache Quantization
Reduce memory usage during inference:
mlx_vlm.server \
--model mlx-community/Qwen3.5-4B-4bit \
--kv-bits 3.5 \
--kv-quant-scheme turboquant \
--max-kv-size 4096 \
--prefill-step-size 512
--kv-bits: Quantization bits for key-value cache (e.g.,3.5,4,8)--kv-quant-scheme: Quantization algorithm (turboquantoraffine)--max-kv-size: Maximum KV cache entries (limits context window memory)--prefill-step-size: Processing chunk size for prompt encoding
These parameters are passed directly to the underlying generate() and stream_generate() functions in mlx_vlm/generate.py.
Summary
- Installation: Install via
pip install mlx-vlmto get themlx_vlm.serverCLI - Startup: Run
mlx_vlm.server --model <path>to start the FastAPI application on port 8080 - Architecture: The server uses a
model_cachedictionary andVisionFeatureCacheto optimize inference performance - API: OpenAI-compatible endpoints at
/chat/completionssupport streaming, images, and audio - Management: Use
/healthto monitor status and/unloadto free memory - Optimization: Configure KV-cache quantization using
--kv-bitsand--kv-quant-schemefor efficient memory usage
Frequently Asked Questions
How do I preload a model when starting the MLX-VLM FastAPI server?
Pass the --model flag followed by a Hugging Face model ID or local path. The lifespan() function in mlx_vlm/server.py loads this model into the model_cache during startup, eliminating the cold-start latency on the first request. You can also set the PRELOAD_MODEL environment variable instead of using the CLI flag.
What is the difference between the /chat/completions and /responses endpoints?
The /chat/completions endpoint implements the modern OpenAI Chat API format and supports multimodal inputs (text, images, audio), streaming via SSE, and tool calling. The /responses endpoint provides an older response format primarily for backward compatibility. New integrations should use /chat/completions as defined in mlx_vlm/server.py lines 510–545.
How does the server handle repeated image inputs efficiently?
The server creates a VisionFeatureCache instance when loading a model, which stores image embeddings in an LRU cache. When the same image URL or base64 string appears in subsequent requests, the server retrieves cached embeddings from mlx_vlm/vision_cache.py instead of re-running the vision encoder, significantly reducing latency for repeated queries.
Can I run the server without preloading a model?
Yes. If you omit the --model flag, the server starts without loading a model into memory. The first incoming request to /chat/completions or /responses triggers lazy loading via get_cached_model(). While this saves initial memory, it causes a delay on the first request while the model downloads (if remote) or loads from disk.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →