How to Use LoRA Adapters with MLX Omni Server: A Complete Guide

MLX Omni Server supports LoRA adapters for image generation by accepting lora-paths and lora-scales as extra fields in the OpenAI-compatible images endpoint, automatically forwarding them to the MFlux backend.

The mlx-omni-server repository provides an OpenAI-compatible API server for Apple's MLX framework, enabling local inference for both language and diffusion models. When generating images through the MFlux backend, you can apply Low-Rank Adaptation (LoRA) fine-tuned weights to customize output styles without retraining the base model. This guide explains how to pass LoRA parameters through the standard OpenAI images generation interface.

How LoRA Adapters Work in MLX Omni Server

The server extends the standard OpenAI image generation schema to accept additional parameters specific to the MLX ecosystem. When you send a request to the /v1/images/generations endpoint, the ImageGenerationRequest model in src/mlx_omni_server/images/schema.py captures both standard fields (prompt, size, model) and custom LoRA configurations.

The implementation relies on Pydantic's extra = "allow" configuration (lines 30-60), which permits arbitrary additional fields beyond the OpenAI specification. The get_extra_params() method (lines 47-60) then extracts these custom fields—specifically lora-paths and lora-scales—and returns them as a dictionary for downstream processing.

In src/mlx_omni_server/images/images_service.py, the ImagesService class forwards these parameters to MFluxImageGenerator._get_flux(). At lines 71-73, the method passes the LoRA arguments directly to the Flux1 constructor from the mflux library, which loads the adapter weights and applies them during the diffusion process.

Required Parameters for LoRA Configuration

To activate LoRA adapters, include two specific fields in your JSON payload. These fields are not part of the standard OpenAI API but are recognized by the MLX Omni Server's MFlux integration.

lora-paths

The lora-paths field accepts a list of filesystem paths or URLs pointing to LoRA weight files (typically .safetensors format). Each path represents a distinct fine-tuned adapter that modifies the base model's behavior. For example:

"lora-paths": [
    "/Users/models/loras/anime_style.safetensors",
    "/Users/models/loras/vibrant_colors.safetensors"
]

lora-scales

The lora-scales field provides a list of floating-point values that control the strength of each corresponding adapter in lora-paths. Values typically range from 0.0 (no effect) to 1.0 (full strength), though some adapters may work with values up to 2.0. The scale list must match the length of the paths list:

"lora-scales": [0.8, 0.5]

Implementation Details in the Source Code

The parameter flow follows a specific pipeline through the server's architecture. First, src/mlx_omni_server/images/schema.py defines the request validation:

class ImageGenerationRequest(BaseModel):
    model_config = ConfigDict(extra="allow")  # Allows lora-paths and lora-scales

    
    prompt: str
    model: Optional[str] = None
    n: Optional[int] = 1
    # ... standard OpenAI fields

    
    def get_extra_params(self) -> dict:
        # Returns only the non-standard fields including LoRA configs

        return self.model_dump(exclude={"prompt", "model", "n", "size", "response_format"})

Next, src/mlx_omni_server/images/images_service.py handles the integration with the underlying generation engine:

def _get_flux(self, model_name: str, extra_params: dict):
    # Extract LoRA configurations from extra_params

    lora_paths = extra_params.get("lora-paths", [])
    lora_scales = extra_params.get("lora-scales", [])
    
    # Initialize Flux1 with LoRA support

    flux = Flux1(
        model_config=model_name,
        lora_paths=lora_paths,      # Lines 71-73 in source

        lora_scales=lora_scales
    )
    return flux

The Flux1 class from the mflux library then manages the actual weight merging and inference-time application of these adapters.

Complete API Request Example

The following Python client demonstrates how to generate images with multiple LoRA adapters applied. This example targets the default server address and uses the MFlux backend with quantized FLUX models:

import requests
import json

url = "http://localhost:8000/v1/images/generations"

payload = {
    "prompt": "A futuristic cityscape at sunset, photorealistic",
    "size": "1024x1024",
    "model": "dhairyashil/FLUX.1-schnell-mflux-4bit",
    # LoRA adapter configuration

    "lora-paths": [
        "/path/to/lora_adapter_1.safetensors",
        "/path/to/lora_adapter_2.safetensors"
    ],
    "lora-scales": [0.8, 0.5],
    # Optional MFlux-specific parameters

    "steps": 8,
    "guidance": 5.0,
    "seed": 42
}

headers = {"Content-Type": "application/json"}
response = requests.post(url, data=json.dumps(payload), headers=headers)

if response.status_code == 200:
    data = response.json()
    # Response contains base64-encoded images or URLs depending on response_format

    print(json.dumps(data, indent=2))
else:
    print(f"Error {response.status_code}: {response.text}")

You can also test this using curl from the command line:

curl -X POST http://localhost:8000/v1/images/generations \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "Cyberpunk portrait of a samurai",
    "model": "dhairyashil/FLUX.1-schnell-mflux-4bit",
    "lora-paths": ["/models/cyberpunk_style.safetensors"],
    "lora-scales": [0.9],
    "steps": 4
  }'

Summary

  • MLX Omni Server accepts LoRA adapters through the standard OpenAI images endpoint by allowing extra fields in the request schema.
  • Configuration requires two fields: lora-paths (list of file paths) and lora-scales (list of strength values).
  • Implementation location: src/mlx_omni_server/images/schema.py handles parameter extraction, while src/mlx_omni_server/images/images_service.py forwards values to the MFlux Flux1 constructor.
  • File format: Standard .safetensors files containing LoRA weights are supported.
  • Multiple adapters: You can chain several LoRA layers by providing parallel lists of paths and their corresponding scale factors.

Frequently Asked Questions

What file formats are supported for LoRA weights?

MLX Omni Server supports LoRA weights in the Safetensors format (.safetensors), which is the standard for MLX and PyTorch model distribution. The server passes these file paths directly to the underlying mflux library, which handles the actual tensor loading and merging with the base FLUX model weights.

Can I use multiple LoRA adapters simultaneously?

Yes, you can apply multiple LoRA layers by providing a list of paths in lora-paths and corresponding scales in lora-scales. According to the implementation in src/mlx_omni_server/images/images_service.py, both lists are forwarded to the Flux1 constructor, which composites the adaptations. Ensure both arrays have matching lengths; otherwise, the MFlux backend may raise a configuration error.

Do LoRA parameters work with the chat completions endpoint?

No, the lora-paths and lora-scales parameters are currently implemented only for the images generation endpoint (/v1/images/generations) when using the MFlux backend. The chat completion endpoints in src/mlx_omni_server/chat/ follow different model loading patterns and do not expose these specific fields for LLM inference.

How do I determine the correct scale values for my LoRA adapters?

Start with a scale of 1.0 for single adapters, then adjust downward (0.5-0.8) if the style is too strong or upward (1.2-1.5) for subtle adapters. When combining multiple LoRAs in mlx-omni-server, ensure the cumulative effect does not saturate the model; the total sum of scales generally should not exceed 2.0 to maintain generation stability. Experimentation is essential as optimal values depend on the specific adapter training and the base model quantization level.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →