# How to Use LoRA Adapters with MLX Omni Server: A Complete Guide

> Learn to integrate LoRA adapters with MLX Omni Server for advanced image generation. This guide explains how to use Lora paths and scales with the OpenAI compatible images endpoint.

- Repository: [madroid/mlx-omni-server](https://github.com/madroidmaq/mlx-omni-server)
- Tags: how-to-guide
- Published: 2026-03-06

---

**MLX Omni Server supports LoRA adapters for image generation by accepting `lora-paths` and `lora-scales` as extra fields in the OpenAI-compatible images endpoint, automatically forwarding them to the MFlux backend.**

The **mlx-omni-server** repository provides an OpenAI-compatible API server for Apple's MLX framework, enabling local inference for both language and diffusion models. When generating images through the MFlux backend, you can apply Low-Rank Adaptation (LoRA) fine-tuned weights to customize output styles without retraining the base model. This guide explains how to pass LoRA parameters through the standard OpenAI images generation interface.

## How LoRA Adapters Work in MLX Omni Server

The server extends the standard OpenAI image generation schema to accept additional parameters specific to the MLX ecosystem. When you send a request to the `/v1/images/generations` endpoint, the `ImageGenerationRequest` model in [`src/mlx_omni_server/images/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/images/schema.py) captures both standard fields (prompt, size, model) and custom LoRA configurations.

The implementation relies on Pydantic's `extra = "allow"` configuration (lines 30-60), which permits arbitrary additional fields beyond the OpenAI specification. The `get_extra_params()` method (lines 47-60) then extracts these custom fields—specifically `lora-paths` and `lora-scales`—and returns them as a dictionary for downstream processing.

In [`src/mlx_omni_server/images/images_service.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/images/images_service.py), the `ImagesService` class forwards these parameters to `MFluxImageGenerator._get_flux()`. At lines 71-73, the method passes the LoRA arguments directly to the **Flux1** constructor from the `mflux` library, which loads the adapter weights and applies them during the diffusion process.

## Required Parameters for LoRA Configuration

To activate LoRA adapters, include two specific fields in your JSON payload. These fields are not part of the standard OpenAI API but are recognized by the MLX Omni Server's MFlux integration.

### lora-paths

The `lora-paths` field accepts a list of filesystem paths or URLs pointing to LoRA weight files (typically `.safetensors` format). Each path represents a distinct fine-tuned adapter that modifies the base model's behavior. For example:

```json
"lora-paths": [
    "/Users/models/loras/anime_style.safetensors",
    "/Users/models/loras/vibrant_colors.safetensors"
]

```

### lora-scales

The `lora-scales` field provides a list of floating-point values that control the strength of each corresponding adapter in `lora-paths`. Values typically range from 0.0 (no effect) to 1.0 (full strength), though some adapters may work with values up to 2.0. The scale list must match the length of the paths list:

```json
"lora-scales": [0.8, 0.5]

```

## Implementation Details in the Source Code

The parameter flow follows a specific pipeline through the server's architecture. First, [`src/mlx_omni_server/images/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/images/schema.py) defines the request validation:

```python
class ImageGenerationRequest(BaseModel):
    model_config = ConfigDict(extra="allow")  # Allows lora-paths and lora-scales

    
    prompt: str
    model: Optional[str] = None
    n: Optional[int] = 1
    # ... standard OpenAI fields

    
    def get_extra_params(self) -> dict:
        # Returns only the non-standard fields including LoRA configs

        return self.model_dump(exclude={"prompt", "model", "n", "size", "response_format"})

```

Next, [`src/mlx_omni_server/images/images_service.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/images/images_service.py) handles the integration with the underlying generation engine:

```python
def _get_flux(self, model_name: str, extra_params: dict):
    # Extract LoRA configurations from extra_params

    lora_paths = extra_params.get("lora-paths", [])
    lora_scales = extra_params.get("lora-scales", [])
    
    # Initialize Flux1 with LoRA support

    flux = Flux1(
        model_config=model_name,
        lora_paths=lora_paths,      # Lines 71-73 in source

        lora_scales=lora_scales
    )
    return flux

```

The **Flux1** class from the `mflux` library then manages the actual weight merging and inference-time application of these adapters.

## Complete API Request Example

The following Python client demonstrates how to generate images with multiple LoRA adapters applied. This example targets the default server address and uses the MFlux backend with quantized FLUX models:

```python
import requests
import json

url = "http://localhost:8000/v1/images/generations"

payload = {
    "prompt": "A futuristic cityscape at sunset, photorealistic",
    "size": "1024x1024",
    "model": "dhairyashil/FLUX.1-schnell-mflux-4bit",
    # LoRA adapter configuration

    "lora-paths": [
        "/path/to/lora_adapter_1.safetensors",
        "/path/to/lora_adapter_2.safetensors"
    ],
    "lora-scales": [0.8, 0.5],
    # Optional MFlux-specific parameters

    "steps": 8,
    "guidance": 5.0,
    "seed": 42
}

headers = {"Content-Type": "application/json"}
response = requests.post(url, data=json.dumps(payload), headers=headers)

if response.status_code == 200:
    data = response.json()
    # Response contains base64-encoded images or URLs depending on response_format

    print(json.dumps(data, indent=2))
else:
    print(f"Error {response.status_code}: {response.text}")

```

You can also test this using `curl` from the command line:

```bash
curl -X POST http://localhost:8000/v1/images/generations \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "Cyberpunk portrait of a samurai",
    "model": "dhairyashil/FLUX.1-schnell-mflux-4bit",
    "lora-paths": ["/models/cyberpunk_style.safetensors"],
    "lora-scales": [0.9],
    "steps": 4
  }'

```

## Summary

- **MLX Omni Server** accepts LoRA adapters through the standard OpenAI images endpoint by allowing extra fields in the request schema.
- **Configuration requires two fields**: `lora-paths` (list of file paths) and `lora-scales` (list of strength values).
- **Implementation location**: [`src/mlx_omni_server/images/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/images/schema.py) handles parameter extraction, while [`src/mlx_omni_server/images/images_service.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/images/images_service.py) forwards values to the MFlux `Flux1` constructor.
- **File format**: Standard `.safetensors` files containing LoRA weights are supported.
- **Multiple adapters**: You can chain several LoRA layers by providing parallel lists of paths and their corresponding scale factors.

## Frequently Asked Questions

### What file formats are supported for LoRA weights?

MLX Omni Server supports LoRA weights in the **Safetensors** format (`.safetensors`), which is the standard for MLX and PyTorch model distribution. The server passes these file paths directly to the underlying `mflux` library, which handles the actual tensor loading and merging with the base FLUX model weights.

### Can I use multiple LoRA adapters simultaneously?

Yes, you can apply multiple LoRA layers by providing a list of paths in `lora-paths` and corresponding scales in `lora-scales`. According to the implementation in [`src/mlx_omni_server/images/images_service.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/images/images_service.py), both lists are forwarded to the **Flux1** constructor, which composites the adaptations. Ensure both arrays have matching lengths; otherwise, the MFlux backend may raise a configuration error.

### Do LoRA parameters work with the chat completions endpoint?

No, the `lora-paths` and `lora-scales` parameters are currently implemented **only for the images generation endpoint** (`/v1/images/generations`) when using the MFlux backend. The chat completion endpoints in `src/mlx_omni_server/chat/` follow different model loading patterns and do not expose these specific fields for LLM inference.

### How do I determine the correct scale values for my LoRA adapters?

Start with a scale of **1.0** for single adapters, then adjust downward (0.5-0.8) if the style is too strong or upward (1.2-1.5) for subtle adapters. When combining multiple LoRAs in `mlx-omni-server`, ensure the cumulative effect does not saturate the model; the total sum of scales generally should not exceed 2.0 to maintain generation stability. Experimentation is essential as optimal values depend on the specific adapter training and the base model quantization level.