How to Use the MLX Backend for Apple Silicon Training with Soup

Soup supports native Apple Silicon fine-tuning through its MLX backend by setting backend: mlx in your soup.yaml configuration file and installing the optional dependencies with pip install "soup-cli[mlx]".

The Soup CLI leverages Apple's open-source MLX framework to deliver Metal-accelerated training on M-series chips without CUDA dependencies. By configuring the MLX backend for Apple Silicon training, you can fine-tune quantized models locally while maintaining Soup's standard YAML-driven workflow.

Installing MLX Dependencies

Before training, install the MLX extras to pull in mlx and mlx-lm:

pip install "soup-cli[mlx]"

This installation provides the core libraries required by the MLXSFTTrainerWrapper located in src/soup_cli/trainer/mlx_sft.py.

Configuring the MLX Backend in soup.yaml

The MLX backend is activated by setting backend: mlx in your configuration. Note that this backend only supports the Supervised Fine-Tuning (SFT) task; attempting to use DPO or GRPO will trigger a validation error in src/soup_cli/config/schema.py.

base: mlx-community/Llama-3.2-3B-Instruct-4bit
task: sft
backend: mlx        # Apple Silicon only

data:
  train: ./data/train.jsonl
  format: alpaca

training:
  epochs: 3
  lr: 2e-5
  lora:
    r: 16
    alpha: 32

When backend: mlx is specified, Soup routes training to the MLXSFTTrainerWrapper class in src/soup_cli/trainer/mlx_sft.py, bypassing the standard CUDA-based trainers.

Internal Architecture and Model Loading

The MLX trainer follows a distinct execution path optimized for Metal performance:

  • Model Loading: The load_mlx_model function in src/soup_cli/utils/mlx.py loads 4-bit quantized models from the mlx-community Hugging Face organization. Attempts to request 8-bit quantization are ignored with a warning, as the backend only supports 4-bit MLX checkpoints.
  • LoRA Injection: After freezing the base model, the trainer applies LoRA layers via mlx_lm.tuner.utils.linear_to_lora_layers. The resolve_mlx_target_keys function maps target_modules: auto to the default Q and V projection keys (self_attn.q_proj, self_attn.v_proj).
  • Dataset Preparation: Training data is handled by CacheDataset, built using mlx_lm.tuner.datasets.create_dataset for efficient memory utilization on unified memory architectures.

Resuming Training from Checkpoints

The MLX backend supports resuming via adapter weights only. When using --resume auto, Soup detects MLX-style checkpoint files named NNNNNNN_adapters.safetensors and warm-starts the LoRA weights. Unlike CUDA backends, resume does not restore optimizer state—only the adapter parameters are reloaded.

soup train --resume auto

Device detection logic in src/soup_cli/utils/gpu.py treats backend: mlx as a first-class device, bypassing VRAM pre-flight checks used for NVIDIA GPUs.

Current Limitations

Understanding these constraints ensures successful deployment:

  • Task Restrictions: Only task: sft is permitted. The src/soup_cli/trainer/mlx_dpo.py file explicitly raises errors for DPO and GRPO attempts, and src/soup_cli/config/schema.py validates backend-task compatibility at load time.
  • RNG Seeding: MLX utilizes its own random number generator (mx.random). Any training.seed or training.data_seed values in your configuration are silently ignored, and Soup prints a warning to stdout.
  • Quantization: The trainer enforces 4-bit quantization. The validation logic in mlx_sft.py checks quantization settings and warns if 8-bit is requested.

Summary

  • Install MLX support via pip install "soup-cli[mlx]"
  • Set backend: mlx in soup.yaml with task: sft for Apple Silicon training
  • The MLXSFTTrainerWrapper in src/soup_cli/trainer/mlx_sft.py orchestrates model loading, LoRA application, and training
  • Use mlx-community 4-bit quantized models for optimal performance
  • Resume training detects *_adapters.safetensors files automatically
  • DPO, GRPO, and custom RNG seeds are currently unsupported

Frequently Asked Questions

Which Apple Silicon chips work with the Soup MLX backend?

Any Mac with Apple Silicon—including M1, M2, and M3 series chips—supports the MLX backend. The detect_device() function in src/soup_cli/utils/gpu.py automatically recognizes these devices and configures Metal acceleration without manual intervention.

Can I use DPO or GRPO training with the MLX backend?

No. The MLX backend currently restricts training to supervised fine-tuning (SFT) only. If you specify task: dpo or task: grpo with backend: mlx, the configuration validator in src/soup_cli/config/schema.py rejects the job, and src/soup_cli/trainer/mlx_dpo.py raises a runtime guard error.

How do I resume training from an MLX checkpoint?

Use the --resume auto flag when calling soup train. The trainer scans the output directory for files matching the pattern NNNNNNN_adapters.safetensors (where N represents digits) and loads only the LoRA adapter weights. Optimizer states are not preserved in MLX checkpoints, so training resumes with fresh optimizer initialization.

Does the MLX backend support full fine-tuning or only LoRA?

The current implementation in src/soup_cli/trainer/mlx_sft.py supports LoRA fine-tuning exclusively. The linear_to_lora_layers utility from mlx_lm injects trainable low-rank matrices into frozen base models, and there is no code path for full parameter updates in the MLX trainer wrapper.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →