How to Use the MLX Backend for Apple Silicon Training with Soup
Soup supports native Apple Silicon fine-tuning through its MLX backend by setting backend: mlx in your soup.yaml configuration file and installing the optional dependencies with pip install "soup-cli[mlx]".
The Soup CLI leverages Apple's open-source MLX framework to deliver Metal-accelerated training on M-series chips without CUDA dependencies. By configuring the MLX backend for Apple Silicon training, you can fine-tune quantized models locally while maintaining Soup's standard YAML-driven workflow.
Installing MLX Dependencies
Before training, install the MLX extras to pull in mlx and mlx-lm:
pip install "soup-cli[mlx]"
This installation provides the core libraries required by the MLXSFTTrainerWrapper located in src/soup_cli/trainer/mlx_sft.py.
Configuring the MLX Backend in soup.yaml
The MLX backend is activated by setting backend: mlx in your configuration. Note that this backend only supports the Supervised Fine-Tuning (SFT) task; attempting to use DPO or GRPO will trigger a validation error in src/soup_cli/config/schema.py.
base: mlx-community/Llama-3.2-3B-Instruct-4bit
task: sft
backend: mlx # Apple Silicon only
data:
train: ./data/train.jsonl
format: alpaca
training:
epochs: 3
lr: 2e-5
lora:
r: 16
alpha: 32
When backend: mlx is specified, Soup routes training to the MLXSFTTrainerWrapper class in src/soup_cli/trainer/mlx_sft.py, bypassing the standard CUDA-based trainers.
Internal Architecture and Model Loading
The MLX trainer follows a distinct execution path optimized for Metal performance:
- Model Loading: The
load_mlx_modelfunction insrc/soup_cli/utils/mlx.pyloads 4-bit quantized models from themlx-communityHugging Face organization. Attempts to request 8-bit quantization are ignored with a warning, as the backend only supports 4-bit MLX checkpoints. - LoRA Injection: After freezing the base model, the trainer applies LoRA layers via
mlx_lm.tuner.utils.linear_to_lora_layers. Theresolve_mlx_target_keysfunction mapstarget_modules: autoto the default Q and V projection keys (self_attn.q_proj,self_attn.v_proj). - Dataset Preparation: Training data is handled by
CacheDataset, built usingmlx_lm.tuner.datasets.create_datasetfor efficient memory utilization on unified memory architectures.
Resuming Training from Checkpoints
The MLX backend supports resuming via adapter weights only. When using --resume auto, Soup detects MLX-style checkpoint files named NNNNNNN_adapters.safetensors and warm-starts the LoRA weights. Unlike CUDA backends, resume does not restore optimizer state—only the adapter parameters are reloaded.
soup train --resume auto
Device detection logic in src/soup_cli/utils/gpu.py treats backend: mlx as a first-class device, bypassing VRAM pre-flight checks used for NVIDIA GPUs.
Current Limitations
Understanding these constraints ensures successful deployment:
- Task Restrictions: Only
task: sftis permitted. Thesrc/soup_cli/trainer/mlx_dpo.pyfile explicitly raises errors for DPO and GRPO attempts, andsrc/soup_cli/config/schema.pyvalidates backend-task compatibility at load time. - RNG Seeding: MLX utilizes its own random number generator (
mx.random). Anytraining.seedortraining.data_seedvalues in your configuration are silently ignored, and Soup prints a warning to stdout. - Quantization: The trainer enforces 4-bit quantization. The validation logic in
mlx_sft.pychecks quantization settings and warns if 8-bit is requested.
Summary
- Install MLX support via
pip install "soup-cli[mlx]" - Set
backend: mlxinsoup.yamlwithtask: sftfor Apple Silicon training - The
MLXSFTTrainerWrapperinsrc/soup_cli/trainer/mlx_sft.pyorchestrates model loading, LoRA application, and training - Use
mlx-community4-bit quantized models for optimal performance - Resume training detects
*_adapters.safetensorsfiles automatically - DPO, GRPO, and custom RNG seeds are currently unsupported
Frequently Asked Questions
Which Apple Silicon chips work with the Soup MLX backend?
Any Mac with Apple Silicon—including M1, M2, and M3 series chips—supports the MLX backend. The detect_device() function in src/soup_cli/utils/gpu.py automatically recognizes these devices and configures Metal acceleration without manual intervention.
Can I use DPO or GRPO training with the MLX backend?
No. The MLX backend currently restricts training to supervised fine-tuning (SFT) only. If you specify task: dpo or task: grpo with backend: mlx, the configuration validator in src/soup_cli/config/schema.py rejects the job, and src/soup_cli/trainer/mlx_dpo.py raises a runtime guard error.
How do I resume training from an MLX checkpoint?
Use the --resume auto flag when calling soup train. The trainer scans the output directory for files matching the pattern NNNNNNN_adapters.safetensors (where N represents digits) and loads only the LoRA adapter weights. Optimizer states are not preserved in MLX checkpoints, so training resumes with fresh optimizer initialization.
Does the MLX backend support full fine-tuning or only LoRA?
The current implementation in src/soup_cli/trainer/mlx_sft.py supports LoRA fine-tuning exclusively. The linear_to_lora_layers utility from mlx_lm injects trainable low-rank matrices into frozen base models, and there is no code path for full parameter updates in the MLX trainer wrapper.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →