Apple Silicon Specific Runtime Options (MLX) in llmfit: Complete Configuration Guide
Apple Silicon machines can force or auto-detect the MLX runtime using the --force-runtime mlx CLI flag, MLX_LM_HOST environment variable, and built-in platform detection that hides MLX-only models on non-Apple hardware.
The llmfit project provides first-class support for MLX, Apple's Metal-based inference engine exclusive to M1-M4 series Macs. This guide examines the Apple Silicon specific runtime options implemented across the codebase, from automatic hardware detection to manual runtime overrides.
How llmfit Detects Apple Silicon and MLX Availability
The detection pipeline runs at startup and determines whether MLX can be used on the current machine.
Platform Detection in hardware.rs
The hardware.rs module identifies Apple Silicon via system_profiler and flags unified memory architecture:
// From llmfit-core/src/hardware.rs
// Apple Silicon detection enables unified memory treatment
let is_apple_silicon = detect_apple_silicon(); // uses system_profiler
let unified_memory = is_apple_silicon; // Apple Silicon uses unified memory
This detection result feeds into fit calculations, causing the MLX path to treat RAM and VRAM as a single pool—eliminating CPU-offload logic required for discrete GPU machines.
MLX Provider Probe in providers.rs
The providers.rs file implements dual verification for MLX availability:
// From llmfit-core/src/providers.rs
const MLX_DEFAULT_SERVER_URL: &str = "http://localhost:8080";
fn detect_mlx_installed() -> bool {
// Check 1: Is MLX Python package available?
let python_available = check_env_var("MLX_PYTHON_AVAILABLE");
// Check 2: Can we reach the MLX server?
let server_reachable = probe_server(get_mlx_url());
python_available && server_reachable
}
Both conditions must pass for MLX to be marked as an installed runtime.
CLI Flag: Force MLX Runtime Selection
The --force-runtime flag overrides automatic runtime selection, explicitly requesting MLX even when other runtimes are available or preferred.
Usage and Syntax
# Force MLX runtime for recommendation
llmfit recommend --force-runtime mlx
# Force MLX for a specific model file
llmfit fit model.gguf --force-runtime mlx
The flag accepts any runtime variant defined in the InferenceRuntime enum. From main.rs:
// From llmfit-tui/src/main.rs
#[derive(Parser)]
struct Cli {
/// Force a specific inference runtime
#[arg(long = "force-runtime")]
force_runtime: Option<String>, // "mlx", "llamacpp", "ollama", etc.
}
The string is parsed against InferenceRuntime::Mlx and other variants in fit.rs.
Environment Variable: MLX_LM_HOST
Override the default MLX server endpoint without modifying configuration files.
# Point to remote MLX server
export MLX_LM_HOST=http://my-mlserver.local:8080
llmfit recommend
# One-shot override
MLX_LM_HOST=http://10.0.1.50:9000 llmfit fit model.gguf
From providers.rs, the resolution order is:
MLX_LM_HOSTenvironment variable (if valid URL)- Default
http://localhost:8080
Model Filtering: Hiding MLX-Only Models on Incompatible Hardware
The TUI layer prevents confusion by hiding -MLX suffixed models on non-Apple systems.
Detection Logic in models.rs
// From llmfit-core/src/models.rs
impl Model {
/// Returns true for models like "Llama-3-8B-MLX" or "Qwen-7B-MLX-4bit"
pub fn is_mlx_specific(&self) -> bool {
self.name.ends_with("-MLX") || self.name.contains("-MLX-")
}
}
UI Filtering in tui_app.rs
// From llmfit-tui/src/tui_app.rs
fn filter_models_for_platform(models: Vec<Model>) -> Vec<Model> {
let is_apple = hardware::is_apple_silicon();
models.into_iter()
.filter(|m| !m.is_mlx_specific() || is_apple)
.collect()
}
On Intel Macs or Linux/Windows systems, MLX-only models never appear in selection lists.
Runtime Selection Flow: How Options Interact
The precedence order when multiple options are present:
| Priority | Mechanism | Effect |
|---|---|---|
| 1 | --force-runtime mlx |
Unconditional MLX selection, skips detection |
| 2 | MLX_LM_HOST + successful probe |
MLX available, may be auto-selected |
| 3 | Auto-detection | MLX used if Apple Silicon + server reachable + Python package present |
| 4 | Fallback | CPU or other GPU runtime |
Complete Configuration Examples
Development Workflow on M3 MacBook Pro
# Verify MLX is detected automatically
llmfit recommend
# Output: "Using runtime: MLX (Apple Silicon)"
# Force MLX for benchmarking against llama.cpp
llmfit fit llama-3-8b.gguf --force-runtime mlx --benchmark
llmfit fit llama-3-8b.gguf --force-runtime llamacpp --benchmark
# Remote MLX server for distributed testing
MLX_LM_HOST=http://mac-studio.local:8080 llmfit recommend
Programmatic Runtime Selection
use llmfit_core::fit::{InferenceRuntime, FitEngine};
use llmfit_core::providers::Provider;
// Runtime-agnostic: auto-detect best option
let engine = FitEngine::auto().unwrap();
// Force MLX explicitly
let engine = FitEngine::with_runtime(InferenceRuntime::Mlx);
// Check availability before forcing
let installed = Provider::detect_installed();
if !installed.mlx.is_empty() {
println!("MLX ready: {} models", installed.mlx.len());
}
Summary
- Automatic detection in
providers.rsandhardware.rsverifies Apple Silicon hardware, MLX Python package, and server reachability before enabling MLX --force-runtime mlxinmain.rsbypasses auto-selection for explicit MLX controlMLX_LM_HOSTenvironment variable redirects MLX server connections-MLXmodel suffix inmodels.rstriggers platform-specific filtering intui_app.rs- Unified memory handling in
fit.rsoptimizes memory calculations exclusive to Apple Silicon
Frequently Asked Questions
What happens if I run --force-runtime mlx on Intel Mac or Linux?
The command attempts MLX initialization, fails the platform check, and llmfit returns an error indicating MLX is unavailable on non-Apple Silicon hardware. The error originates from the runtime constructor in fit.rs before any model loading occurs.
Can I use MLX on a remote Apple Silicon machine from my Intel Mac?
Yes. Set MLX_LM_HOST to the remote machine's address. The local llmfit instance treats it as a standard MLX provider endpoint. Model filtering still hides -MLX models locally, but you can force them via --force-runtime mlx since the runtime check validates server reachability rather than local hardware.
How does llmfit distinguish between MLX Python package and MLX server?
providers.rs checks MLX_PYTHON_AVAILABLE environment variable for local package installation, while separately probing MLX_LM_HOST or default localhost for server responsiveness. Both can be true (local MLX with server), or only server (remote MLX), or only package (Python API without server).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →