Apple Silicon Specific Runtime Options (MLX) in llmfit: Complete Configuration Guide

Apple Silicon machines can force or auto-detect the MLX runtime using the --force-runtime mlx CLI flag, MLX_LM_HOST environment variable, and built-in platform detection that hides MLX-only models on non-Apple hardware.

The llmfit project provides first-class support for MLX, Apple's Metal-based inference engine exclusive to M1-M4 series Macs. This guide examines the Apple Silicon specific runtime options implemented across the codebase, from automatic hardware detection to manual runtime overrides.


How llmfit Detects Apple Silicon and MLX Availability

The detection pipeline runs at startup and determines whether MLX can be used on the current machine.

Platform Detection in hardware.rs

The hardware.rs module identifies Apple Silicon via system_profiler and flags unified memory architecture:

// From llmfit-core/src/hardware.rs
// Apple Silicon detection enables unified memory treatment
let is_apple_silicon = detect_apple_silicon(); // uses system_profiler
let unified_memory = is_apple_silicon; // Apple Silicon uses unified memory

This detection result feeds into fit calculations, causing the MLX path to treat RAM and VRAM as a single pool—eliminating CPU-offload logic required for discrete GPU machines.

MLX Provider Probe in providers.rs

The providers.rs file implements dual verification for MLX availability:

// From llmfit-core/src/providers.rs
const MLX_DEFAULT_SERVER_URL: &str = "http://localhost:8080";

fn detect_mlx_installed() -> bool {
    // Check 1: Is MLX Python package available?
    let python_available = check_env_var("MLX_PYTHON_AVAILABLE");
    
    // Check 2: Can we reach the MLX server?
    let server_reachable = probe_server(get_mlx_url());
    
    python_available && server_reachable
}

Both conditions must pass for MLX to be marked as an installed runtime.


CLI Flag: Force MLX Runtime Selection

The --force-runtime flag overrides automatic runtime selection, explicitly requesting MLX even when other runtimes are available or preferred.

Usage and Syntax


# Force MLX runtime for recommendation

llmfit recommend --force-runtime mlx

# Force MLX for a specific model file

llmfit fit model.gguf --force-runtime mlx

The flag accepts any runtime variant defined in the InferenceRuntime enum. From main.rs:

// From llmfit-tui/src/main.rs
#[derive(Parser)]
struct Cli {
    /// Force a specific inference runtime
    #[arg(long = "force-runtime")]
    force_runtime: Option<String>, // "mlx", "llamacpp", "ollama", etc.
}

The string is parsed against InferenceRuntime::Mlx and other variants in fit.rs.


Environment Variable: MLX_LM_HOST

Override the default MLX server endpoint without modifying configuration files.


# Point to remote MLX server

export MLX_LM_HOST=http://my-mlserver.local:8080
llmfit recommend

# One-shot override

MLX_LM_HOST=http://10.0.1.50:9000 llmfit fit model.gguf

From providers.rs, the resolution order is:

  1. MLX_LM_HOST environment variable (if valid URL)
  2. Default http://localhost:8080

Model Filtering: Hiding MLX-Only Models on Incompatible Hardware

The TUI layer prevents confusion by hiding -MLX suffixed models on non-Apple systems.

Detection Logic in models.rs

// From llmfit-core/src/models.rs
impl Model {
    /// Returns true for models like "Llama-3-8B-MLX" or "Qwen-7B-MLX-4bit"
    pub fn is_mlx_specific(&self) -> bool {
        self.name.ends_with("-MLX") || self.name.contains("-MLX-")
    }
}

UI Filtering in tui_app.rs

// From llmfit-tui/src/tui_app.rs
fn filter_models_for_platform(models: Vec<Model>) -> Vec<Model> {
    let is_apple = hardware::is_apple_silicon();
    
    models.into_iter()
        .filter(|m| !m.is_mlx_specific() || is_apple)
        .collect()
}

On Intel Macs or Linux/Windows systems, MLX-only models never appear in selection lists.


Runtime Selection Flow: How Options Interact

The precedence order when multiple options are present:

Priority Mechanism Effect
1 --force-runtime mlx Unconditional MLX selection, skips detection
2 MLX_LM_HOST + successful probe MLX available, may be auto-selected
3 Auto-detection MLX used if Apple Silicon + server reachable + Python package present
4 Fallback CPU or other GPU runtime

Complete Configuration Examples

Development Workflow on M3 MacBook Pro


# Verify MLX is detected automatically

llmfit recommend

# Output: "Using runtime: MLX (Apple Silicon)"

# Force MLX for benchmarking against llama.cpp

llmfit fit llama-3-8b.gguf --force-runtime mlx --benchmark
llmfit fit llama-3-8b.gguf --force-runtime llamacpp --benchmark

# Remote MLX server for distributed testing

MLX_LM_HOST=http://mac-studio.local:8080 llmfit recommend

Programmatic Runtime Selection

use llmfit_core::fit::{InferenceRuntime, FitEngine};
use llmfit_core::providers::Provider;

// Runtime-agnostic: auto-detect best option
let engine = FitEngine::auto().unwrap();

// Force MLX explicitly
let engine = FitEngine::with_runtime(InferenceRuntime::Mlx);

// Check availability before forcing
let installed = Provider::detect_installed();
if !installed.mlx.is_empty() {
    println!("MLX ready: {} models", installed.mlx.len());
}

Summary

  • Automatic detection in providers.rs and hardware.rs verifies Apple Silicon hardware, MLX Python package, and server reachability before enabling MLX
  • --force-runtime mlx in main.rs bypasses auto-selection for explicit MLX control
  • MLX_LM_HOST environment variable redirects MLX server connections
  • -MLX model suffix in models.rs triggers platform-specific filtering in tui_app.rs
  • Unified memory handling in fit.rs optimizes memory calculations exclusive to Apple Silicon

Frequently Asked Questions

What happens if I run --force-runtime mlx on Intel Mac or Linux?

The command attempts MLX initialization, fails the platform check, and llmfit returns an error indicating MLX is unavailable on non-Apple Silicon hardware. The error originates from the runtime constructor in fit.rs before any model loading occurs.

Can I use MLX on a remote Apple Silicon machine from my Intel Mac?

Yes. Set MLX_LM_HOST to the remote machine's address. The local llmfit instance treats it as a standard MLX provider endpoint. Model filtering still hides -MLX models locally, but you can force them via --force-runtime mlx since the runtime check validates server reachability rather than local hardware.

How does llmfit distinguish between MLX Python package and MLX server?

providers.rs checks MLX_PYTHON_AVAILABLE environment variable for local package installation, while separately probing MLX_LM_HOST or default localhost for server responsiveness. Both can be true (local MLX with server), or only server (remote MLX), or only package (Python API without server).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →