How Apple Silicon Unified Memory Affects llmfit Run Modes: Detection & Execution Logic

Apple Silicon's unified memory architecture causes llmfit to bypass the CPU-offload path and choose between GPU-only or CPU-only execution modes.

The llmfit crate by AlexsJones dynamically adapts its model execution strategy based on hardware detection. On Apple Silicon Macs, where CPU and GPU share a single memory pool, the traditional CPU-offload optimization becomes redundant. This article examines how llmfit-core detects unified memory and alters its run mode selection accordingly.

Detecting Apple Silicon Unified Memory in llmfit

The hardware detection logic lives in llmfit-core/src/hardware.rs. During SystemSpecs::detect(), the code executes macOS-specific system profiling to identify unified memory capabilities.

Detection Mechanism

The detection occurs at lines 78-91, where llmfit shells out to system_profiler and parses the output:

// Detect Apple Silicon unified memory.
let unified_memory = if cfg!(target_os = "macos") {
    // Use system_profiler to detect if unified memory is present.
    if let Ok(output) = std::process::Command::new("system_profiler")
        .arg("SPDisplaysDataType")
        .output()
    {
        let stdout = String::from_utf8_lossy(&output.stdout);
        stdout.contains("Unified Memory")
    } else {
        false
    }
} else {
    false
};

When unified memory is detected, llmfit remaps the VRAM metric to system RAM (lines 93-95):

// On unified memory Macs, VRAM is effectively the same as system RAM.
if unified_memory {
    total_vram_gb = ram_gb;
}

This remapping ensures downstream calculations treat the entire memory pool as available GPU-accessible memory.

How Unified Memory Alters Run Mode Selection

The run mode selection logic resides in llmfit-core/src/fit.rs, specifically within ModelFit::choose_best_mode() at lines 86-114.

Standard Non-Unified Memory Path

On traditional systems with discrete GPUs, llmfit evaluates this fallback chain:

  1. Gpu mode — if model.min_vram_gb <= sys.total_vram_gb
  2. CpuOffload mode — if model.min_vram_gb + model.min_ram_gb * 0.5 <= sys.total_vram_gb
  3. CpuOnly mode — if model.min_ram_gb <= sys.ram_gb

Unified Memory Bypass

The Apple Silicon path short-circuits this logic at lines 91-96:

// If unified memory is present (Apple Silicon), we treat VRAM as RAM.
if sys.unified_memory {
    // On unified memory Macs, CPU-offload path is skipped because VRAM == RAM.
    if model.min_ram_gb <= sys.ram_gb {
        return (RunMode::CpuOnly, model.min_ram_gb);
    }
}

Key implications:

  • RunMode::CpuOffload is never selected — the hybrid split between VRAM and RAM serves no purpose when both subsystems share identical physical memory
  • Memory calculation simplifies — only min_ram_gb is evaluated against ram_gb
  • Execution falls through to CPU-only or continues to other fallback candidates

Complete Run Mode Selection Flow

The RunMode enum defines five execution strategies in llmfit-core/src/fit.rs:

RunMode Description Unified Memory Behavior
Gpu Weights fully resident in GPU memory Available if model fits in unified pool
MoeOffload Mixture-of-Experts partial loading Available via forced mode or fallback
CpuOffload Model split between GPU and system RAM Skipped entirely
CpuOnly Entire model in system RAM Primary fallback on unified memory
TensorParallel Multi-node tensor parallelism Available for distributed setups

Practical Code Example

The following demonstrates detection and mode selection on Apple Silicon:

use llmfit_core::hardware::SystemSpecs;
use llmfit_core::fit::{ModelFit, RunMode};
use llmfit_core::models::ModelInfo;

fn main() {
    // Detect hardware — unified_memory true on Apple Silicon Macs
    let specs = SystemSpecs::detect();
    println!("Apple Silicon unified memory: {}", specs.unified_memory);
    println!("Available memory pool: {:.1} GB", specs.ram_gb);

    // Model requiring 16 GB VRAM, 8 GB RAM
    let model = ModelInfo::dummy_with_vram(16.0, 8.0);

    // llmfit selects optimal execution mode
    let fit = ModelFit::analyze_with_forced_runtime(model, &specs, None);
    
    println!("Selected run mode: {}", fit.run_mode.name());
    println!("Estimated memory: {:.1} GB", fit.memory_gb);
}

Sample output on MacBook Pro (M3, 36 GB unified memory):

Apple Silicon unified memory: true
Available memory pool: 36.0 GB
Selected run mode: CPU-Only
Estimated memory: 8.0 GB

Notice that despite the model's 16 GB "VRAM" requirement, the unified memory detection routes execution through CpuOnly rather than attempting CPU-offload split calculations.

Performance Characteristics by Mode

The throughput estimation function at lines 116-125 assigns efficiency multipliers:

fn estimate_throughput(model: &ModelInfo, mode: RunMode, sys: &SystemSpecs) -> f64 {
    match mode {
        RunMode::Gpu => model.recommended_throughput_gb() * 1.0,
        RunMode::MoeOffload => model.recommended_throughput_gb() * 0.9,
        RunMode::CpuOffload => model.recommended_throughput_gb() * 0.6,
        RunMode::CpuOnly => model.recommended_throughput_gb() * 0.4,
        RunMode::TensorParallel => model.recommended_throughput_gb() * 0.8,
    }
}

On Apple Silicon, the CpuOnly mode (0.4x baseline) may underperform relative to Metal GPU acceleration. However, llmfit allows forced mode selection via analyze_with_forced_runtime() for users who explicitly want GPU execution through frameworks like MLX or PyTorch with Metal backend.

Summary

  • Detection source: llmfit-core/src/hardware.rs uses system_profiler SPDisplaysDataType to identify "Unified Memory" string on macOS
  • Memory remapping: total_vram_gb becomes ram_gb when unified memory is detected
  • Execution impact: RunMode::CpuOffload is bypassed; selection proceeds directly to CpuOnly or other fallbacks
  • User control: forced parameter in analyze_with_forced_runtime() overrides automatic selection

Frequently Asked Questions

How does llmfit detect Apple Silicon specifically?

llmfit does not explicitly check for the M1/M2/M3 chip identifier. Instead, it detects the unified memory capability through system_profiler output parsing. This approach correctly identifies both Apple Silicon Macs and any future macOS systems with unified memory architectures.

Can I force GPU mode on Apple Silicon despite llmfit selecting CPU-only?

Yes. Pass Some(RunMode::Gpu) to analyze_with_forced_runtime(). The viability check at line 63 (needed <= sys.total_vram_gb) will pass because unified memory detection remaps total_vram_gb to your full system RAM. This enables Metal-backed GPU execution when the underlying inference framework supports it.

Why does llmfit skip CPU-offload on unified memory systems?

CPU-offload optimizes for the bandwidth and latency differential between discrete GPU VRAM and slower system RAM. On Apple Silicon, CPU and GPU access the same physical memory pool at identical speeds, eliminating any benefit from partial GPU residence. The CpuOnly mode provides equivalent memory access patterns without the complexity of layer splitting.

What happens if a model exceeds available unified memory?

The selection algorithm falls through to the minimal-memory candidate sorting at lines 108-113. llmfit will still recommend a run mode, but actual execution will likely trigger system-level memory swapping or framework-specific out-of-memory errors. The tool provides memory estimates; enforcement depends on the downstream inference engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →