# How Apple Silicon Unified Memory Affects llmfit Run Modes: Detection & Execution Logic

> Discover how Apple Silicon unified memory impacts llmfit run modes, forcing GPU-only or CPU-only execution and bypassing CPU-offload.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-20

---

**Apple Silicon's unified memory architecture causes llmfit to bypass the CPU-offload path and choose between GPU-only or CPU-only execution modes.**

The `llmfit` crate by AlexsJones dynamically adapts its model execution strategy based on hardware detection. On Apple Silicon Macs, where CPU and GPU share a single memory pool, the traditional CPU-offload optimization becomes redundant. This article examines how `llmfit-core` detects unified memory and alters its run mode selection accordingly.

## Detecting Apple Silicon Unified Memory in llmfit

The hardware detection logic lives in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs). During `SystemSpecs::detect()`, the code executes macOS-specific system profiling to identify unified memory capabilities.

### Detection Mechanism

The detection occurs at lines 78-91, where `llmfit` shells out to `system_profiler` and parses the output:

```rust
// Detect Apple Silicon unified memory.
let unified_memory = if cfg!(target_os = "macos") {
    // Use system_profiler to detect if unified memory is present.
    if let Ok(output) = std::process::Command::new("system_profiler")
        .arg("SPDisplaysDataType")
        .output()
    {
        let stdout = String::from_utf8_lossy(&output.stdout);
        stdout.contains("Unified Memory")
    } else {
        false
    }
} else {
    false
};

```

When unified memory is detected, `llmfit` remaps the VRAM metric to system RAM (lines 93-95):

```rust
// On unified memory Macs, VRAM is effectively the same as system RAM.
if unified_memory {
    total_vram_gb = ram_gb;
}

```

This remapping ensures downstream calculations treat the entire memory pool as available GPU-accessible memory.

## How Unified Memory Alters Run Mode Selection

The run mode selection logic resides in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), specifically within `ModelFit::choose_best_mode()` at lines 86-114.

### Standard Non-Unified Memory Path

On traditional systems with discrete GPUs, `llmfit` evaluates this fallback chain:

1. **Gpu mode** — if `model.min_vram_gb <= sys.total_vram_gb`
2. **CpuOffload mode** — if `model.min_vram_gb + model.min_ram_gb * 0.5 <= sys.total_vram_gb`
3. **CpuOnly mode** — if `model.min_ram_gb <= sys.ram_gb`

### Unified Memory Bypass

The Apple Silicon path short-circuits this logic at lines 91-96:

```rust
// If unified memory is present (Apple Silicon), we treat VRAM as RAM.
if sys.unified_memory {
    // On unified memory Macs, CPU-offload path is skipped because VRAM == RAM.
    if model.min_ram_gb <= sys.ram_gb {
        return (RunMode::CpuOnly, model.min_ram_gb);
    }
}

```

**Key implications:**
- **`RunMode::CpuOffload` is never selected** — the hybrid split between VRAM and RAM serves no purpose when both subsystems share identical physical memory
- **Memory calculation simplifies** — only `min_ram_gb` is evaluated against `ram_gb`
- **Execution falls through to CPU-only** or continues to other fallback candidates

## Complete Run Mode Selection Flow

The `RunMode` enum defines five execution strategies in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs):

| RunMode | Description | Unified Memory Behavior |
|---------|-------------|------------------------|
| `Gpu` | Weights fully resident in GPU memory | Available if model fits in unified pool |
| `MoeOffload` | Mixture-of-Experts partial loading | Available via forced mode or fallback |
| `CpuOffload` | Model split between GPU and system RAM | **Skipped entirely** |
| `CpuOnly` | Entire model in system RAM | Primary fallback on unified memory |
| `TensorParallel` | Multi-node tensor parallelism | Available for distributed setups |

## Practical Code Example

The following demonstrates detection and mode selection on Apple Silicon:

```rust
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::fit::{ModelFit, RunMode};
use llmfit_core::models::ModelInfo;

fn main() {
    // Detect hardware — unified_memory true on Apple Silicon Macs
    let specs = SystemSpecs::detect();
    println!("Apple Silicon unified memory: {}", specs.unified_memory);
    println!("Available memory pool: {:.1} GB", specs.ram_gb);

    // Model requiring 16 GB VRAM, 8 GB RAM
    let model = ModelInfo::dummy_with_vram(16.0, 8.0);

    // llmfit selects optimal execution mode
    let fit = ModelFit::analyze_with_forced_runtime(model, &specs, None);
    
    println!("Selected run mode: {}", fit.run_mode.name());
    println!("Estimated memory: {:.1} GB", fit.memory_gb);
}

```

**Sample output on MacBook Pro (M3, 36 GB unified memory):**

```text
Apple Silicon unified memory: true
Available memory pool: 36.0 GB
Selected run mode: CPU-Only
Estimated memory: 8.0 GB

```

Notice that despite the model's 16 GB "VRAM" requirement, the unified memory detection routes execution through `CpuOnly` rather than attempting CPU-offload split calculations.

## Performance Characteristics by Mode

The throughput estimation function at lines 116-125 assigns efficiency multipliers:

```rust
fn estimate_throughput(model: &ModelInfo, mode: RunMode, sys: &SystemSpecs) -> f64 {
    match mode {
        RunMode::Gpu => model.recommended_throughput_gb() * 1.0,
        RunMode::MoeOffload => model.recommended_throughput_gb() * 0.9,
        RunMode::CpuOffload => model.recommended_throughput_gb() * 0.6,
        RunMode::CpuOnly => model.recommended_throughput_gb() * 0.4,
        RunMode::TensorParallel => model.recommended_throughput_gb() * 0.8,
    }
}

```

On Apple Silicon, the `CpuOnly` mode (0.4x baseline) may underperform relative to Metal GPU acceleration. However, `llmfit` allows forced mode selection via `analyze_with_forced_runtime()` for users who explicitly want GPU execution through frameworks like MLX or PyTorch with Metal backend.

## Summary

- **Detection source**: [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) uses `system_profiler SPDisplaysDataType` to identify "Unified Memory" string on macOS
- **Memory remapping**: `total_vram_gb` becomes `ram_gb` when unified memory is detected
- **Execution impact**: `RunMode::CpuOffload` is bypassed; selection proceeds directly to `CpuOnly` or other fallbacks
- **User control**: `forced` parameter in `analyze_with_forced_runtime()` overrides automatic selection

## Frequently Asked Questions

### How does llmfit detect Apple Silicon specifically?

`llmfit` does not explicitly check for the M1/M2/M3 chip identifier. Instead, it detects the **unified memory** capability through `system_profiler` output parsing. This approach correctly identifies both Apple Silicon Macs and any future macOS systems with unified memory architectures.

### Can I force GPU mode on Apple Silicon despite llmfit selecting CPU-only?

Yes. Pass `Some(RunMode::Gpu)` to `analyze_with_forced_runtime()`. The viability check at line 63 (`needed <= sys.total_vram_gb`) will pass because unified memory detection remaps `total_vram_gb` to your full system RAM. This enables Metal-backed GPU execution when the underlying inference framework supports it.

### Why does llmfit skip CPU-offload on unified memory systems?

CPU-offload optimizes for the bandwidth and latency differential between discrete GPU VRAM and slower system RAM. On Apple Silicon, CPU and GPU access the same physical memory pool at identical speeds, eliminating any benefit from partial GPU residence. The `CpuOnly` mode provides equivalent memory access patterns without the complexity of layer splitting.

### What happens if a model exceeds available unified memory?

The selection algorithm falls through to the minimal-memory candidate sorting at lines 108-113. `llmfit` will still recommend a run mode, but actual execution will likely trigger system-level memory swapping or framework-specific out-of-memory errors. The tool provides memory estimates; enforcement depends on the downstream inference engine.