# Why Apple Silicon Unified Memory Skips CpuOffload in llmfit

> Discover why Apple Silicon unified memory skips CpuOffload in llmfit. Learn how direct GPU access to system RAM boosts LLM performance and eliminates data copying.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-11

---

**On Apple Silicon systems, llmfit skips the CpuOffload mode because unified memory architectures allow the GPU to access system RAM directly, eliminating the need to copy data between distinct VRAM and RAM pools.**

The llmfit repository contains a Rust-based inference engine that optimizes model execution based on hardware capabilities. When running on Apple Silicon devices, the framework detects the unified memory architecture and deliberately bypasses the CpuOffload path used by discrete GPU systems. This optimization prevents unnecessary memory copies and ensures accurate resource planning on Metal-enabled Macs.

## The Architectural Difference Between Discrete and Unified Memory

**CpuOffload** is a run mode designed for discrete GPU systems where VRAM is physically separate from system RAM. On these systems, when GPU memory is exhausted, the framework must copy tensors to system RAM and perform computations on the CPU. 

Apple Silicon devices use a **unified memory** architecture where the GPU and CPU share the same physical memory pool. Because the Metal backend can already address the full system RAM directly, the concept of "offloading" from GPU memory to CPU memory becomes redundant. The data remains in the same memory location regardless of which processor accesses it.

## Hardware Detection in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs)

The detection logic resides in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs), where the `SystemSpecs` struct captures the hardware configuration. The code queries the primary GPU's properties to determine if it reports `unified_memory` as true.

```rust
let primary = gpus.first();
let unified_memory = primary.map(|g| g.unified_memory).unwrap_or(false);
let specs = SystemSpecs {
    // ... other fields ...
    unified_memory,
    backend: primary.map(|g| g.backend).unwrap_or(cpu_backend),
    // ... other fields ...
};

```

When running on Apple Silicon, the Metal backend reports `unified_memory: true`, setting the `unified_memory` boolean field in `SystemSpecs` to `true`. This flag propagates through the core analysis pipeline and influences all subsequent planning decisions.

## Run Mode Selection in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs)

The core analysis in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) uses the `unified_memory` flag to short-circuit the run mode selection. When this flag is detected, the function immediately returns `RunMode::Gpu` and skips evaluation of the CpuOffload path.

```rust
fn select_run_mode(specs: &SystemSpecs) -> RunMode {
    if specs.unified_memory {
        // Apple Silicon – GPU can already touch system RAM.
        RunMode::Gpu
    } else if specs.has_gpu {
        // Discrete-GPU path – may need CpuOffload evaluation.
        // ... detailed logic for VRAM-constrained scenarios ...
    } else {
        RunMode::CpuOnly
    }
}

```

By returning `RunMode::Gpu` for unified memory systems, the code ensures that the model remains on the GPU execution path. This avoids the performance penalty and complexity of managing separate memory pools that the CpuOffload mode is designed to handle.

## Execution Planning in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs)

The planning phase in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) contains the final branch that omits the offload logic entirely. When `system.unified_memory` is true, the planner generates a GPU-only execution plan without creating fallback layers for CPU execution.

```rust
match system.unified_memory {
    true => {
        // Direct GPU path – no offload needed.
        let plan = Plan::new_gpu(...);
    }
    false => {
        // Evaluate Gpu, CpuOffload, CpuOnly, and hybrid strategies.
        // ... logic for discrete GPU memory management ...
    }
}

```

This branch ensures that Apple Silicon devices do not create execution plans that attempt to move data between memory pools. Since the GPU and CPU share the same memory, attempting to "offload" would duplicate work and potentially produce incorrect memory usage estimates by double-counting the same physical RAM.

## Summary

- **CpuOffload** is designed exclusively for discrete GPUs with limited VRAM that must spill to system RAM.
- Apple Silicon reports `unified_memory: true` through the Metal backend, captured in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs).
- The [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) module detects this flag and returns `RunMode::Gpu`, bypassing CpuOffload evaluation.
- The [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) module generates GPU-only execution plans for unified memory systems, avoiding redundant memory copy operations.
- Skipping CpuOffload on Apple Silicon prevents performance overhead and ensures accurate memory accounting.

## Frequently Asked Questions

### What is CpuOffload in llmfit?

**CpuOffload** is a run mode in llmfit designed for discrete GPU systems where VRAM is physically separate from system RAM. When GPU memory is exhausted, this mode copies model layers to system RAM and executes them on the CPU, allowing models larger than VRAM to run at reduced speed.

### Does skipping CpuOffload affect performance on Apple Silicon?

No, skipping CpuOffload does not harm performance. Because Apple Silicon uses unified memory, the GPU can access the full system RAM directly without copying data. The model stays on the GPU execution path, avoiding the latency that discrete GPUs incur when moving data between VRAM and RAM.

### How does llmfit detect unified memory?

The detection occurs in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) by querying the primary GPU's `unified_memory` property. When running on macOS with Metal, Apple Silicon devices report `true` for this property, which sets the `unified_memory` flag in the `SystemSpecs` struct used throughout the framework.

### Can I force CpuOffload on Apple Silicon?

The source code in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) and [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) explicitly bypasses CpuOffload logic when `unified_memory` is true. There is no configuration option to enable CpuOffload on Apple Silicon because the architecture makes it unnecessary—the GPU already accesses the same memory pool that would serve as the "offload" destination.