# How llmfit Detects Unified Memory Systems Like Apple Silicon

> Discover how llmfit detects unified memory systems like Apple Silicon. Learn about RAM mapping, unified memory flags, and Metal's working set size for optimal GPU utilization.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-08-22

---

**llmfit treats Apple Silicon as a unified memory system by mapping total system RAM to GPU VRAM, setting the `unified_memory` flag to `true`, and querying Metal's "recommendedMaxWorkingSetSize" to determine the GPU's available working set.**

The `llmfit` repository implements specialized hardware detection for unified memory systems like Apple Silicon in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs). Unlike discrete GPUs with dedicated VRAM, these systems require distinct logic to handle shared physical memory pools between CPU and GPU.

## Detection Flow in SystemSpecs::detect()

The hardware detection process begins in `SystemSpecs::detect()` at line 272 of [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs). When the constructor runs, it first attempts to detect Apple Silicon hardware by calling `Self::detect_apple_gpu(total_ram_gb)`.

If the machine is running on Apple Silicon, this function returns the total system RAM in GiB; otherwise, it returns `None`. This check determines whether the system uses a unified memory architecture.

## The detect_apple_gpu Function Implementation

The `detect_apple_gpu` function, defined around line 1600, identifies Apple Silicon devices and returns the total system RAM as the available GPU memory. This approach recognizes that on unified memory systems, the GPU shares the same physical memory pool as the CPU.

When a value is returned, the detection logic constructs a `GpuInfo` struct and pushes it to the GPU list between lines 272-285:

```rust
gpus.push(GpuInfo {
    name,
    vram_gb: Some(vram),          // = total RAM
    backend: GpuBackend::Metal,   // Metal is the only backend on macOS
    count: 1,
    unified_memory: true,
});

```

The `unified_memory` field is set to **true** to signal that the GPU and CPU share physical memory, while the `backend` is forced to **Metal** because Apple Silicon devices expose their GPU exclusively through the Metal API.

## Querying Available GPU Memory with Metal

After constructing the GPU list, `SystemSpecs::detect()` checks the `unified_memory` flag. When true, it calls `detect_gpu_available_gb()` at lines 334-338 to query the actual GPU-available memory using the Metal API.

This function reads **"recommendedMaxWorkingSetSize"**, which represents the amount of the shared pool the GPU is allowed to wire up. This value is typically lower than total system RAM and represents the practical limit for GPU memory usage. The result is stored in `gpu_available_gb` within the final `SystemSpecs` struct.

## Resulting SystemSpecs Structure

The final `SystemSpecs` struct contains the following fields for unified memory systems:

- `total_ram_gb` – Physical RAM size
- `gpu_vram_gb` – Identical to `total_ram_gb` for Apple Silicon  
- `unified_memory: true` – Signals shared memory architecture
- `backend: GpuBackend::Metal` – Inference backend for model execution
- `gpu_available_gb` – Metal's recommended working set size

This information enables the fit-scoring logic to treat the GPU as having VRAM-first memory while applying correct capacity limits when estimating model fit.

## Practical Implementation Example

To detect unified memory systems in your own code using llmfit:

```rust
use llmfit_core::hardware::SystemSpecs;

fn main() {
    // Detect the current machine's hardware.
    let specs = SystemSpecs::detect();

    // Unified-memory detection (Apple Silicon, NVIDIA Grace, AMD APU, etc.)
    if specs.unified_memory {
        println!("Unified-memory system detected!");
        println!("Total RAM (and GPU VRAM): {:.2} GiB", specs.total_ram_gb);
        println!(
            "GPU-available (Metal's working-set limit): {:?} GiB",
            specs.gpu_available_gb
        );
        println!("Backend used for inference: {:?}", specs.backend);
    } else {
        println!("Discrete-GPU system – normal VRAM handling.");
    }
}

```

Running this on an Apple Silicon Mac outputs:

```

Unified-memory system detected!
Total RAM (and GPU VRAM): 32.00 GiB
GPU-available (Metal's working-set limit): Some(24.00) GiB
Backend used for inference: Metal

```

## Summary

- **llmfit** detects Apple Silicon by calling `detect_apple_gpu()` in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs), which returns total system RAM when running on macOS ARM architecture.
- The system sets `unified_memory: true` in the `GpuInfo` struct to indicate shared physical memory between CPU and GPU.
- The **Metal** backend is forced for all Apple Silicon detection, as Metal is the exclusive GPU API on macOS.
- Available GPU memory is determined by querying Metal's **"recommendedMaxWorkingSetSize"** rather than total RAM, providing accurate working set limits.
- The resulting `SystemSpecs` enables the fit-scoring engine to correctly evaluate model compatibility on unified memory architectures.

## Frequently Asked Questions

### How does llmfit distinguish between Apple Silicon and discrete GPUs?

llmfit distinguishes Apple Silicon by calling the `detect_apple_gpu` function at line 1600 in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs). This function checks for macOS ARM architecture and returns the total system RAM as VRAM if detected. For discrete GPUs, different detection paths populate separate VRAM values without setting the `unified_memory` flag.

### Why does llmfit use Metal's "recommendedMaxWorkingSetSize" instead of total RAM?

The **"recommendedMaxWorkingSetSize"** query provides the amount of unified memory the GPU is actually allowed to wire up, which is typically lower than total system RAM. This value represents the practical GPU memory limit available for model inference, preventing overallocation that could starve the system or other processes.

### Can llmfit detect other unified memory systems besides Apple Silicon?

While the current implementation in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) specifically targets Apple Silicon through the `detect_apple_gpu` function, the architecture supports other unified memory systems. The `unified_memory` boolean flag in `GpuInfo` is designed to handle any unified memory architecture, though additional detection logic would be required for systems like NVIDIA Grace or AMD APUs.

### What happens if recommendedMaxWorkingSetSize is not available?

If the Metal API cannot provide the **"recommendedMaxWorkingSetSize"** value, the `detect_gpu_available_gb()` function returns `None` for `gpu_available_gb`. In this case, the fit-scoring logic falls back to using the total system RAM value, though this represents the maximum theoretical memory rather than the GPU's practical working set limit.