# Unified Memory Detection Mechanism in LLMFIT: Apple Silicon vs NVIDIA Grace Blackwell

> Discover LLMFIT's unified memory detection for Apple Silicon vs NVIDIA Grace Blackwell. Our mechanism uses system APIs to identify shared RAM and optimize VRAM calculations.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-21

---

**LLMFIT detects unified memory architectures by querying system-specific APIs—`system_profiler` for Apple Silicon and `nvidia-smi` addressing modes for NVIDIA Grace—to flag shared RAM pools and adjust VRAM calculations accordingly.**

The [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) module implements a cross-platform unified memory detection mechanism that determines whether a GPU shares physical RAM with the CPU. This distinction is critical for the AlexsJones/llmfit repository because model fit scoring treats the entire system memory as VRAM when unified memory is detected, avoiding incorrect capacity assumptions on Apple Silicon and NVIDIA Grace Blackwell SoCs.

## How Unified Memory Detection Works in LLMFIT

The detection flow begins in `SystemSpecs::detect()`, which orchestrates GPU enumeration through `detect_all_gpus()`. This function assesses each hardware platform sequentially, applying platform-specific heuristics to identify unified memory architectures before constructing `GpuInfo` structs.

### Apple Silicon Detection via system_profiler

For macOS systems, the code checks for Apple Silicon GPUs after querying other backends. The `detect_apple_gpu(total_ram_gb)` function executes `system_profiler SPDisplaysDataType` to read the `recommendedMaxWorkingSetSize` value.

```rust
// Inside detect_all_gpus() in llmfit-core/src/hardware.rs
if let Some(vram) = Self::detect_apple_gpu(total_ram_gb) {
    let name = if cpu_name.to_lowercase().contains("apple") {
        cpu_name.to_string()
    } else {
        "Apple Silicon".to_string()
    };
    gpus.push(GpuInfo {
        name,
        vram_gb: Some(vram),          // VRAM equals total system RAM
        backend: GpuBackend::Metal,   // Metal is the only backend on macOS
        count: 1,
        unified_memory: true,         // Flag shared memory architecture
    });
}

```

Key implementation details include:

- **Physical memory pool**: The GPU draws from the same RAM as the CPU, so `detect_apple_gpu` assigns `total_ram_gb` as the VRAM value.
- **Backend assignment**: `GpuBackend::Metal` is hardcoded for Apple Silicon.
- **Detection location**: The `system_profiler` parsing logic resides around lines 710–750 in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs).

### NVIDIA Grace Blackwell Detection via nvidia-smi

For NVIDIA Grace SoCs, `detect_nvidia_gpus()` issues an extended `nvidia-smi` query that includes the `addressing_mode` column.

```rust
// Extended query in detect_nvidia_gpus()
let output = std::process::Command::new("nvidia-smi")
    .arg("--query-gpu=addressing_mode,memory.total,name")
    .arg("--format=csv,noheader,nounits")
    .output()
    .ok()?;

```

The `parse_nvidia_smi_extended()` function interprets the `addressing_mode` field. If the mode equals **ATS** (Address Translation Services), the GPU uses unified memory:

```rust
let addr_mode = parts[0].trim();
let is_unified = addr_mode.eq_ignore_ascii_case("ATS");

// Aggregation logic later in the function
let entry = grouped.entry(name).or_insert((0, 0.0, false));
entry.0 += 1;
if vram_mb > entry.1 { entry.1 = vram_mb; }
if is_unified { entry.2 = true; }   // Set unified flag

```

After building the GPU list, LLMFIT performs a final pass to correct VRAM values for unified memory devices:

```rust
let is_nvidia_unified = gpus.iter().any(|g| is_nvidia_unified_memory_gpu(&g.name));
if is_nvidia_unified {
    for gpu in &mut gpus {
        if is_nvidia_unified_memory_gpu(&gpu.name) {
            gpu.unified_memory = true;
            gpu.vram_gb = Some(total_ram_gb);   // Treat RAM as VRAM
        }
    }
}

```

The helper `is_nvidia_unified_memory_gpu` (lines ~340–350) matches known Grace GPU identifiers such as "Grace" or specific device IDs like `10de:2e12`.

## Impact on Model Fit Scoring

The `unified_memory` boolean flag drives critical logic downstream. When `true`, the fit-scoring algorithm skips CPU-offload paths designed for discrete VRAM scenarios. This optimization recognizes that "spilling" to CPU RAM is meaningless when the GPU already accesses the same memory pool.

Additionally, `gpu_available_gb` queries (lines ~130–138) only execute for Apple Silicon devices, as macOS provides specific APIs for available GPU working set size that NVIDIA Grace platforms lack.

## Practical Implementation Example

```rust
use llmfit_core::hardware::SystemSpecs;

let specs = SystemSpecs::detect();

for gpu in specs.gpus {
    println!("GPU: {}", gpu.name);
    println!("  Backend: {}", gpu.backend.label());
    println!("  VRAM (GB): {:.2}", gpu.vram_gb.unwrap_or(0.0));
    println!("  Unified memory? {}", gpu.unified_memory);
}

```

Running this snippet on an Apple Silicon Mac produces output similar to:

```

GPU: Apple M2 Pro
  Backend: Metal
  VRAM (GB): 32.00
  Unified memory? true

```

On a Grace-based DGX Spark node, the output reflects the larger unified memory pool:

```

GPU: NVIDIA Grace
  Backend: CUDA
  VRAM (GB): 256.00
  Unified memory? true

```

## Summary

- **Apple Silicon path**: Uses `system_profiler SPDisplaysDataType` to read `recommendedMaxWorkingSetSize`, assigns total RAM as VRAM, and sets `backend: Metal` with `unified_memory: true`.
- **NVIDIA Grace path**: Parses `nvidia-smi` addressing mode for "ATS", flags unified memory via `is_nvidia_unified_memory_gpu`, and overwrites VRAM with total system RAM.
- **Convergence**: Both platforms populate `GpuInfo` with `unified_memory: true`, enabling transparent handling of shared memory pools in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs).
- **Downstream effects**: The flag prevents incorrect CPU-offload calculations and ensures accurate model sizing on unified memory architectures.

## Frequently Asked Questions

### How does LLMFIT distinguish between discrete and unified memory GPUs?

LLMFIT checks platform-specific indicators. For Apple Silicon, it detects the Metal backend and queries macOS system profiler data. For NVIDIA, it checks the `addressing_mode` field for "ATS" (Address Translation Services) in `nvidia-smi` output. Discrete GPUs report separate VRAM values and lack these unified memory indicators.

### Why does VRAM equal total RAM on unified memory systems?

Because the GPU and CPU share the same physical memory pool. On Apple Silicon and NVIDIA Grace, there is no separate VRAM chip—the GPU allocates from system RAM. LLMFIT reflects this hardware reality by assigning `total_ram_gb` to the `vram_gb` field when `unified_memory` is true.

### Where is the unified memory detection logic located in the source code?

The primary implementation resides in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs). Key functions include `detect_apple_gpu()` around lines 710–750, `parse_nvidia_smi_extended()` for ATS detection, and `is_nvidia_unified_memory_gpu()` around lines 340–350.

### Does unified memory detection affect model loading strategy?

Yes. When `unified_memory` is true, fit scoring skips CPU-offload paths designed for discrete VRAM scenarios. This prevents incorrect capacity calculations during model fitting because there is no separate RAM pool to spill to when the GPU already accesses system memory.