Unified Memory Detection Mechanism in LLMFIT: Apple Silicon vs NVIDIA Grace Blackwell

LLMFIT detects unified memory architectures by querying system-specific APIs—system_profiler for Apple Silicon and nvidia-smi addressing modes for NVIDIA Grace—to flag shared RAM pools and adjust VRAM calculations accordingly.

The llmfit-core/src/hardware.rs module implements a cross-platform unified memory detection mechanism that determines whether a GPU shares physical RAM with the CPU. This distinction is critical for the AlexsJones/llmfit repository because model fit scoring treats the entire system memory as VRAM when unified memory is detected, avoiding incorrect capacity assumptions on Apple Silicon and NVIDIA Grace Blackwell SoCs.

How Unified Memory Detection Works in LLMFIT

The detection flow begins in SystemSpecs::detect(), which orchestrates GPU enumeration through detect_all_gpus(). This function assesses each hardware platform sequentially, applying platform-specific heuristics to identify unified memory architectures before constructing GpuInfo structs.

Apple Silicon Detection via system_profiler

For macOS systems, the code checks for Apple Silicon GPUs after querying other backends. The detect_apple_gpu(total_ram_gb) function executes system_profiler SPDisplaysDataType to read the recommendedMaxWorkingSetSize value.

// Inside detect_all_gpus() in llmfit-core/src/hardware.rs
if let Some(vram) = Self::detect_apple_gpu(total_ram_gb) {
    let name = if cpu_name.to_lowercase().contains("apple") {
        cpu_name.to_string()
    } else {
        "Apple Silicon".to_string()
    };
    gpus.push(GpuInfo {
        name,
        vram_gb: Some(vram),          // VRAM equals total system RAM
        backend: GpuBackend::Metal,   // Metal is the only backend on macOS
        count: 1,
        unified_memory: true,         // Flag shared memory architecture
    });
}

Key implementation details include:

  • Physical memory pool: The GPU draws from the same RAM as the CPU, so detect_apple_gpu assigns total_ram_gb as the VRAM value.
  • Backend assignment: GpuBackend::Metal is hardcoded for Apple Silicon.
  • Detection location: The system_profiler parsing logic resides around lines 710–750 in hardware.rs.

NVIDIA Grace Blackwell Detection via nvidia-smi

For NVIDIA Grace SoCs, detect_nvidia_gpus() issues an extended nvidia-smi query that includes the addressing_mode column.

// Extended query in detect_nvidia_gpus()
let output = std::process::Command::new("nvidia-smi")
    .arg("--query-gpu=addressing_mode,memory.total,name")
    .arg("--format=csv,noheader,nounits")
    .output()
    .ok()?;

The parse_nvidia_smi_extended() function interprets the addressing_mode field. If the mode equals ATS (Address Translation Services), the GPU uses unified memory:

let addr_mode = parts[0].trim();
let is_unified = addr_mode.eq_ignore_ascii_case("ATS");

// Aggregation logic later in the function
let entry = grouped.entry(name).or_insert((0, 0.0, false));
entry.0 += 1;
if vram_mb > entry.1 { entry.1 = vram_mb; }
if is_unified { entry.2 = true; }   // Set unified flag

After building the GPU list, LLMFIT performs a final pass to correct VRAM values for unified memory devices:

let is_nvidia_unified = gpus.iter().any(|g| is_nvidia_unified_memory_gpu(&g.name));
if is_nvidia_unified {
    for gpu in &mut gpus {
        if is_nvidia_unified_memory_gpu(&gpu.name) {
            gpu.unified_memory = true;
            gpu.vram_gb = Some(total_ram_gb);   // Treat RAM as VRAM
        }
    }
}

The helper is_nvidia_unified_memory_gpu (lines ~340–350) matches known Grace GPU identifiers such as "Grace" or specific device IDs like 10de:2e12.

Impact on Model Fit Scoring

The unified_memory boolean flag drives critical logic downstream. When true, the fit-scoring algorithm skips CPU-offload paths designed for discrete VRAM scenarios. This optimization recognizes that "spilling" to CPU RAM is meaningless when the GPU already accesses the same memory pool.

Additionally, gpu_available_gb queries (lines ~130–138) only execute for Apple Silicon devices, as macOS provides specific APIs for available GPU working set size that NVIDIA Grace platforms lack.

Practical Implementation Example

use llmfit_core::hardware::SystemSpecs;

let specs = SystemSpecs::detect();

for gpu in specs.gpus {
    println!("GPU: {}", gpu.name);
    println!("  Backend: {}", gpu.backend.label());
    println!("  VRAM (GB): {:.2}", gpu.vram_gb.unwrap_or(0.0));
    println!("  Unified memory? {}", gpu.unified_memory);
}

Running this snippet on an Apple Silicon Mac produces output similar to:


GPU: Apple M2 Pro
  Backend: Metal
  VRAM (GB): 32.00
  Unified memory? true

On a Grace-based DGX Spark node, the output reflects the larger unified memory pool:


GPU: NVIDIA Grace
  Backend: CUDA
  VRAM (GB): 256.00
  Unified memory? true

Summary

  • Apple Silicon path: Uses system_profiler SPDisplaysDataType to read recommendedMaxWorkingSetSize, assigns total RAM as VRAM, and sets backend: Metal with unified_memory: true.
  • NVIDIA Grace path: Parses nvidia-smi addressing mode for "ATS", flags unified memory via is_nvidia_unified_memory_gpu, and overwrites VRAM with total system RAM.
  • Convergence: Both platforms populate GpuInfo with unified_memory: true, enabling transparent handling of shared memory pools in llmfit-core/src/hardware.rs.
  • Downstream effects: The flag prevents incorrect CPU-offload calculations and ensures accurate model sizing on unified memory architectures.

Frequently Asked Questions

How does LLMFIT distinguish between discrete and unified memory GPUs?

LLMFIT checks platform-specific indicators. For Apple Silicon, it detects the Metal backend and queries macOS system profiler data. For NVIDIA, it checks the addressing_mode field for "ATS" (Address Translation Services) in nvidia-smi output. Discrete GPUs report separate VRAM values and lack these unified memory indicators.

Why does VRAM equal total RAM on unified memory systems?

Because the GPU and CPU share the same physical memory pool. On Apple Silicon and NVIDIA Grace, there is no separate VRAM chip—the GPU allocates from system RAM. LLMFIT reflects this hardware reality by assigning total_ram_gb to the vram_gb field when unified_memory is true.

Where is the unified memory detection logic located in the source code?

The primary implementation resides in llmfit-core/src/hardware.rs. Key functions include detect_apple_gpu() around lines 710–750, parse_nvidia_smi_extended() for ATS detection, and is_nvidia_unified_memory_gpu() around lines 340–350.

Does unified memory detection affect model loading strategy?

Yes. When unified_memory is true, fit scoring skips CPU-offload paths designed for discrete VRAM scenarios. This prevents incorrect capacity calculations during model fitting because there is no separate RAM pool to spill to when the GPU already accesses system memory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →