Unified Memory Detection Mechanism in LLMFIT: Apple Silicon vs NVIDIA Grace Blackwell
LLMFIT detects unified memory architectures by querying system-specific APIs—system_profiler for Apple Silicon and nvidia-smi addressing modes for NVIDIA Grace—to flag shared RAM pools and adjust VRAM calculations accordingly.
The llmfit-core/src/hardware.rs module implements a cross-platform unified memory detection mechanism that determines whether a GPU shares physical RAM with the CPU. This distinction is critical for the AlexsJones/llmfit repository because model fit scoring treats the entire system memory as VRAM when unified memory is detected, avoiding incorrect capacity assumptions on Apple Silicon and NVIDIA Grace Blackwell SoCs.
How Unified Memory Detection Works in LLMFIT
The detection flow begins in SystemSpecs::detect(), which orchestrates GPU enumeration through detect_all_gpus(). This function assesses each hardware platform sequentially, applying platform-specific heuristics to identify unified memory architectures before constructing GpuInfo structs.
Apple Silicon Detection via system_profiler
For macOS systems, the code checks for Apple Silicon GPUs after querying other backends. The detect_apple_gpu(total_ram_gb) function executes system_profiler SPDisplaysDataType to read the recommendedMaxWorkingSetSize value.
// Inside detect_all_gpus() in llmfit-core/src/hardware.rs
if let Some(vram) = Self::detect_apple_gpu(total_ram_gb) {
let name = if cpu_name.to_lowercase().contains("apple") {
cpu_name.to_string()
} else {
"Apple Silicon".to_string()
};
gpus.push(GpuInfo {
name,
vram_gb: Some(vram), // VRAM equals total system RAM
backend: GpuBackend::Metal, // Metal is the only backend on macOS
count: 1,
unified_memory: true, // Flag shared memory architecture
});
}
Key implementation details include:
- Physical memory pool: The GPU draws from the same RAM as the CPU, so
detect_apple_gpuassignstotal_ram_gbas the VRAM value. - Backend assignment:
GpuBackend::Metalis hardcoded for Apple Silicon. - Detection location: The
system_profilerparsing logic resides around lines 710–750 inhardware.rs.
NVIDIA Grace Blackwell Detection via nvidia-smi
For NVIDIA Grace SoCs, detect_nvidia_gpus() issues an extended nvidia-smi query that includes the addressing_mode column.
// Extended query in detect_nvidia_gpus()
let output = std::process::Command::new("nvidia-smi")
.arg("--query-gpu=addressing_mode,memory.total,name")
.arg("--format=csv,noheader,nounits")
.output()
.ok()?;
The parse_nvidia_smi_extended() function interprets the addressing_mode field. If the mode equals ATS (Address Translation Services), the GPU uses unified memory:
let addr_mode = parts[0].trim();
let is_unified = addr_mode.eq_ignore_ascii_case("ATS");
// Aggregation logic later in the function
let entry = grouped.entry(name).or_insert((0, 0.0, false));
entry.0 += 1;
if vram_mb > entry.1 { entry.1 = vram_mb; }
if is_unified { entry.2 = true; } // Set unified flag
After building the GPU list, LLMFIT performs a final pass to correct VRAM values for unified memory devices:
let is_nvidia_unified = gpus.iter().any(|g| is_nvidia_unified_memory_gpu(&g.name));
if is_nvidia_unified {
for gpu in &mut gpus {
if is_nvidia_unified_memory_gpu(&gpu.name) {
gpu.unified_memory = true;
gpu.vram_gb = Some(total_ram_gb); // Treat RAM as VRAM
}
}
}
The helper is_nvidia_unified_memory_gpu (lines ~340–350) matches known Grace GPU identifiers such as "Grace" or specific device IDs like 10de:2e12.
Impact on Model Fit Scoring
The unified_memory boolean flag drives critical logic downstream. When true, the fit-scoring algorithm skips CPU-offload paths designed for discrete VRAM scenarios. This optimization recognizes that "spilling" to CPU RAM is meaningless when the GPU already accesses the same memory pool.
Additionally, gpu_available_gb queries (lines ~130–138) only execute for Apple Silicon devices, as macOS provides specific APIs for available GPU working set size that NVIDIA Grace platforms lack.
Practical Implementation Example
use llmfit_core::hardware::SystemSpecs;
let specs = SystemSpecs::detect();
for gpu in specs.gpus {
println!("GPU: {}", gpu.name);
println!(" Backend: {}", gpu.backend.label());
println!(" VRAM (GB): {:.2}", gpu.vram_gb.unwrap_or(0.0));
println!(" Unified memory? {}", gpu.unified_memory);
}
Running this snippet on an Apple Silicon Mac produces output similar to:
GPU: Apple M2 Pro
Backend: Metal
VRAM (GB): 32.00
Unified memory? true
On a Grace-based DGX Spark node, the output reflects the larger unified memory pool:
GPU: NVIDIA Grace
Backend: CUDA
VRAM (GB): 256.00
Unified memory? true
Summary
- Apple Silicon path: Uses
system_profiler SPDisplaysDataTypeto readrecommendedMaxWorkingSetSize, assigns total RAM as VRAM, and setsbackend: Metalwithunified_memory: true. - NVIDIA Grace path: Parses
nvidia-smiaddressing mode for "ATS", flags unified memory viais_nvidia_unified_memory_gpu, and overwrites VRAM with total system RAM. - Convergence: Both platforms populate
GpuInfowithunified_memory: true, enabling transparent handling of shared memory pools inllmfit-core/src/hardware.rs. - Downstream effects: The flag prevents incorrect CPU-offload calculations and ensures accurate model sizing on unified memory architectures.
Frequently Asked Questions
How does LLMFIT distinguish between discrete and unified memory GPUs?
LLMFIT checks platform-specific indicators. For Apple Silicon, it detects the Metal backend and queries macOS system profiler data. For NVIDIA, it checks the addressing_mode field for "ATS" (Address Translation Services) in nvidia-smi output. Discrete GPUs report separate VRAM values and lack these unified memory indicators.
Why does VRAM equal total RAM on unified memory systems?
Because the GPU and CPU share the same physical memory pool. On Apple Silicon and NVIDIA Grace, there is no separate VRAM chip—the GPU allocates from system RAM. LLMFIT reflects this hardware reality by assigning total_ram_gb to the vram_gb field when unified_memory is true.
Where is the unified memory detection logic located in the source code?
The primary implementation resides in llmfit-core/src/hardware.rs. Key functions include detect_apple_gpu() around lines 710–750, parse_nvidia_smi_extended() for ATS detection, and is_nvidia_unified_memory_gpu() around lines 340–350.
Does unified memory detection affect model loading strategy?
Yes. When unified_memory is true, fit scoring skips CPU-offload paths designed for discrete VRAM scenarios. This prevents incorrect capacity calculations during model fitting because there is no separate RAM pool to spill to when the GPU already accesses system memory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →