How SystemSpecs::detect Distinguishes Apple Silicon Unified Memory from Discrete VRAM in llmfit
SystemSpecs::detect() identifies Apple Silicon by querying Metal's recommendedMaxWorkingSetSize via system_profiler, marks the GPU with unified_memory: true, and automatically skips the CpuOffload execution path because unified memory systems lack separate RAM and VRAM pools.
The SystemSpecs::detect() method in the AlexsJones/llmfit crate performs hardware introspection to determine optimal model execution strategies. When running on macOS, it specifically differentiates Apple Silicon's unified memory architecture from traditional discrete GPU setups. This distinction directly impacts which run-modes the planner considers viable, particularly excluding CpuOffload on systems where CPU and GPU share the same memory pool.
Detecting Apple Silicon Unified Memory
The detection logic resides in llmfit-core/src/hardware.rs, where SystemSpecs::detect() delegates GPU enumeration to detect_all_gpus(). This function treats Apple Silicon as a special case distinct from discrete GPUs.
Querying Metal via system_profiler
For macOS targets, the code invokes detect_apple_gpu(), which executes the macOS system_profiler SPDisplaysDataType command. It extracts the recommendedMaxWorkingSetSize reported by Metal, representing the portion of the shared memory pool that the GPU may actually use.
// Apple Silicon (unified memory)
if let Some(vram) = Self::detect_apple_gpu(total_ram_gb) {
let name = if cpu_name.to_lowercase().contains("apple") {
cpu_name.to_string()
} else {
"Apple Silicon".to_string()
};
gpus.push(GpuInfo {
name,
vram_gb: Some(vram),
backend: GpuBackend::Metal,
count: 1,
unified_memory: true,
});
}
Because Apple Silicon utilizes a single unified memory pool, the function returns the total system RAM as the "VRAM" size and explicitly sets unified_memory: true in the GpuInfo struct.
Populating the unified_memory Flag
After detecting the primary GPU, the code propagates the unified_memory flag to the top-level SystemSpecs struct. As implemented in hardware.rs at lines 445-458, this flag indicates whether the system uses a shared memory architecture:
let unified_memory = primary.map(|g| g.unified_memory).unwrap_or(false);
…
SystemSpecs {
…,
unified_memory,
…
}
Impact on CpuOffload Skip Path
The presence of unified memory eliminates the CpuOffload execution path because there is no distinct RAM pool to spill model weights to. The planner handles this logic in llmfit-core/src/plan.rs.
The Planning Logic Guard in plan.rs
Around line 400 in plan.rs, a guard clause skips the CpuOffload evaluation when system.unified_memory is true:
// Skip CpuOffload on unified‑memory systems (Apple Silicon, AMD/APU, NVIDIA Grace)
if !system.unified_memory {
// consider CpuOffload …
}
According to the llmfit source code, this check ensures that CpuOffload only activates when VRAM and RAM are separate physical resources, which is not the case on Apple Silicon.
Available Run-Modes on Apple Silicon
Consequently, on Apple Silicon systems the planner evaluates only two execution strategies:
- Gpu – The model resides entirely in the unified memory pool (treated as VRAM)
- CpuOnly – The model runs purely on CPU without utilizing the Metal backend
The CpuOffload mode, which typically splits work between discrete GPU VRAM and system RAM, is omitted from consideration.
Verification Examples
You can verify the detection behavior at runtime using the following patterns:
// Example: printing whether the current machine uses unified memory
let specs = llmfit_core::hardware::SystemSpecs::detect();
println!(
"Unified memory: {}, backend: {}",
specs.unified_memory,
specs.backend.label()
);
// Fit‑planning: the planner automatically avoids CpuOffload on Apple Silicon
let plan = llmfit_core::plan::Plan::new(&model, &specs);
assert!(!plan.run_modes().contains(&RunMode::CpuOffload));
Summary
- SystemSpecs::detect() calls
detect_apple_gpu()on macOS to query Metal's memory statistics viasystem_profiler SPDisplaysDataType. - Unified memory detection sets the
unified_memory: trueflag inGpuInfo, which propagates to theSystemSpecsstruct. - CpuOffload exclusion occurs in
plan.rswhensystem.unified_memoryis true, preventing the planner from considering execution modes that require separate RAM and VRAM pools. - Apple Silicon systems only support
GpuandCpuOnlyrun-modes, utilizingGpuBackend::Metalfor GPU acceleration.
Frequently Asked Questions
How does llmfit distinguish Apple Silicon from discrete GPUs?
The detect_apple_gpu() function executes system_profiler SPDisplaysDataType and parses the recommendedMaxWorkingSetSize value from Metal's report. It returns the total system RAM as the VRAM capacity and sets unified_memory: true, whereas discrete GPUs would report separate VRAM values and unified_memory: false.
Why is CpuOffload skipped on Apple Silicon Macs?
CpuOffload requires distinct memory pools to transfer weights between CPU RAM and GPU VRAM. Since Apple Silicon uses a unified memory architecture where the CPU and GPU share the same physical memory, there is no benefit to offloading between separate pools. The planner in plan.rs explicitly skips this mode when system.unified_memory is true.
What execution modes are available on Apple Silicon?
On Apple Silicon systems, the planner only considers Gpu mode (running the model entirely in unified memory using Metal) and CpuOnly mode (running purely on CPU without GPU acceleration). The CpuOffload mode is excluded from the available run-modes list.
How can I verify if my system is detected as unified memory?
Instantiate SystemSpecs::detect() and check the unified_memory boolean field. If true, and the backend is GpuBackend::Metal, your system is correctly identified as an Apple Silicon unified memory architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →