How llmfit Handles Unified Memory on Apple Silicon for LLM Inference
llmfit detects Apple Silicon by parsing system_profiler output, sets unified_memory: true, and treats total system RAM as the shared GPU/CPU memory pool to disable CPU-offload paths and optimize quantization selection.
The llmfit project implements specialized logic for Apple Silicon devices where the CPU and GPU share a single memory architecture. Unlike discrete GPU systems that separate VRAM from system RAM, llmfit recognizes the unified memory on Apple Silicon and adapts its model fitting algorithms to treat the entire memory pool as available for GPU inference via the Metal backend.
Detecting Apple Silicon Hardware
Parsing system_profiler for Apple GPU
The detection pipeline begins in llmfit-core/src/hardware.rs within the SystemSpecs::detect() method. When initializing hardware detection, llmfit calls detect_apple_gpu() at lines 72-80 to identify Apple Silicon chips:
// From llmfit-core/src/hardware.rs
pub fn detect_apple_gpu() -> Option<GpuInfo> {
// Executes: system_profiler SPDisplaysDataType
// Searches for "Apple M" or "Apple GPU" in output
}
This function executes system_profiler SPDisplaysDataType and scans the output for lines containing "Apple M" or "Apple GPU". When detected, the function returns the total system RAM (total_ram_gb) as the GPU's VRAM size and marks the device with unified_memory: true.
Setting the Unified Memory Flag
Upon successful detection at lines 87-94 in hardware.rs, llmfit constructs the SystemSpecs struct with three critical properties for Apple Silicon:
unified_memory: true— Indicates the shared memory architecturegpu_vram_gb: Some(total_ram_gb)— Treats total RAM as available GPU memorybackend: GpuBackend::Metal— Sets the compute backend to Metal
Additionally, when unified_memory is true, the system queries detect_gpu_available_gb() at lines 27-31 to retrieve the precise available memory figure from Metal's recommendedMaxWorkingSetSize, providing an accurate view of current memory pressure.
Memory Pool Management on Apple Silicon
Unified Pool Architecture
On Apple Silicon (M1/M2/M3 series), llmfit treats the physical RAM as a single shared resource. Because the GPU and CPU access the same memory pool, there is no physical distinction between "system RAM" and "GPU VRAM" from the application's perspective. The hardware detection logic ensures that gpu_vram_gb equals total_ram_gb, preventing artificial memory limits that would constrain model loading on devices with ample unified memory.
Model Fit Analysis and CPU Offload Disabling
Disabling CPU Offload Paths
In llmfit-core/src/fit.rs, the ModelFit::analyze() routine at lines 18-22 explicitly checks the system.unified_memory flag:
// From llmfit-core/src/fit.rs
if system.unified_memory {
// CPU-offload disabled: no separate RAM pool to spill to
// GPU and CPU share the same pool
}
When unified memory is present, llmfit disables the CPU-offload path because there is no secondary memory pool to spill model layers into. Traditional discrete GPU systems offload layers to system RAM when VRAM is exhausted, but this optimization is irrelevant on Apple Silicon where both processors share identical physical memory.
Quantization Selection
The fit analysis proceeds to select the optimal quantization level at lines 18-28. Because the GPU reports a VRAM size equal to total system RAM, the algorithm uses this unified pool size to determine which model quantization (Q4_0, Q5_K_M, etc.) will fit while leaving sufficient headroom for the operating system and other processes. If the unified pool is unavailable (an unlikely edge case on Apple Silicon), the routine falls back to pure-CPU inference.
Fit Scoring Against Unified Memory
The score_fit() function at lines 4-10 evaluates memory requirements against the unified pool by comparing mem_required versus mem_available. The resulting fit level (Perfect, Good, Tight, or TooLarge) reflects whether the model can execute within the total RAM/VRAM constraints of the Apple Silicon device, ensuring accurate performance predictions for Metal-accelerated inference.
Code Implementation Examples
Displaying System Specifications
The following Rust code retrieves and displays the unified memory configuration detected by llmfit:
let specs = llmfit_core::hardware::SystemSpecs::detect();
println!("Total RAM: {:.1} GB", specs.total_ram_gb);
println!("GPU backend: {:?}", specs.backend);
println!("Unified memory: {}", specs.unified_memory);
if let Some(vram) = specs.gpu_vram_gb {
println!("GPU VRAM (shared): {:.1} GB", vram);
}
Running Fit Analysis on Apple Silicon
This example demonstrates how llmfit analyzes model compatibility using the unified memory pool:
let specs = llmfit_core::hardware::SystemSpecs::detect();
let model = llmfit_core::models::LlmModel::from_name("tinyllama-1b").unwrap();
let fit = llmfit_core::fit::ModelFit::analyze(&model, &specs, None);
println!("Fit level: {:?}", fit.fit_level);
println!("Run mode: {:?}", fit.run_mode); // Returns RunMode::Gpu on unified memory
The output will indicate RunMode::Gpu when sufficient unified memory exists, bypassing CPU-offload considerations entirely.
Summary
- Hardware detection occurs via
detect_apple_gpu()inllmfit-core/src/hardware.rs, parsingsystem_profilerto identify Apple Silicon and setunified_memory: true - Memory representation treats
total_ram_gbas GPU VRAM, eliminating artificial boundaries between CPU and GPU memory pools - CPU-offload disabling happens automatically in
ModelFit::analyze()whenunified_memoryis detected, as there is no separate RAM pool for offloading - Fit scoring evaluates models against the total unified memory pool using
score_fit(), ensuring accurate capacity planning for Metal-backed inference
Frequently Asked Questions
How does llmfit detect Apple Silicon hardware?
llmfit executes system_profiler SPDisplaysDataType in the detect_apple_gpu() function and scans for "Apple M" or "Apple GPU" strings in the output. When found, it sets the unified_memory flag to true and configures the Metal backend.
Why does llmfit disable CPU offload on Apple Silicon?
CPU offload moves model layers from GPU VRAM to system RAM when memory is constrained. Because Apple Silicon shares a single physical memory pool between CPU and GPU, there is no distinct "system RAM" to offload to—both processors access the same memory. Disabling this path prevents unnecessary data movement overhead.
What happens if the unified memory pool is unavailable?
If detect_gpu_available_gb() returns no available memory (an edge case on Apple Silicon), ModelFit::analyze() falls back to pure-CPU inference mode. This ensures the application remains functional even if Metal memory reporting fails.
How does llmfit calculate available GPU memory on Metal?
When unified_memory is true, llmfit calls detect_gpu_available_gb() which queries Metal's recommendedMaxWorkingSetSize to determine the current available memory within the unified pool, providing accurate headroom calculations for model fitting algorithms.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →