How Apple Silicon Unified Memory Affects llmfit Run Modes: Detection & Execution Logic
Apple Silicon's unified memory architecture causes llmfit to bypass the CPU-offload path and choose between GPU-only or CPU-only execution modes.
The llmfit crate by AlexsJones dynamically adapts its model execution strategy based on hardware detection. On Apple Silicon Macs, where CPU and GPU share a single memory pool, the traditional CPU-offload optimization becomes redundant. This article examines how llmfit-core detects unified memory and alters its run mode selection accordingly.
Detecting Apple Silicon Unified Memory in llmfit
The hardware detection logic lives in llmfit-core/src/hardware.rs. During SystemSpecs::detect(), the code executes macOS-specific system profiling to identify unified memory capabilities.
Detection Mechanism
The detection occurs at lines 78-91, where llmfit shells out to system_profiler and parses the output:
// Detect Apple Silicon unified memory.
let unified_memory = if cfg!(target_os = "macos") {
// Use system_profiler to detect if unified memory is present.
if let Ok(output) = std::process::Command::new("system_profiler")
.arg("SPDisplaysDataType")
.output()
{
let stdout = String::from_utf8_lossy(&output.stdout);
stdout.contains("Unified Memory")
} else {
false
}
} else {
false
};
When unified memory is detected, llmfit remaps the VRAM metric to system RAM (lines 93-95):
// On unified memory Macs, VRAM is effectively the same as system RAM.
if unified_memory {
total_vram_gb = ram_gb;
}
This remapping ensures downstream calculations treat the entire memory pool as available GPU-accessible memory.
How Unified Memory Alters Run Mode Selection
The run mode selection logic resides in llmfit-core/src/fit.rs, specifically within ModelFit::choose_best_mode() at lines 86-114.
Standard Non-Unified Memory Path
On traditional systems with discrete GPUs, llmfit evaluates this fallback chain:
- Gpu mode — if
model.min_vram_gb <= sys.total_vram_gb - CpuOffload mode — if
model.min_vram_gb + model.min_ram_gb * 0.5 <= sys.total_vram_gb - CpuOnly mode — if
model.min_ram_gb <= sys.ram_gb
Unified Memory Bypass
The Apple Silicon path short-circuits this logic at lines 91-96:
// If unified memory is present (Apple Silicon), we treat VRAM as RAM.
if sys.unified_memory {
// On unified memory Macs, CPU-offload path is skipped because VRAM == RAM.
if model.min_ram_gb <= sys.ram_gb {
return (RunMode::CpuOnly, model.min_ram_gb);
}
}
Key implications:
RunMode::CpuOffloadis never selected — the hybrid split between VRAM and RAM serves no purpose when both subsystems share identical physical memory- Memory calculation simplifies — only
min_ram_gbis evaluated againstram_gb - Execution falls through to CPU-only or continues to other fallback candidates
Complete Run Mode Selection Flow
The RunMode enum defines five execution strategies in llmfit-core/src/fit.rs:
| RunMode | Description | Unified Memory Behavior |
|---|---|---|
Gpu |
Weights fully resident in GPU memory | Available if model fits in unified pool |
MoeOffload |
Mixture-of-Experts partial loading | Available via forced mode or fallback |
CpuOffload |
Model split between GPU and system RAM | Skipped entirely |
CpuOnly |
Entire model in system RAM | Primary fallback on unified memory |
TensorParallel |
Multi-node tensor parallelism | Available for distributed setups |
Practical Code Example
The following demonstrates detection and mode selection on Apple Silicon:
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::fit::{ModelFit, RunMode};
use llmfit_core::models::ModelInfo;
fn main() {
// Detect hardware — unified_memory true on Apple Silicon Macs
let specs = SystemSpecs::detect();
println!("Apple Silicon unified memory: {}", specs.unified_memory);
println!("Available memory pool: {:.1} GB", specs.ram_gb);
// Model requiring 16 GB VRAM, 8 GB RAM
let model = ModelInfo::dummy_with_vram(16.0, 8.0);
// llmfit selects optimal execution mode
let fit = ModelFit::analyze_with_forced_runtime(model, &specs, None);
println!("Selected run mode: {}", fit.run_mode.name());
println!("Estimated memory: {:.1} GB", fit.memory_gb);
}
Sample output on MacBook Pro (M3, 36 GB unified memory):
Apple Silicon unified memory: true
Available memory pool: 36.0 GB
Selected run mode: CPU-Only
Estimated memory: 8.0 GB
Notice that despite the model's 16 GB "VRAM" requirement, the unified memory detection routes execution through CpuOnly rather than attempting CPU-offload split calculations.
Performance Characteristics by Mode
The throughput estimation function at lines 116-125 assigns efficiency multipliers:
fn estimate_throughput(model: &ModelInfo, mode: RunMode, sys: &SystemSpecs) -> f64 {
match mode {
RunMode::Gpu => model.recommended_throughput_gb() * 1.0,
RunMode::MoeOffload => model.recommended_throughput_gb() * 0.9,
RunMode::CpuOffload => model.recommended_throughput_gb() * 0.6,
RunMode::CpuOnly => model.recommended_throughput_gb() * 0.4,
RunMode::TensorParallel => model.recommended_throughput_gb() * 0.8,
}
}
On Apple Silicon, the CpuOnly mode (0.4x baseline) may underperform relative to Metal GPU acceleration. However, llmfit allows forced mode selection via analyze_with_forced_runtime() for users who explicitly want GPU execution through frameworks like MLX or PyTorch with Metal backend.
Summary
- Detection source:
llmfit-core/src/hardware.rsusessystem_profiler SPDisplaysDataTypeto identify "Unified Memory" string on macOS - Memory remapping:
total_vram_gbbecomesram_gbwhen unified memory is detected - Execution impact:
RunMode::CpuOffloadis bypassed; selection proceeds directly toCpuOnlyor other fallbacks - User control:
forcedparameter inanalyze_with_forced_runtime()overrides automatic selection
Frequently Asked Questions
How does llmfit detect Apple Silicon specifically?
llmfit does not explicitly check for the M1/M2/M3 chip identifier. Instead, it detects the unified memory capability through system_profiler output parsing. This approach correctly identifies both Apple Silicon Macs and any future macOS systems with unified memory architectures.
Can I force GPU mode on Apple Silicon despite llmfit selecting CPU-only?
Yes. Pass Some(RunMode::Gpu) to analyze_with_forced_runtime(). The viability check at line 63 (needed <= sys.total_vram_gb) will pass because unified memory detection remaps total_vram_gb to your full system RAM. This enables Metal-backed GPU execution when the underlying inference framework supports it.
Why does llmfit skip CPU-offload on unified memory systems?
CPU-offload optimizes for the bandwidth and latency differential between discrete GPU VRAM and slower system RAM. On Apple Silicon, CPU and GPU access the same physical memory pool at identical speeds, eliminating any benefit from partial GPU residence. The CpuOnly mode provides equivalent memory access patterns without the complexity of layer splitting.
What happens if a model exceeds available unified memory?
The selection algorithm falls through to the minimal-memory candidate sorting at lines 108-113. llmfit will still recommend a run mode, but actual execution will likely trigger system-level memory swapping or framework-specific out-of-memory errors. The tool provides memory estimates; enforcement depends on the downstream inference engine.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →