How llmfit Detects Local Hardware Specifications: RAM, CPU, and GPU VRAM
llmfit obtains the host's complete hardware profile through a single call to SystemSpecs::detect() in llmfit-core/src/hardware.rs, combining the cross-platform sysinfo crate for system memory and CPU detection with vendor-specific command-line probes for GPU VRAM discovery.
Understanding how llmfit detect local hardware specifications works is essential for developers optimizing LLM inference across diverse hardware. The Rust-based llmfit repository uses a unified detection pipeline that aggregates RAM, CPU cores, and GPU capabilities into a single SystemSpecs struct. This architecture enables automatic backend selection and memory-constrained model fitting across platforms including Linux, Windows, macOS, and specialized hardware like Ascend NPUs.
SystemSpecs::detect() Entry Point
The detection process begins in llmfit-core/src/hardware.rs with the SystemSpecs::detect() method. This function orchestrates a three-stage pipeline: first gathering CPU and RAM statistics via the sysinfo crate, then executing platform-specific GPU discovery routines, and finally normalizing unified-memory systems where GPUs share system RAM.
RAM and CPU Detection Using sysinfo
For system memory and processor information, llmfit relies on the sysinfo crate. The implementation calls System::new_all() to refresh all system information, then extracts total_memory() and available_memory() values.
On macOS, where available_memory() may return zero, the code falls back to Self::available_ram_fallback() (lines 80-84). The detection also captures the CPU name and total core count from the same sysinfo query (lines 74-90).
GPU VRAM Detection Pipeline
GPU discovery follows a prioritized fallback chain. The detect_all_gpus() function (line 98) aggregates results from multiple vendor-specific methods, returning a Vec<GpuInfo> sorted by VRAM capacity (largest first).
NVIDIA GPU Detection (nvidia-smi and SysFS)
For NVIDIA hardware, llmfit first attempts detect_nvidia_gpus(), which executes nvidia-smi with CSV output. The try_nvidia_smi_with_addressing_mode() function checks the addressing_mode column to detect unified-memory configurations like ATS (Address Translation Services) on NVIDIA Grace SoCs (lines 123-131).
If the CLI tool is unavailable, the code falls back to detect_nvidia_gpu_sysfs_info(), reading /sys/class/drm/card*/device/mem_info_vram_total directly from the Linux kernel driver. This path also uses lspci to resolve GPU names when the driver is present (lines 95-119).
AMD GPU Detection (ROCm and SysFS)
AMD graphics cards are detected through detect_amd_gpu_rocm_info(), which parses rocm-smi output for VRAM utilization (--showmeminfo vram) and product names (--showproductname). The parser handles both block-style and tabular output formats, then groups identical models to count multi-GPU setups.
Without ROCm, the system falls back to detect_amd_gpu_sysfs_info(), scanning /sys/class/drm for vendor ID 0x1002, reading VRAM values via sysfs, and filtering out integrated iGPUs when discrete cards are present (lines 132-150).
Intel and Windows GPU Queries
On Windows, detect_gpu_windows_info() executes PowerShell queries against Win32_VideoController to retrieve adapter names and AdapterRAM values. The implementation applies registry-based VRAM fixes via apply_registry_vram() and explicitly filters out integrated Intel HD graphics as unsupported (lines 156-166).
Apple Silicon and macOS Metal Detection
For macOS, detect_macos_metal_gpus() invokes system_profiler SPDisplaysDataType to enumerate Metal-compatible discrete GPUs from AMD and Intel. On Apple Silicon systems, detect_apple_gpu() reads the "recommendedMaxWorkingSetSize" from Metal and treats the entire system RAM as VRAM, setting unified_memory = true. The GPU name defaults to the CPU name when it contains "Apple" (lines 176-184).
Ascend NPU and Vulkan Fallback
Specialized hardware like Huawei Ascend NPUs is handled by detect_ascend_npus(), which executes npu-smi and maps results to GpuInfo with GpuBackend::Ascend (lines 188-192). When all vendor-specific methods fail, the system uses a detect_vulkan_gpu_info() fallback to ensure at least one generic GPU backend is reported (lines 194-196).
Unified Memory Handling
Modern SoCs like Apple Silicon M-series, AMD APUs, and NVIDIA Grace chips implement unified memory architectures where the GPU shares the system RAM pool. The detection logic identifies these configurations through specialized checks: is_nvidia_unified_memory_gpu() for Grace SoCs, is_amd_unified_memory_apu() for AMD integrated graphics, and specific Apple Silicon detection logic.
When unified memory is detected, the code sets unified_memory = true and substitutes the system RAM size for the GPU VRAM value, ensuring memory calculations treat the entire RAM pool as available for model inference.
GPU Aggregation and Selection Logic
After collecting candidate GPUs, detect_all_gpus() performs three normalization steps:
- Merges duplicate sources via
merge_gpu_sources()to eliminate redundant entries from multiple detection methods - Prefers discrete GPUs over integrated graphics using
prefer_discrete_gpus() - Calculates totals by summing VRAM across all cards into
total_gpu_vram_gb
The final vector is sorted in descending VRAM order, making the first entry the primary GPU used for inference (lines 306-321).
Code Example: Retrieving System Specs
The following Rust code demonstrates how to consume the hardware detection API:
use llmfit_core::hardware::SystemSpecs;
// Retrieve the complete hardware specification
let specs = SystemSpecs::detect();
println!("🧮 RAM: {:.2} GB total, {:.2} GB available",
specs.total_ram_gb, specs.available_ram_gb);
println!("⚙️ CPU: {} ({} cores)", specs.cpu_name, specs.total_cpu_cores);
if specs.has_gpu {
println!("🎮 Primary GPU: {}", specs.gpu_name.unwrap_or_else(|| "unknown".into()));
println!("💾 VRAM (primary): {:.2} GB", specs.gpu_vram_gb.unwrap_or(0.0));
if let Some(total) = specs.total_gpu_vram_gb {
println!("📊 Total VRAM across all GPUs: {:.2} GB", total);
}
if specs.unified_memory {
println!("🔗 Unified memory system – GPU can use the full RAM pool");
}
} else {
println!("🚫 No discrete GPU detected");
}
On a typical laptop, this outputs discrete VRAM values like 6.00 GB. On Apple Silicon, it reports unified_memory = true with VRAM equal to total system RAM.
Summary
- Centralized detection: All hardware profiling flows through
SystemSpecs::detect()inllmfit-core/src/hardware.rs - Cross-platform RAM/CPU: Uses the
sysinfocrate with macOS-specific fallbacks for available memory - Multi-vendor GPU support: Implements specific detection paths for NVIDIA (nvidia-smi/sysfs), AMD (ROCm/sysfs), Intel (Windows WMI), and Apple (Metal/system_profiler)
- Unified memory awareness: Detects shared-memory architectures on Apple Silicon, AMD APUs, and NVIDIA Grace SoCs, substituting system RAM for VRAM calculations
- Intelligent aggregation: Merges duplicate sources, prefers discrete GPUs, sorts by VRAM capacity, and calculates total available GPU memory across multi-card setups
Frequently Asked Questions
How does llmfit handle macOS memory reporting limitations?
On macOS, the sysinfo crate may return zero for available memory. llmfit implements available_ram_fallback() (lines 80-84 in hardware.rs) to provide accurate available RAM calculations when the standard API fails, ensuring system specs remain accurate across Darwin platforms.
Can llmfit detect multiple GPUs in a single system?
Yes. The detect_all_gpus() function aggregates all discovered GPUs into a Vec<GpuInfo>, groups identical models, sums total VRAM across all cards, and sorts results by memory capacity. This enables llmfit to calculate total_gpu_vram_gb for multi-GPU inference scenarios while selecting the largest VRAM card as the primary backend.
What happens if no GPU is detected on the system?
If all vendor-specific detection methods (NVIDIA, AMD, Intel, Apple, Ascend) fail, llmfit falls back to detect_vulkan_gpu_info() to provide a generic GPU entry. This ensures the SystemSpecs struct always contains at least one backend identifier, preventing runtime panics while clearly indicating limited GPU acceleration availability.
How does llmfit distinguish between discrete and integrated graphics?
The detection pipeline calls prefer_discrete_gpus() after merging sources. For AMD systems, detect_amd_gpu_sysfs_info() explicitly filters out iGPUs when discrete cards are present. On Windows, Intel HD integrated graphics are filtered out as unsupported, while discrete AMD and NVIDIA cards are prioritized in the final GPU selection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →