How llmfit Detects GPU VRAM Using nvidia-smi, rocm-smi, system_profiler, and objc2-metal
llmfit determines available GPU VRAM by executing platform-specific system commands—nvidia-smi for NVIDIA GPUs, rocm-smi for AMD GPUs on Linux, and system_profiler or the objc2-metal crate for macOS—then parsing the output to populate a GpuInfo struct used throughout the application.
The llmfit repository (AlexsJones/llmfit) implements robust cross-platform GPU memory detection in llmfit-core/src/hardware.rs. The detection engine shells out to native system utilities, parses CSV or structured text output, and aggregates results into a SystemSpecs struct that downstream components use for model-fit calculations and hardware validation.
NVIDIA GPU VRAM Detection with nvidia-smi
On Linux and Windows systems with NVIDIA hardware, llmfit relies on the nvidia-smi utility to query dedicated video memory.
Standard and Extended CSV Queries
The primary detection path executes nvidia-smi with specific query parameters to retrieve memory totals and device names:
nvidia-smi --query-gpu=memory.total,name --format=csv,noheader,nounits
For edge cases like DGX or Jetson devices where the standard query reports 0 GB, the code falls back to an extended query that includes the addressing_mode field. This extended form helps identify legacy drivers or unusual memory configurations that cannot expose VRAM through standard APIs.
Parsing Implementation
In llmfit-core/src/hardware.rs, the SystemSpecs::parse_nvidia_smi_list and parse_nvidia_smi_extended functions handle CSV parsing. These methods extract the memory.total value, convert it to gigabytes, and instantiate GpuInfo structs with the GPU name and VRAM capacity.
// From llmfit-core/src/hardware.rs
let output = Command::new("nvidia-smi")
.args(&["--query-gpu=memory.total,name", "--format=csv,noheader,nounits"])
.output()?;
let nvidia_gpus = SystemSpecs::parse_nvidia_smi_list(
&String::from_utf8_lossy(&output.stdout)
);
Linux Sysfs Fallback
When nvidia-smi is unavailable or returns zero values, llmfit invokes detect_nvidia_gpu_sysfs_info to read PCI sysfs entries directly from /sys/class/pci_bus/. This fallback traverses the PCI device tree to synthesize a VRAM estimate based on the GPU's BAR memory reporting, ensuring detection works on headless servers without full NVIDIA driver utilities installed.
AMD GPU VRAM Detection with rocm-smi
For AMD GPUs on Linux, llmfit queries the ROCm management interface via rocm-smi. The detection logic shells out to:
rocm-smi -i
The parser scans the command output for the "GPU Memory Total" field. Because rocm-smi reports memory in MiB, hardware.rs converts the value to GB before storing it in the GpuInfo struct.
// Internal implementation in llmfit-core/src/hardware.rs
let output = Command::new("rocm-smi").arg("-i").output()?;
// Parse "GPU Memory Total" field, convert MiB to GB
This approach supports modern AMD datacenter and consumer GPUs running the ROCm stack, providing accurate VRAM figures for model batch-size calculations.
macOS GPU VRAM Detection
macOS detection uses two complementary strategies depending on the hardware generation and driver availability.
system_profiler SPDisplaysDataType
On Intel-based Macs and older systems, llmfit executes:
system_profiler SPDisplaysDataType
The hardware.rs parser scans the structured text output for the "VRAM (Total)" line. When found, the value is parsed and recorded as the discrete GPU's memory capacity.
objc2-metal Working-Set Queries
For Metal-enabled GPUs or when system_profiler provides insufficient detail, llmfit uses the objc2-metal crate (declared in Cargo.toml) to query the Metal API directly. This method retrieves the "effective Metal working-set limit"—the maximum buffer size the Metal driver will allocate—which serves as a reliable VRAM estimate for Intel integrated graphics and discrete AMD GPUs in Mac systems.
// Example: macOS detection path
let mac_gpu = SystemSpecs::detect_macos_gpu_vram()?;
// Uses objc2-metal to query Metal working-set limit when available
Unified Memory Handling on Apple Silicon
Apple Silicon Macs (M1, M2, M3 series) use unified memory architecture where system RAM serves as both CPU and GPU memory. In this case, llmfit treats the reported system RAM as VRAM and sets the GpuInfo.unified_memory flag to true. The detection logic skips the CpuOffload path because no separate VRAM pool exists, preventing unnecessary memory-copy optimizations.
Integration and Aggregation in SystemSpecs
All detection routines feed into SystemSpecs::detect(), the central aggregation method in hardware.rs. This function:
- Probes the operating system to identify the platform (NVIDIA, AMD, or macOS).
- Executes the appropriate platform-specific tool (
nvidia-smi,rocm-smi, orsystem_profiler/objc2-metal). - Parses the raw output into
GpuInfostructs containingname,vram_gb, and architecture flags. - Falls back to sysfs parsing or extended queries when primary tools report zero values.
- Aggregates CPU, RAM, and GPU information into a single
SystemSpecsinstance consumed by the doctor diagnostics (llmfit-core/src/doctor.rs) and model-fit analysis modules.
Unit tests in llmfit-core/tests/fixtures/hardware/nvidia/ validate the CSV parsing logic using fixture files like standard-identical.csv and extended-unified.csv, ensuring robustness across different nvidia-smi output formats.
Summary
llmfit-core/src/hardware.rscontains the complete VRAM detection implementation for NVIDIA, AMD, and macOS platforms.- NVIDIA detection uses
nvidia-smiCSV output withparse_nvidia_smi_list, falling back todetect_nvidia_gpu_sysfs_infowhen necessary. - AMD detection relies on
rocm-smi -ioutput parsing to extract "GPU Memory Total" values. - macOS detection combines
system_profiler SPDisplaysDataTypescanning withobjc2-metalworking-set queries for Metal-enabled devices. - Apple Silicon sets the
unified_memoryflag and uses system RAM as VRAM, skipping discrete GPU offload logic. SystemSpecs::detect()aggregates all hardware discovery into a unified struct for downstream consumption.
Frequently Asked Questions
What happens if nvidia-smi is not installed on a Linux system?
If nvidia-smi is missing or returns zero values, llmfit invokes detect_nvidia_gpu_sysfs_info to read PCI sysfs entries in /sys/class/pci_bus/. This fallback parses BAR memory information directly from the kernel's PCI subsystem, allowing VRAM detection on headless servers or minimal container environments without the full NVIDIA driver toolkit.
How does llmfit handle unified memory architectures like Apple Silicon?
On Apple Silicon Macs, llmfit detects the unified memory architecture and sets GpuInfo.unified_memory to true. The system treats the total system RAM as available GPU memory, and the application logic skips the CpuOffload optimization path since there is no separate VRAM pool requiring explicit memory management.
Why does llmfit use objc2-metal instead of system_profiler on some macOS systems?
The objc2-metal crate queries the Metal API for the effective working-set limit, which provides a more accurate representation of available GPU memory than system_profiler on certain Intel Macs with discrete AMD GPUs. This method reflects the actual buffer allocation limits imposed by the Metal driver, rather than the theoretical hardware maximum.
Can llmfit detect VRAM on headless servers without display drivers?
Yes. On Linux systems without display drivers but with NVIDIA hardware exposed via PCI, the detect_nvidia_gpu_sysfs_info fallback reads kernel sysfs entries to determine VRAM capacity. Similarly, AMD detection via rocm-smi does not require an active display connection, only the ROCm kernel modules. MacOS detection requires either system_profiler or Metal drivers, which are present in standard macOS installations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →