How llmfit Detects VRAM Across CUDA, ROCm, Vulkan, and SYCL GPU Backends
llmfit detects VRAM by probing vendor-specific utilities and OS interfaces, then merging and deduplicating results from multiple sources to handle cases where one backend reports names but another provides memory sizes.
The llmfit open-source project provides unified GPU detection for large language model inference. According to the AlexsJones/llmfit source code, its VRAM detection system in [llmfit-core/src/hardware.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) follows a consistent pipeline: detect all GPUs from every available source, merge overlapping entries, and sort by available memory. This architecture handles the reality that different backends expose different information—some report VRAM directly, others only device names.
NVIDIA CUDA VRAM Detection
The CUDA backend uses nvidia-smi as its primary detection path, with multiple fallback layers.
Primary Query: Extended NVIDIA-SMI
The detect_nvidia_gpus function first attempts an extended query that captures unified-memory GPUs:
nvidia-smi --query-gpu=addressing_mode,memory.total,name --format=csv,noheader,nounits
The parser parse_nvidia_smi_extended processes this output. The addressing_mode field reveals ATS (Address Translation Services) for unified-memory devices like NVIDIA Grace or DGX-Spark SoCs.
Fallback Queries
If the extended query fails, llmfit falls back to:
- Standard
nvidia-smi --query-gpu=memory.total,name - Linux sysfs:
/sys/class/drm/card*/device/mem_info_vram_total
Unified Memory Handling
When addressing_mode is "ATS" or VRAM reads as zero, the code substitutes total system RAM via read_proc_meminfo_total_gb. The GpuInfo struct marks unified_memory = true for these devices.
// Conceptual flow in hardware.rs
if addressing_mode == "ATS" || vram_mib == 0 {
vram_gb = total_system_ram_gb;
unified_memory = true;
}
AMD ROCm VRAM Detection
The ROCm backend pairs memory queries with product name lookups, then filters integrated GPUs.
Dual-Command Approach
The detect_amd_gpu_rocm_info function executes:
# VRAM totals in bytes
rocm-smi --showmeminfo vram
# Product names
rocm-smi --showproductname
The parsers parse_rocm_vram_bytes and parse_rocm_product_names handle both block and tabular output formats. The combiner parse_rocm_smi_output aligns these vectors by GPU index, groups identical models via is_same_gpu_name, and records both count and max_vram per model.
iGPU Filtering
The ROCm path includes a specific heuristic: if a device reports ≤2 GiB VRAM while a discrete GPU is present, the entry is discarded. This avoids reporting AMD APUs alongside dedicated cards when only the discrete VRAM is relevant for model sizing.
Name-Based Estimation Fallback
When rocm-smi is unavailable or reports invalid values, llmfit calls estimate_vram_from_name to derive VRAM from model strings like "Radeon RX 7900 XTX".
Vulkan VRAM Detection
Vulkan detection is unique: it discovers device names but never VRAM values directly.
Name-Only Query
The detect_vulkan_gpu_info function calls:
vulkaninfo --summary
The parse_vulkan_gpus parser extracts deviceName= entries and creates GpuInfo records with vram_gb: None.
Merge-Time Resolution
VRAM for Vulkan-detected GPUs is populated during the merging phase. The merge_gpu_sources function copies VRAM values from other detection methods (CUDA, ROCm, or WMI) when device names match. The is_same_gpu_name normalizer handles vendor-specific naming differences—for example, matching "AMD Radeon RX 7600" (ROCm) with "AMD Radeon RX 7600 XT (RADV NAVI33)" (Vulkan).
Intel SYCL VRAM Detection
SYCL detection is indirect because oneAPI provides no native VRAM query utility.
Platform-Specific Paths
On Linux, SYCL-labeled GPUs are discovered through /sys/class/drm sysfs entries. On Windows, WMI (Win32_VideoController) provides the detection path. Both leverage the same VRAM extraction code used for CUDA and ROCm fallbacks:
- Linux:
mem_info_vram_totalsysfs reads - Windows:
AdapterRAMproperty from WMI
Unified Memory for Intel
Intel integrated and some discrete GPUs share system memory. The same unified-memory logic applies: when VRAM reads as zero or the device class indicates integrated graphics, estimate_vram_from_name or total system RAM substitutes the value.
Cross-Backend Merge and Deduplication
The detect_all_gpus orchestration function implements a seven-stage pipeline:
let gpus = Self::detect_all_gpus(total_ram_gb, &cpu_name);
// 1. NVIDIA → nvidia-smi → sysfs fallback
// 2. AMD → rocm-smi → sysfs fallback
// 3. Windows WMI (NVIDIA, AMD, Intel, SYCL coverage)
// 4. Intel macOS Metal, Apple Silicon
// 5. Vulkan → vulkaninfo (name only, merged later)
// 6. Ascend → npu-smi
// 7. Merge → deduplicate → sort by VRAM descending
Deduplication Logic
The merge_gpu_sources function:
- Groups entries where
is_same_gpu_namereturns true - Preserves the maximum VRAM when sources disagree
- Carries
unified_memoryflags forward - Sums per-card VRAM into
total_gpu_vram_gbfor multi-GPU configurations
Windows WMI Universal Fallback
The detect_gpu_windows_info function provides cross-vendor coverage when vendor tools are missing:
Get-CimInstance Win32_VideoController |
Select-Object Name,AdapterRAM |
ForEach-Object { "$($_.Name)|$($_.AdapterRAM)" }
This powers SYCL detection on Windows and serves as a generic fallback for all vendors.
VRAM Estimation from Model Names
The estimate_vram_from_name function provides final fallback coverage. It pattern-matches GPU model strings against known configurations—extracting numeric series identifiers and mapping to typical VRAM capacities. This handles:
- New GPUs without driver support
- Containers missing
nvidia-smiorrocm-smi - Virtualized environments with passthrough devices
Summary
- CUDA:
nvidia-smiwith ATS detection for unified memory, sysfs fallback, system RAM substitution for Grace/DGX-Spark - ROCm: Dual
rocm-smicommands with iGPU filtering and name-based estimation - Vulkan: Name-only discovery, VRAM resolved via cross-source merging
- SYCL: Indirect detection through sysfs/WMI, VRAM from shared platform paths
- Unified architecture: Merge pipeline with
merge_gpu_sourcesandis_same_gpu_namededuplication produces canonical GPU inventory sorted by available VRAM
Frequently Asked Questions
Why does llmfit use multiple detection methods for the same GPU?
Different backends expose different information. nvidia-smi provides precise VRAM but requires the NVIDIA driver. Vulkan is universally available but reports only device names. By querying all sources and merging results, llmfit maximizes hardware coverage while maintaining accuracy when precise data is available. The merge logic in hardware.rs prioritizes actual VRAM readings over estimates.
How does llmfit handle unified memory GPUs like Apple Silicon or NVIDIA Grace?
These devices set unified_memory = true in their GpuInfo struct. When the detection path identifies ATS addressing mode (NVIDIA) or reads zero dedicated VRAM (Apple/Intel), the code substitutes total system RAM via read_proc_meminfo_total_gb. This correctly reflects that the GPU shares the main memory pool.
What happens when no vendor tools are installed?
llmfit cascades through fallback layers: vendor CLI → OS sysfs/WMI → Vulkan names → model-name estimation. On a Linux system without nvidia-smi or rocm-smi, the code reads /sys/class/drm/card*/device/ and parses PCI IDs. If all else fails, estimate_vram_from_name provides a best-guess based on device name patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →