# How llmfit Detects GPU VRAM Using nvidia-smi, rocm-smi, system_profiler, and objc2-metal

> Discover how llmfit detects GPU VRAM using nvidia-smi, rocm-smi, system_profiler, and objc2-metal. Learn about cross-platform GPU memory detection and analysis.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-11

---

**llmfit determines available GPU VRAM by executing platform-specific system commands—`nvidia-smi` for NVIDIA GPUs, `rocm-smi` for AMD GPUs on Linux, and `system_profiler` or the `objc2-metal` crate for macOS—then parsing the output to populate a `GpuInfo` struct used throughout the application.**

The `llmfit` repository (AlexsJones/llmfit) implements robust cross-platform GPU memory detection in **[`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs)**. The detection engine shells out to native system utilities, parses CSV or structured text output, and aggregates results into a `SystemSpecs` struct that downstream components use for model-fit calculations and hardware validation.

## NVIDIA GPU VRAM Detection with nvidia-smi

On Linux and Windows systems with NVIDIA hardware, `llmfit` relies on the `nvidia-smi` utility to query dedicated video memory.

### Standard and Extended CSV Queries

The primary detection path executes `nvidia-smi` with specific query parameters to retrieve memory totals and device names:

```bash
nvidia-smi --query-gpu=memory.total,name --format=csv,noheader,nounits

```

For edge cases like DGX or Jetson devices where the standard query reports **0 GB**, the code falls back to an extended query that includes the `addressing_mode` field. This extended form helps identify legacy drivers or unusual memory configurations that cannot expose VRAM through standard APIs.

### Parsing Implementation

In [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs), the `SystemSpecs::parse_nvidia_smi_list` and `parse_nvidia_smi_extended` functions handle CSV parsing. These methods extract the `memory.total` value, convert it to gigabytes, and instantiate `GpuInfo` structs with the GPU name and VRAM capacity.

```rust
// From llmfit-core/src/hardware.rs
let output = Command::new("nvidia-smi")
    .args(&["--query-gpu=memory.total,name", "--format=csv,noheader,nounits"])
    .output()?;

let nvidia_gpus = SystemSpecs::parse_nvidia_smi_list(
    &String::from_utf8_lossy(&output.stdout)
);

```

### Linux Sysfs Fallback

When `nvidia-smi` is unavailable or returns zero values, `llmfit` invokes `detect_nvidia_gpu_sysfs_info` to read PCI sysfs entries directly from `/sys/class/pci_bus/`. This fallback traverses the PCI device tree to synthesize a VRAM estimate based on the GPU's BAR memory reporting, ensuring detection works on headless servers without full NVIDIA driver utilities installed.

## AMD GPU VRAM Detection with rocm-smi

For AMD GPUs on Linux, `llmfit` queries the ROCm management interface via `rocm-smi`. The detection logic shells out to:

```bash
rocm-smi -i

```

The parser scans the command output for the **"GPU Memory Total"** field. Because `rocm-smi` reports memory in MiB, [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) converts the value to GB before storing it in the `GpuInfo` struct.

```rust
// Internal implementation in llmfit-core/src/hardware.rs
let output = Command::new("rocm-smi").arg("-i").output()?;
// Parse "GPU Memory Total" field, convert MiB to GB

```

This approach supports modern AMD datacenter and consumer GPUs running the ROCm stack, providing accurate VRAM figures for model batch-size calculations.

## macOS GPU VRAM Detection

macOS detection uses two complementary strategies depending on the hardware generation and driver availability.

### system_profiler SPDisplaysDataType

On Intel-based Macs and older systems, `llmfit` executes:

```bash
system_profiler SPDisplaysDataType

```

The [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) parser scans the structured text output for the **"VRAM (Total)"** line. When found, the value is parsed and recorded as the discrete GPU's memory capacity.

### objc2-metal Working-Set Queries

For Metal-enabled GPUs or when `system_profiler` provides insufficient detail, `llmfit` uses the **`objc2-metal`** crate (declared in [`Cargo.toml`](https://github.com/AlexsJones/llmfit/blob/main/Cargo.toml)) to query the Metal API directly. This method retrieves the **"effective Metal working-set limit"**—the maximum buffer size the Metal driver will allocate—which serves as a reliable VRAM estimate for Intel integrated graphics and discrete AMD GPUs in Mac systems.

```rust
// Example: macOS detection path
let mac_gpu = SystemSpecs::detect_macos_gpu_vram()?;
// Uses objc2-metal to query Metal working-set limit when available

```

### Unified Memory Handling on Apple Silicon

Apple Silicon Macs (M1, M2, M3 series) use unified memory architecture where system RAM serves as both CPU and GPU memory. In this case, `llmfit` treats the reported system RAM as VRAM and sets the `GpuInfo.unified_memory` flag to `true`. The detection logic skips the `CpuOffload` path because no separate VRAM pool exists, preventing unnecessary memory-copy optimizations.

## Integration and Aggregation in SystemSpecs

All detection routines feed into **`SystemSpecs::detect()`**, the central aggregation method in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs). This function:

1. **Probes** the operating system to identify the platform (NVIDIA, AMD, or macOS).
2. **Executes** the appropriate platform-specific tool (`nvidia-smi`, `rocm-smi`, or `system_profiler`/`objc2-metal`).
3. **Parses** the raw output into `GpuInfo` structs containing `name`, `vram_gb`, and architecture flags.
4. **Falls back** to sysfs parsing or extended queries when primary tools report zero values.
5. **Aggregates** CPU, RAM, and GPU information into a single `SystemSpecs` instance consumed by the doctor diagnostics ([`llmfit-core/src/doctor.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/doctor.rs)) and model-fit analysis modules.

Unit tests in **`llmfit-core/tests/fixtures/hardware/nvidia/`** validate the CSV parsing logic using fixture files like `standard-identical.csv` and `extended-unified.csv`, ensuring robustness across different `nvidia-smi` output formats.

## Summary

- **[`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs)** contains the complete VRAM detection implementation for NVIDIA, AMD, and macOS platforms.
- **NVIDIA detection** uses `nvidia-smi` CSV output with `parse_nvidia_smi_list`, falling back to `detect_nvidia_gpu_sysfs_info` when necessary.
- **AMD detection** relies on `rocm-smi -i` output parsing to extract "GPU Memory Total" values.
- **macOS detection** combines `system_profiler SPDisplaysDataType` scanning with `objc2-metal` working-set queries for Metal-enabled devices.
- **Apple Silicon** sets the `unified_memory` flag and uses system RAM as VRAM, skipping discrete GPU offload logic.
- **`SystemSpecs::detect()`** aggregates all hardware discovery into a unified struct for downstream consumption.

## Frequently Asked Questions

### What happens if nvidia-smi is not installed on a Linux system?

If `nvidia-smi` is missing or returns zero values, `llmfit` invokes `detect_nvidia_gpu_sysfs_info` to read PCI sysfs entries in `/sys/class/pci_bus/`. This fallback parses BAR memory information directly from the kernel's PCI subsystem, allowing VRAM detection on headless servers or minimal container environments without the full NVIDIA driver toolkit.

### How does llmfit handle unified memory architectures like Apple Silicon?

On Apple Silicon Macs, `llmfit` detects the unified memory architecture and sets `GpuInfo.unified_memory` to `true`. The system treats the total system RAM as available GPU memory, and the application logic skips the `CpuOffload` optimization path since there is no separate VRAM pool requiring explicit memory management.

### Why does llmfit use objc2-metal instead of system_profiler on some macOS systems?

The `objc2-metal` crate queries the Metal API for the **effective working-set limit**, which provides a more accurate representation of available GPU memory than `system_profiler` on certain Intel Macs with discrete AMD GPUs. This method reflects the actual buffer allocation limits imposed by the Metal driver, rather than the theoretical hardware maximum.

### Can llmfit detect VRAM on headless servers without display drivers?

Yes. On Linux systems without display drivers but with NVIDIA hardware exposed via PCI, the `detect_nvidia_gpu_sysfs_info` fallback reads kernel sysfs entries to determine VRAM capacity. Similarly, AMD detection via `rocm-smi` does not require an active display connection, only the ROCm kernel modules. MacOS detection requires either `system_profiler` or Metal drivers, which are present in standard macOS installations.