# How to Troubleshoot GPU Detection Issues with llmfit doctor

> Troubleshoot GPU detection issues with llmfit doctor. Run llmfit doctor --verbose to check drivers, permissions, and environment variables, common causes of GPU failures.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: how-to-guide
- Published: 2026-08-20

---

**Run `llmfit doctor --verbose` and verify that platform-specific utilities (`nvidia-smi`, `rocm-smi`, or `system_profiler`) exist in your `$PATH` and return valid output—missing drivers, permission errors, or environment variables like `LLMFIT_SKIP_GPU=1` are the most common causes of detection failures.**

The `llmfit doctor` command in [AlexsJones/llmfit](https://github.com/AlexsJones/llmfit) validates your runtime environment and determines whether GPU acceleration is available. When GPU detection fails, the tool silently falls back to CPU-only mode, which can severely impact model training performance. Understanding how the detection mechanism works—and where it can break—lets you diagnose problems quickly.

## How llmfit doctor Detects GPUs

The diagnostic flow originates in **[`llmfit-core/src/doctor.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/doctor.rs)**, which orchestrates system checks by calling **[`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs)**. The hardware module executes platform-specific external commands to query GPU status:

| Platform | Utility | Information Gathered |
|----------|---------|----------------------|
| NVIDIA | `nvidia-smi` | Driver version, GPU model, VRAM capacity, temperature, and health status |
| AMD | `rocm-smi` | ROCm driver version, GPU metadata, and memory statistics |
| Apple Silicon | `system_profiler SPDisplaysDataType` | Unified memory allocation (treated as VRAM on M-series chips) |

If any utility is absent from `$PATH` or exits with a non-zero status, [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) logs a warning and signals to [`doctor.rs`](https://github.com/AlexsJones/llmfit/blob/main/doctor.rs) that GPU acceleration is unavailable. This design prioritizes portability over deep library dependencies but creates several external failure points.

## Common Causes of GPU Detection Failures

### Missing System Utilities

The most frequent issue occurs when GPU management tools are not installed or not discoverable. Unlike Python-based ML frameworks that might use `pycuda` or direct driver APIs, `llmfit` relies on shell-out commands.

Check availability with:

```bash
which nvidia-smi   # NVIDIA systems

which rocm-smi     # AMD systems

```

If either returns nothing, install the appropriate driver package:
- **NVIDIA Ubuntu**: `sudo apt install nvidia-driver-<version>`
- **AMD Ubuntu**: `sudo apt install rocm-dev`

### Driver Mismatch or Failure

Even when utilities exist, they may report driver errors. Run manually to verify:

```bash
nvidia-smi
rocm-smi

```

Healthy output shows a table with GPU names, temperatures, and memory usage. Common error responses include "NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver" or ROCm initialization failures—these indicate kernel module problems that `llmfit doctor` will surface as detection failures.

### Permission and Container Restrictions

In Docker containers or restricted shell environments, the detection utilities may exist but fail to execute due to:
- Missing `--gpus all` flag in `docker run`
- Seccomp profiles blocking driver IOCTL calls
- Read-only filesystems preventing temporary file creation

Verify permissions by running the utilities as the same user that will execute `llmfit`.

### Environment Variable Overrides

The detection logic respects `LLMFIT_SKIP_GPU=1`, which forces CPU-only mode regardless of hardware availability. Check your environment:

```bash
echo $LLMFIT_SKIP_GPU
unset LLMFIT_SKIP_GPU  # Remove if set

```

Also verify that `PATH` does not contain stale or incompatible utility versions earlier in the search order.

## Step-by-Step Troubleshooting Workflow

### 1. Capture Verbose Diagnostic Output

Start with maximum visibility into the detection process:

```bash
llmfit doctor --verbose

```

This flag prints the exact shell commands being executed, their exit codes, and captured stderr—critical for identifying which stage fails.

### 2. Validate Platform Utilities Manually

Execute the same commands `llmfit` uses:

```bash

# NVIDIA verification

nvidia-smi --query-gpu=name,driver_version,memory.total --format=csv

# AMD verification

rocm-smi --showproductname --showdriverversion --showmeminfo vram

# Apple Silicon verification

system_profiler SPDisplaysDataType | grep -A5 "Chipset Model"

```

Compare output against expected formats. Malformed JSON or CSV from these tools will cause [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) parsing to fail.

### 3. Inspect the Diagnostics Artifact

[`llmfit-core/src/doctor.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/doctor.rs) writes structured results to [`diagnostics.json`](https://github.com/AlexsJones/llmfit/blob/main/diagnostics.json) in the working directory (or the path specified by `--temp-dir`). Examine the GPU section:

```bash
cat diagnostics.json | jq '.gpu_detection'

```

Look for fields like `"status": "failed"` and `"reason"` strings that explain the fallback decision.

### 4. Clear Cached Results

If you've recently changed drivers or hardware, stale cache entries may interfere:

```bash
llmfit doctor --clear-cache --verbose

```

### 5. Verify Against Broader CUDA/ROCm Environment

Isolate whether the issue is `llmfit`-specific or systemic:

```bash

# PyTorch CUDA check

python -c "import torch; print(f'CUDA available: {torch.cuda.is_available()}'); print(f'Device count: {torch.cuda.device_count()}')"

# Direct ROCm check (HIP)

python -c "import torch; print(f'HIP available: {torch.backends.hip.is_available() if hasattr(torch.backends, \"hip\") else \"N/A\"}')"

```

If these also fail, the problem lies below `llmfit` in driver or kernel configuration.

### 6. Check for GPU Power Management States

Some servers disable GPUs via ACPI or vendor-specific tools. Verify the GPU is not in a sleep state:

```bash

# NVIDIA persistence mode

nvidia-smi -q -d PERFORMANCE | grep "Persistence-M"

# Force persistence mode on if disabled

sudo nvidia-smi -pm 1

```

Review BIOS/UEFI settings for "Above 4G Decoding" and "Resizable BAR" options that affect GPU visibility.

### 7. Update and Rebuild

Detection logic evolves with new hardware support. Ensure you're running current source:

```bash
cd llmfit
git pull origin main
cargo update && cargo build --release

```

## Key Source Files for Deep Debugging

| File | Purpose | Direct Link |
|------|---------|-------------|
| [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) | Low-level hardware probing, GPU detection command execution | [View source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) |
| [`llmfit-core/src/doctor.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/doctor.rs) | Diagnostic orchestration, report formatting, cache management | [View source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/doctor.rs) |
| [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs) | CLI argument parsing, subcommand dispatch | [View source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs) |

Reading [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) reveals the exact command-line arguments passed to `nvidia-smi` and `rocm-smi`, useful for reproducing failures outside `llmfit`.

## Summary

- **`llmfit doctor --verbose`** is your first diagnostic tool—use it to see exactly which detection command fails
- GPU detection depends entirely on external utilities (`nvidia-smi`, `rocm-smi`, `system_profiler`) being present and functional in `$PATH`
- Check **[`diagnostics.json`](https://github.com/AlexsJones/llmfit/blob/main/diagnostics.json)** for structured failure reasons after any doctor run
- Environment variable **`LLMFIT_SKIP_GPU=1`** silently disables GPU detection—verify it is unset
- Manual execution of platform utilities isolates driver issues from `llmfit`-specific bugs
- Update to latest source with `cargo update && cargo build --release` before reporting issues

## Frequently Asked Questions

### Why does llmfit doctor report "GPU not detected" when nvidia-smi works fine?

The parsing logic in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) expects specific output fields. If `nvidia-smi` returns unexpected formats (older driver versions, corrupt GPU EEPROM, or non-standard locale settings), the regex or CSV parser may fail. Run `llmfit doctor --verbose` to see the raw capture and compare against expected patterns in the source code.

### Can I force llmfit to use GPU even when doctor detects none?

No—`llmfit doctor`'s detection result gates GPU execution paths throughout the tool. Attempting to override this risks runtime failures during model initialization. Instead, resolve the underlying detection issue by ensuring utilities execute successfully and return valid data.

### Where are llmfit doctor logs stored?

By default, [`diagnostics.json`](https://github.com/AlexsJones/llmfit/blob/main/diagnostics.json) is written to the current working directory. Use `--temp-dir /path/to/dir` to redirect output. The file contains structured JSON with sections for `gpu_detection`, `cpu_info`, `memory_info`, and `runtime_environment`—examine `gpu_detection.status` and `gpu_detection.reason` for specific failure explanations.

### Does llmfit support GPU pass-through in Docker containers?

Yes, provided the container is launched with GPU access flags (`--gpus all` for NVIDIA, or `--device /dev/kfd` with ROCm bind-mounts for AMD). Verify by running `nvidia-smi` or `rocm-smi` inside the container before invoking `llmfit doctor`. Restricted containers without device access will report CPU-only mode.