How to Troubleshoot GPU Detection Issues with llmfit doctor

Run llmfit doctor --verbose and verify that platform-specific utilities (nvidia-smi, rocm-smi, or system_profiler) exist in your $PATH and return valid output—missing drivers, permission errors, or environment variables like LLMFIT_SKIP_GPU=1 are the most common causes of detection failures.

The llmfit doctor command in AlexsJones/llmfit validates your runtime environment and determines whether GPU acceleration is available. When GPU detection fails, the tool silently falls back to CPU-only mode, which can severely impact model training performance. Understanding how the detection mechanism works—and where it can break—lets you diagnose problems quickly.

How llmfit doctor Detects GPUs

The diagnostic flow originates in llmfit-core/src/doctor.rs, which orchestrates system checks by calling llmfit-core/src/hardware.rs. The hardware module executes platform-specific external commands to query GPU status:

Platform Utility Information Gathered
NVIDIA nvidia-smi Driver version, GPU model, VRAM capacity, temperature, and health status
AMD rocm-smi ROCm driver version, GPU metadata, and memory statistics
Apple Silicon system_profiler SPDisplaysDataType Unified memory allocation (treated as VRAM on M-series chips)

If any utility is absent from $PATH or exits with a non-zero status, hardware.rs logs a warning and signals to doctor.rs that GPU acceleration is unavailable. This design prioritizes portability over deep library dependencies but creates several external failure points.

Common Causes of GPU Detection Failures

Missing System Utilities

The most frequent issue occurs when GPU management tools are not installed or not discoverable. Unlike Python-based ML frameworks that might use pycuda or direct driver APIs, llmfit relies on shell-out commands.

Check availability with:

which nvidia-smi   # NVIDIA systems

which rocm-smi     # AMD systems

If either returns nothing, install the appropriate driver package:

  • NVIDIA Ubuntu: sudo apt install nvidia-driver-<version>
  • AMD Ubuntu: sudo apt install rocm-dev

Driver Mismatch or Failure

Even when utilities exist, they may report driver errors. Run manually to verify:

nvidia-smi
rocm-smi

Healthy output shows a table with GPU names, temperatures, and memory usage. Common error responses include "NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver" or ROCm initialization failures—these indicate kernel module problems that llmfit doctor will surface as detection failures.

Permission and Container Restrictions

In Docker containers or restricted shell environments, the detection utilities may exist but fail to execute due to:

  • Missing --gpus all flag in docker run
  • Seccomp profiles blocking driver IOCTL calls
  • Read-only filesystems preventing temporary file creation

Verify permissions by running the utilities as the same user that will execute llmfit.

Environment Variable Overrides

The detection logic respects LLMFIT_SKIP_GPU=1, which forces CPU-only mode regardless of hardware availability. Check your environment:

echo $LLMFIT_SKIP_GPU
unset LLMFIT_SKIP_GPU  # Remove if set

Also verify that PATH does not contain stale or incompatible utility versions earlier in the search order.

Step-by-Step Troubleshooting Workflow

1. Capture Verbose Diagnostic Output

Start with maximum visibility into the detection process:

llmfit doctor --verbose

This flag prints the exact shell commands being executed, their exit codes, and captured stderr—critical for identifying which stage fails.

2. Validate Platform Utilities Manually

Execute the same commands llmfit uses:


# NVIDIA verification

nvidia-smi --query-gpu=name,driver_version,memory.total --format=csv

# AMD verification

rocm-smi --showproductname --showdriverversion --showmeminfo vram

# Apple Silicon verification

system_profiler SPDisplaysDataType | grep -A5 "Chipset Model"

Compare output against expected formats. Malformed JSON or CSV from these tools will cause hardware.rs parsing to fail.

3. Inspect the Diagnostics Artifact

llmfit-core/src/doctor.rs writes structured results to diagnostics.json in the working directory (or the path specified by --temp-dir). Examine the GPU section:

cat diagnostics.json | jq '.gpu_detection'

Look for fields like "status": "failed" and "reason" strings that explain the fallback decision.

4. Clear Cached Results

If you've recently changed drivers or hardware, stale cache entries may interfere:

llmfit doctor --clear-cache --verbose

5. Verify Against Broader CUDA/ROCm Environment

Isolate whether the issue is llmfit-specific or systemic:


# PyTorch CUDA check

python -c "import torch; print(f'CUDA available: {torch.cuda.is_available()}'); print(f'Device count: {torch.cuda.device_count()}')"

# Direct ROCm check (HIP)

python -c "import torch; print(f'HIP available: {torch.backends.hip.is_available() if hasattr(torch.backends, \"hip\") else \"N/A\"}')"

If these also fail, the problem lies below llmfit in driver or kernel configuration.

6. Check for GPU Power Management States

Some servers disable GPUs via ACPI or vendor-specific tools. Verify the GPU is not in a sleep state:


# NVIDIA persistence mode

nvidia-smi -q -d PERFORMANCE | grep "Persistence-M"

# Force persistence mode on if disabled

sudo nvidia-smi -pm 1

Review BIOS/UEFI settings for "Above 4G Decoding" and "Resizable BAR" options that affect GPU visibility.

7. Update and Rebuild

Detection logic evolves with new hardware support. Ensure you're running current source:

cd llmfit
git pull origin main
cargo update && cargo build --release

Key Source Files for Deep Debugging

File Purpose Direct Link
llmfit-core/src/hardware.rs Low-level hardware probing, GPU detection command execution View source
llmfit-core/src/doctor.rs Diagnostic orchestration, report formatting, cache management View source
llmfit-tui/src/main.rs CLI argument parsing, subcommand dispatch View source

Reading hardware.rs reveals the exact command-line arguments passed to nvidia-smi and rocm-smi, useful for reproducing failures outside llmfit.

Summary

  • llmfit doctor --verbose is your first diagnostic tool—use it to see exactly which detection command fails
  • GPU detection depends entirely on external utilities (nvidia-smi, rocm-smi, system_profiler) being present and functional in $PATH
  • Check diagnostics.json for structured failure reasons after any doctor run
  • Environment variable LLMFIT_SKIP_GPU=1 silently disables GPU detection—verify it is unset
  • Manual execution of platform utilities isolates driver issues from llmfit-specific bugs
  • Update to latest source with cargo update && cargo build --release before reporting issues

Frequently Asked Questions

Why does llmfit doctor report "GPU not detected" when nvidia-smi works fine?

The parsing logic in llmfit-core/src/hardware.rs expects specific output fields. If nvidia-smi returns unexpected formats (older driver versions, corrupt GPU EEPROM, or non-standard locale settings), the regex or CSV parser may fail. Run llmfit doctor --verbose to see the raw capture and compare against expected patterns in the source code.

Can I force llmfit to use GPU even when doctor detects none?

No—llmfit doctor's detection result gates GPU execution paths throughout the tool. Attempting to override this risks runtime failures during model initialization. Instead, resolve the underlying detection issue by ensuring utilities execute successfully and return valid data.

Where are llmfit doctor logs stored?

By default, diagnostics.json is written to the current working directory. Use --temp-dir /path/to/dir to redirect output. The file contains structured JSON with sections for gpu_detection, cpu_info, memory_info, and runtime_environment—examine gpu_detection.status and gpu_detection.reason for specific failure explanations.

Does llmfit support GPU pass-through in Docker containers?

Yes, provided the container is launched with GPU access flags (--gpus all for NVIDIA, or --device /dev/kfd with ROCm bind-mounts for AMD). Verify by running nvidia-smi or rocm-smi inside the container before invoking llmfit doctor. Restricted containers without device access will report CPU-only mode.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →