How to Troubleshoot GPU Detection Issues with llmfit doctor
Run llmfit doctor --verbose and verify that platform-specific utilities (nvidia-smi, rocm-smi, or system_profiler) exist in your $PATH and return valid output—missing drivers, permission errors, or environment variables like LLMFIT_SKIP_GPU=1 are the most common causes of detection failures.
The llmfit doctor command in AlexsJones/llmfit validates your runtime environment and determines whether GPU acceleration is available. When GPU detection fails, the tool silently falls back to CPU-only mode, which can severely impact model training performance. Understanding how the detection mechanism works—and where it can break—lets you diagnose problems quickly.
How llmfit doctor Detects GPUs
The diagnostic flow originates in llmfit-core/src/doctor.rs, which orchestrates system checks by calling llmfit-core/src/hardware.rs. The hardware module executes platform-specific external commands to query GPU status:
| Platform | Utility | Information Gathered |
|---|---|---|
| NVIDIA | nvidia-smi |
Driver version, GPU model, VRAM capacity, temperature, and health status |
| AMD | rocm-smi |
ROCm driver version, GPU metadata, and memory statistics |
| Apple Silicon | system_profiler SPDisplaysDataType |
Unified memory allocation (treated as VRAM on M-series chips) |
If any utility is absent from $PATH or exits with a non-zero status, hardware.rs logs a warning and signals to doctor.rs that GPU acceleration is unavailable. This design prioritizes portability over deep library dependencies but creates several external failure points.
Common Causes of GPU Detection Failures
Missing System Utilities
The most frequent issue occurs when GPU management tools are not installed or not discoverable. Unlike Python-based ML frameworks that might use pycuda or direct driver APIs, llmfit relies on shell-out commands.
Check availability with:
which nvidia-smi # NVIDIA systems
which rocm-smi # AMD systems
If either returns nothing, install the appropriate driver package:
- NVIDIA Ubuntu:
sudo apt install nvidia-driver-<version> - AMD Ubuntu:
sudo apt install rocm-dev
Driver Mismatch or Failure
Even when utilities exist, they may report driver errors. Run manually to verify:
nvidia-smi
rocm-smi
Healthy output shows a table with GPU names, temperatures, and memory usage. Common error responses include "NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver" or ROCm initialization failures—these indicate kernel module problems that llmfit doctor will surface as detection failures.
Permission and Container Restrictions
In Docker containers or restricted shell environments, the detection utilities may exist but fail to execute due to:
- Missing
--gpus allflag indocker run - Seccomp profiles blocking driver IOCTL calls
- Read-only filesystems preventing temporary file creation
Verify permissions by running the utilities as the same user that will execute llmfit.
Environment Variable Overrides
The detection logic respects LLMFIT_SKIP_GPU=1, which forces CPU-only mode regardless of hardware availability. Check your environment:
echo $LLMFIT_SKIP_GPU
unset LLMFIT_SKIP_GPU # Remove if set
Also verify that PATH does not contain stale or incompatible utility versions earlier in the search order.
Step-by-Step Troubleshooting Workflow
1. Capture Verbose Diagnostic Output
Start with maximum visibility into the detection process:
llmfit doctor --verbose
This flag prints the exact shell commands being executed, their exit codes, and captured stderr—critical for identifying which stage fails.
2. Validate Platform Utilities Manually
Execute the same commands llmfit uses:
# NVIDIA verification
nvidia-smi --query-gpu=name,driver_version,memory.total --format=csv
# AMD verification
rocm-smi --showproductname --showdriverversion --showmeminfo vram
# Apple Silicon verification
system_profiler SPDisplaysDataType | grep -A5 "Chipset Model"
Compare output against expected formats. Malformed JSON or CSV from these tools will cause hardware.rs parsing to fail.
3. Inspect the Diagnostics Artifact
llmfit-core/src/doctor.rs writes structured results to diagnostics.json in the working directory (or the path specified by --temp-dir). Examine the GPU section:
cat diagnostics.json | jq '.gpu_detection'
Look for fields like "status": "failed" and "reason" strings that explain the fallback decision.
4. Clear Cached Results
If you've recently changed drivers or hardware, stale cache entries may interfere:
llmfit doctor --clear-cache --verbose
5. Verify Against Broader CUDA/ROCm Environment
Isolate whether the issue is llmfit-specific or systemic:
# PyTorch CUDA check
python -c "import torch; print(f'CUDA available: {torch.cuda.is_available()}'); print(f'Device count: {torch.cuda.device_count()}')"
# Direct ROCm check (HIP)
python -c "import torch; print(f'HIP available: {torch.backends.hip.is_available() if hasattr(torch.backends, \"hip\") else \"N/A\"}')"
If these also fail, the problem lies below llmfit in driver or kernel configuration.
6. Check for GPU Power Management States
Some servers disable GPUs via ACPI or vendor-specific tools. Verify the GPU is not in a sleep state:
# NVIDIA persistence mode
nvidia-smi -q -d PERFORMANCE | grep "Persistence-M"
# Force persistence mode on if disabled
sudo nvidia-smi -pm 1
Review BIOS/UEFI settings for "Above 4G Decoding" and "Resizable BAR" options that affect GPU visibility.
7. Update and Rebuild
Detection logic evolves with new hardware support. Ensure you're running current source:
cd llmfit
git pull origin main
cargo update && cargo build --release
Key Source Files for Deep Debugging
| File | Purpose | Direct Link |
|---|---|---|
llmfit-core/src/hardware.rs |
Low-level hardware probing, GPU detection command execution | View source |
llmfit-core/src/doctor.rs |
Diagnostic orchestration, report formatting, cache management | View source |
llmfit-tui/src/main.rs |
CLI argument parsing, subcommand dispatch | View source |
Reading hardware.rs reveals the exact command-line arguments passed to nvidia-smi and rocm-smi, useful for reproducing failures outside llmfit.
Summary
llmfit doctor --verboseis your first diagnostic tool—use it to see exactly which detection command fails- GPU detection depends entirely on external utilities (
nvidia-smi,rocm-smi,system_profiler) being present and functional in$PATH - Check
diagnostics.jsonfor structured failure reasons after any doctor run - Environment variable
LLMFIT_SKIP_GPU=1silently disables GPU detection—verify it is unset - Manual execution of platform utilities isolates driver issues from
llmfit-specific bugs - Update to latest source with
cargo update && cargo build --releasebefore reporting issues
Frequently Asked Questions
Why does llmfit doctor report "GPU not detected" when nvidia-smi works fine?
The parsing logic in llmfit-core/src/hardware.rs expects specific output fields. If nvidia-smi returns unexpected formats (older driver versions, corrupt GPU EEPROM, or non-standard locale settings), the regex or CSV parser may fail. Run llmfit doctor --verbose to see the raw capture and compare against expected patterns in the source code.
Can I force llmfit to use GPU even when doctor detects none?
No—llmfit doctor's detection result gates GPU execution paths throughout the tool. Attempting to override this risks runtime failures during model initialization. Instead, resolve the underlying detection issue by ensuring utilities execute successfully and return valid data.
Where are llmfit doctor logs stored?
By default, diagnostics.json is written to the current working directory. Use --temp-dir /path/to/dir to redirect output. The file contains structured JSON with sections for gpu_detection, cpu_info, memory_info, and runtime_environment—examine gpu_detection.status and gpu_detection.reason for specific failure explanations.
Does llmfit support GPU pass-through in Docker containers?
Yes, provided the container is launched with GPU access flags (--gpus all for NVIDIA, or --device /dev/kfd with ROCm bind-mounts for AMD). Verify by running nvidia-smi or rocm-smi inside the container before invoking llmfit doctor. Restricted containers without device access will report CPU-only mode.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →