# How to Fix "Torch not compiled with CUDA" Error on Windows with insanely-fast-whisper

> Fix the Torch not compiled with CUDA error for insanely-fast-whisper on Windows. Resolve PyTorch GPU memory issues and get your Whisper model running on the GPU quickly.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: how-to-guide
- Published: 2026-03-27

---

**The "Torch not compiled with CUDA" error occurs when PyTorch is installed as a CPU-only wheel instead of a CUDA-enabled build, causing the insanely-fast-whisper CLI to fail when it attempts to move the Whisper model to GPU memory.**

When running `insanely-fast-whisper` on Windows with an NVIDIA GPU, you may encounter an `AssertionError: Torch not compiled with CUDA enabled`. This happens because the default `pip` installation often pulls the CPU-only version of PyTorch, while the CLI explicitly requests a CUDA device to accelerate transcription. The error is not a bug in the tool itself, but rather a dependency mismatch that prevents GPU utilization.

## Why the Error Occurs in the Source Code

The `insanely-fast-whisper` CLI constructs a Hugging Face pipeline that defaults to GPU execution unless explicitly told otherwise. In [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py), the code imports PyTorch and builds a device string that resolves to `cuda:0` by default:

```python
import torch

# ...

device="mps" if args.device_id == "mps" else f"cuda:{args.device_id}"

```

When the pipeline is created with `torch_dtype=torch.float16` and the `device` parameter points to a CUDA index, PyTorch attempts to initialize CUDA tensors. If the installed `torch` binary lacks CUDA kernels, this triggers the assertion error before any transcription begins.

This device logic also appears in [`src/insanely_fast_whisper/utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarization_pipeline.py), where the diarization module attempts to create a `torch.device` object using the same conditional logic. Additionally, [`src/insanely_fast_whisper/utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarize.py) relies on `torch.from_numpy()` operations that expect a CUDA-capable runtime when GPU mode is active.

## Step-by-Step Resolution

Follow these steps to replace the CPU-only PyTorch installation with a CUDA-enabled build that matches your NVIDIA driver and CUDA toolkit version.

1. **Check your CUDA version** to ensure you download the correct wheel index:

   ```powershell
   nvcc --version
   ```

   Note the version (e.g., 12.1, 11.8) from the output.

2. **Uninstall the existing CPU-only PyTorch** from your virtual environment:

   ```powershell
   pip uninstall -y torch torchvision torchaudio
   ```

3. **Install the CUDA-compatible PyTorch build** using the official PyTorch index URL. Replace `cu121` with your specific CUDA version (e.g., `cu118` for CUDA 11.8):

   ```powershell
   python -m pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
   ```

4. **Verify the installation** by checking CUDA availability in Python:

   ```python
   import torch
   print(torch.cuda.is_available())  # Should output: True

   print(torch.__version__)          # Should include +cu121 (or your version)

   ```

5. **Run the transcription** with default GPU settings:

   ```powershell
   insanely-fast-whisper --file-name audio.wav
   ```

## Forcing CPU Mode as a Workaround

If you prefer to run transcription without GPU acceleration or lack NVIDIA hardware, bypass the CUDA requirement entirely by specifying the CPU device:

```powershell
insanely-fast-whisper --file-name audio.wav --device-id cpu

```

This prevents the CLI from attempting to construct the `cuda:0` device string and avoids loading CUDA-specific kernels, allowing the CPU-only PyTorch build to function correctly. On macOS, use `--device-id mps` to leverage Apple Silicon instead.

## Summary

- The error stems from a **CPU-only PyTorch installation** attempting to execute CUDA-specific code paths in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py) and the diarization utilities.
- **Resolution requires reinstalling PyTorch** with the matching CUDA wheel from `https://download.pytorch.org/whl/cu{version}`.
- **Verify CUDA availability** with `torch.cuda.is_available()` before running the CLI.
- **CPU fallback** is available via `--device-id cpu` if GPU acceleration is not required.

## Frequently Asked Questions

### How do I know which CUDA version to use for PyTorch?

Check your installed NVIDIA driver and CUDA toolkit by running `nvcc --version` in PowerShell or Command Prompt. Match the PyTorch index URL to this version—use `cu121` for CUDA 12.1, `cu118` for CUDA 11.8, etc. Installing a mismatched version will result in the same CUDA error or driver incompatibility issues.

### Why does pip install the CPU-only version by default on Windows?

PyTorch's default package index serves CPU-only wheels unless explicitly specified because CUDA builds require specific NVIDIA runtime libraries that may not exist on all Windows systems. The project README confirms this behavior and recommends the manual `--index-url` installation method to force the CUDA-enabled binary.

### Can I use insanely-fast-whisper without an NVIDIA GPU?

Yes. Pass `--device-id cpu` to force CPU execution, or use `--device-id mps` on Apple Silicon Macs. While this eliminates the CUDA error, transcription will run significantly slower than the GPU-accelerated path optimized for CUDA in the `insanely-fast-whisper` pipeline.

### Does this error affect the speaker diarization feature specifically?

The diarization modules in [`utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/diarize.py) and [`utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/diarization_pipeline.py) also import PyTorch and construct device tensors. If the base PyTorch installation lacks CUDA support, diarization will fail with the same assertion error when processing audio segments, making the CUDA fix necessary for full GPU-enabled functionality.