How to Fix "Torch not compiled with CUDA" Error on Windows with insanely-fast-whisper

The "Torch not compiled with CUDA" error occurs when PyTorch is installed as a CPU-only wheel instead of a CUDA-enabled build, causing the insanely-fast-whisper CLI to fail when it attempts to move the Whisper model to GPU memory.

When running insanely-fast-whisper on Windows with an NVIDIA GPU, you may encounter an AssertionError: Torch not compiled with CUDA enabled. This happens because the default pip installation often pulls the CPU-only version of PyTorch, while the CLI explicitly requests a CUDA device to accelerate transcription. The error is not a bug in the tool itself, but rather a dependency mismatch that prevents GPU utilization.

Why the Error Occurs in the Source Code

The insanely-fast-whisper CLI constructs a Hugging Face pipeline that defaults to GPU execution unless explicitly told otherwise. In src/insanely_fast_whisper/cli.py, the code imports PyTorch and builds a device string that resolves to cuda:0 by default:

import torch

# ...

device="mps" if args.device_id == "mps" else f"cuda:{args.device_id}"

When the pipeline is created with torch_dtype=torch.float16 and the device parameter points to a CUDA index, PyTorch attempts to initialize CUDA tensors. If the installed torch binary lacks CUDA kernels, this triggers the assertion error before any transcription begins.

This device logic also appears in src/insanely_fast_whisper/utils/diarization_pipeline.py, where the diarization module attempts to create a torch.device object using the same conditional logic. Additionally, src/insanely_fast_whisper/utils/diarize.py relies on torch.from_numpy() operations that expect a CUDA-capable runtime when GPU mode is active.

Step-by-Step Resolution

Follow these steps to replace the CPU-only PyTorch installation with a CUDA-enabled build that matches your NVIDIA driver and CUDA toolkit version.

  1. Check your CUDA version to ensure you download the correct wheel index:

    nvcc --version

    Note the version (e.g., 12.1, 11.8) from the output.

  2. Uninstall the existing CPU-only PyTorch from your virtual environment:

    pip uninstall -y torch torchvision torchaudio
  3. Install the CUDA-compatible PyTorch build using the official PyTorch index URL. Replace cu121 with your specific CUDA version (e.g., cu118 for CUDA 11.8):

    python -m pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
  4. Verify the installation by checking CUDA availability in Python:

    import torch
    print(torch.cuda.is_available())  # Should output: True
    
    print(torch.__version__)          # Should include +cu121 (or your version)
    
  5. Run the transcription with default GPU settings:

    insanely-fast-whisper --file-name audio.wav

Forcing CPU Mode as a Workaround

If you prefer to run transcription without GPU acceleration or lack NVIDIA hardware, bypass the CUDA requirement entirely by specifying the CPU device:

insanely-fast-whisper --file-name audio.wav --device-id cpu

This prevents the CLI from attempting to construct the cuda:0 device string and avoids loading CUDA-specific kernels, allowing the CPU-only PyTorch build to function correctly. On macOS, use --device-id mps to leverage Apple Silicon instead.

Summary

  • The error stems from a CPU-only PyTorch installation attempting to execute CUDA-specific code paths in cli.py and the diarization utilities.
  • Resolution requires reinstalling PyTorch with the matching CUDA wheel from https://download.pytorch.org/whl/cu{version}.
  • Verify CUDA availability with torch.cuda.is_available() before running the CLI.
  • CPU fallback is available via --device-id cpu if GPU acceleration is not required.

Frequently Asked Questions

How do I know which CUDA version to use for PyTorch?

Check your installed NVIDIA driver and CUDA toolkit by running nvcc --version in PowerShell or Command Prompt. Match the PyTorch index URL to this version—use cu121 for CUDA 12.1, cu118 for CUDA 11.8, etc. Installing a mismatched version will result in the same CUDA error or driver incompatibility issues.

Why does pip install the CPU-only version by default on Windows?

PyTorch's default package index serves CPU-only wheels unless explicitly specified because CUDA builds require specific NVIDIA runtime libraries that may not exist on all Windows systems. The project README confirms this behavior and recommends the manual --index-url installation method to force the CUDA-enabled binary.

Can I use insanely-fast-whisper without an NVIDIA GPU?

Yes. Pass --device-id cpu to force CPU execution, or use --device-id mps on Apple Silicon Macs. While this eliminates the CUDA error, transcription will run significantly slower than the GPU-accelerated path optimized for CUDA in the insanely-fast-whisper pipeline.

Does this error affect the speaker diarization feature specifically?

The diarization modules in utils/diarize.py and utils/diarization_pipeline.py also import PyTorch and construct device tensors. If the base PyTorch installation lacks CUDA support, diarization will fail with the same assertion error when processing audio segments, making the CUDA fix necessary for full GPU-enabled functionality.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →