# Troubleshooting CUDA Driver Compatibility with PyTorch Backends for Cosmos 3

> Fix CUDA driver compatibility errors in Cosmos 3 with PyTorch. Ensure your CUDA wheel tag matches your NVIDIA driver for successful GPU acceleration. Avoid torch.cuda.is_available() failing.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: tutorial
- Published: 2026-07-03

---

**Cosmos 3 requires matching the CUDA wheel tag (`cu128` or `cu130`) to your NVIDIA driver version; otherwise, `torch.cuda.is_available()` returns `False` and GPU acceleration fails.**

Cosmos 3 relies on PyTorch, vLLM, and Transformers wheels that are pre-compiled for specific CUDA versions. When the CUDA version baked into the wheel does not match the NVIDIA driver on your host, the CUDA runtime fails to load, forcing all notebooks to fall back to CPU or abort entirely. This guide explains how to align your driver version with the correct PyTorch backend tag using the `uv` package manager as implemented in the NVIDIA/cosmos repository.

## Understanding the CUDA Driver and Wheel Tag Relationship

The `cuXXX` tag in a PyTorch wheel name indicates which CUDA runtime version the binary expects. In [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md), the documentation establishes a clear mapping between driver capabilities and backend tags.

### The Driver-to-Tag Mapping Table

| Driver CUDA | Backend Tag | Installation Group |
|-------------|-------------|-------------------|
| **13.x** | `cu130` | Default for most Cosmos 3 notebooks |
| **12.x** | `cu128` | Required for hosts with CUDA 12 drivers |

Selecting the wrong tag causes the wheel to expect a newer driver than is present, triggering runtime errors. The wheels are hosted on the PyPI index used by `uv`, and each group pins a specific `torch` wheel (e.g., `torch==X.Y.Z+cu130`).

## How Cosmos 3 Selects the PyTorch Wheel

Cosmos 3 uses **uv** to manage dependencies and selects the appropriate CUDA wheel through dependency groups or backend flags.

### Cosmos Framework Backend

For the **Cosmos Framework** backend, use the `uv sync` command with a specific group:

```bash

# For CUDA 13 drivers

uv sync --all-extras --group=cu130-train

# For CUDA 12 drivers

uv sync --all-extras --group=cu128-train

```

You can override the default group using the `COSMOS3_UV_GROUP` environment variable:

```bash
export COSMOS3_UV_GROUP=cu128-train   # Use when host has CUDA 12.x driver

uv sync --all-extras --group=$COSMOS3_UV_GROUP

```

### Diffusers, Transformers, and vLLM Backends

For the **Diffusers**, **Transformers**, **vLLM**, and **vLLM-Omni** backends, specify the backend via the `--torch-backend` flag:

```bash

# CUDA 13 driver

uv pip install --torch-backend=cu130

# CUDA 12 driver

uv pip install --torch-backend=cu128

```

**Note:** When `--torch-backend=auto` is used, `uv` attempts to infer the correct wheel from the driver, but this is **not reliable** for vLLM installations according to the notebook comments in `run_with_vllm.ipynb` (line 42). Explicitly specifying `cu130` or `cu128` avoids ambiguity.

## Common Failure Symptoms and Root Causes

When troubleshooting CUDA driver compatibility with PyTorch backends for Cosmos 3, look for these specific symptoms:

| Symptom | Likely Cause |
|---------|--------------|
| `torch.cuda.is_available()` returns `False` | Mismatch between driver and torch wheel (wrong `cuXXX` tag). |
| `ImportError: libcuda.so` or "CUDA driver version is insufficient for CUDA runtime version" | Using a CUDA 13 wheel (`cu130`) on a driver that only supports CUDA 12. |
| GPU not listed by `torch.cuda.device_count()` | Driver/wheel mismatch or missing `LD_LIBRARY_PATH` when using the NGC PyTorch container. |

## Step-by-Step Troubleshooting Guide

Follow these steps to resolve CUDA compatibility issues in Cosmos 3:

1. **Check your driver version** using `nvidia-smi`:
   
   ```bash
   nvidia-smi
   ```

   
   The "CUDA Version" field indicates the maximum CUDA runtime your driver supports.

2. **Select the matching backend tag** based on the table above (`cu128` for CUDA 12.x, `cu130` for CUDA 13.x).

3. **Install dependencies with the correct group or flag**:
   
   - For **Cosmos Framework**:
     
     ```bash
     export COSMOS3_UV_GROUP=cu128-train   # Adjust for your driver

     uv sync --all-extras --group=$COSMOS3_UV_GROUP
     ```

   
   - For **Diffusers/Transformers/vLLM**:
     
     ```bash
     uv pip install --torch-backend=cu128 ...
     ```

4. **Verify PyTorch detects the GPU** using the verification script documented in the repository.

5. **If verification fails**:
   - Clear the **LD_LIBRARY_PATH** environment variable after activating the virtual environment when using the NGC PyTorch container.
   - Re-run `uv sync` with the correct group; old wheels may be cached in `~/.cache/uv`.

6. **Repeat verification** until `torch.cuda.is_available()` reports `True` and the device name displays correctly.

## Verification Scripts

Use these scripts to confirm your CUDA environment is configured correctly before running Cosmos 3 notebooks.

### Complete Environment Check

Run this Python snippet from the repository's verification section to inspect torch and CUDA versions:

```bash
.venv/bin/python - <<'PY'
import torch

print("torch:", torch.__version__)
print("torch cuda:", torch.version.cuda)
print("cuda available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
if torch.cuda.is_available():
    print("device 0:", torch.cuda.get_device_name(0))
PY

```

### Notebook GPU Inspection

When running Cosmos 3 notebooks (such as `run_with_cosmos_framework.ipynb` at line 250), use this code to list available devices:

```python
print("torch cuda:", torch.version.cuda)
print("cuda available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
for i in range(torch.cuda.device_count()):
    print(f"device {i}:", torch.cuda.get_device_name(i))

```

### Installing Diffusers with Correct Backend

For CUDA 12 hosts installing the Diffusers backend:

```bash
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate

uv pip install --torch-backend=cu128 \
    "diffusers @ git+https://github.com/huggingface/diffusers.git" \
    accelerate av cosmos_guardrail huggingface_hub imageio imageio-ffmpeg \
    torch torchvision transformers

```

This ensures the `torch` wheel matches your CUDA 12 driver capabilities.

## Summary

- **Match the `cuXXX` tag** to your NVIDIA driver version: use `cu128` for CUDA 12.x drivers and `cu130` for CUDA 13.x drivers.
- **Use `uv sync` with the appropriate group** (`cu128-train` or `cu130-train`) for the Cosmos Framework backend.
- **Set `--torch-backend`** explicitly when installing Diffusers, Transformers, or vLLM to avoid automatic detection failures.
- **Clear `LD_LIBRARY_PATH`** after activating the virtual environment when using NGC containers to prevent library conflicts.
- **Verify with the provided Python scripts** before executing notebooks to ensure `torch.cuda.is_available()` returns `True`.

## Frequently Asked Questions

### Why does torch.cuda.is_available() return False even with a GPU installed?

This occurs when the PyTorch wheel's CUDA version exceeds what your NVIDIA driver supports. For example, installing a `cu130` wheel on a host with a CUDA 12 driver causes the CUDA runtime to fail initialization. Check your driver with `nvidia-smi` and reinstall using the matching `cu128` or `cu130` tag via `uv sync` or `uv pip install`.

### What is the difference between cu128 and cu130 wheels?

The `cu128` wheel is compiled against CUDA 12.8 libraries and requires a driver supporting CUDA 12.x, while `cu130` requires CUDA 13.x driver support. The digits correspond to the CUDA toolkit version used during compilation. Using the wrong version results in "insufficient driver" errors when PyTorch attempts to initialize the CUDA context.

### Can I use --torch-backend=auto for vLLM installations?

No, you should not rely on `--torch-backend=auto` for vLLM or vLLM-Omni backends in Cosmos 3. According to the source code comments in `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` (line 42), automatic detection is unreliable for these specific backends. Always specify `--torch-backend=cu128` or `--torch-backend=cu130` explicitly to ensure the correct wheel is installed.

### How do I fix ImportError: libcuda.so errors in Cosmos 3?

This error typically appears when using the NGC PyTorch container alongside Cosmos 3's virtual environment. The solution is to clear the `LD_LIBRARY_PATH` environment variable after activating your `.venv` but before running `uv sync`. This prevents the container's system libraries from conflicting with the `uv`-installed PyTorch wheels. If the error persists, verify you are using the correct `cuXXX` tag for your driver version.