Troubleshooting CUDA Driver Compatibility with PyTorch Backends for Cosmos 3

Cosmos 3 requires matching the CUDA wheel tag (cu128 or cu130) to your NVIDIA driver version; otherwise, torch.cuda.is_available() returns False and GPU acceleration fails.

Cosmos 3 relies on PyTorch, vLLM, and Transformers wheels that are pre-compiled for specific CUDA versions. When the CUDA version baked into the wheel does not match the NVIDIA driver on your host, the CUDA runtime fails to load, forcing all notebooks to fall back to CPU or abort entirely. This guide explains how to align your driver version with the correct PyTorch backend tag using the uv package manager as implemented in the NVIDIA/cosmos repository.

Understanding the CUDA Driver and Wheel Tag Relationship

The cuXXX tag in a PyTorch wheel name indicates which CUDA runtime version the binary expects. In cookbooks/cosmos3/README.md, the documentation establishes a clear mapping between driver capabilities and backend tags.

The Driver-to-Tag Mapping Table

Driver CUDA Backend Tag Installation Group
13.x cu130 Default for most Cosmos 3 notebooks
12.x cu128 Required for hosts with CUDA 12 drivers

Selecting the wrong tag causes the wheel to expect a newer driver than is present, triggering runtime errors. The wheels are hosted on the PyPI index used by uv, and each group pins a specific torch wheel (e.g., torch==X.Y.Z+cu130).

How Cosmos 3 Selects the PyTorch Wheel

Cosmos 3 uses uv to manage dependencies and selects the appropriate CUDA wheel through dependency groups or backend flags.

Cosmos Framework Backend

For the Cosmos Framework backend, use the uv sync command with a specific group:


# For CUDA 13 drivers

uv sync --all-extras --group=cu130-train

# For CUDA 12 drivers

uv sync --all-extras --group=cu128-train

You can override the default group using the COSMOS3_UV_GROUP environment variable:

export COSMOS3_UV_GROUP=cu128-train   # Use when host has CUDA 12.x driver

uv sync --all-extras --group=$COSMOS3_UV_GROUP

Diffusers, Transformers, and vLLM Backends

For the Diffusers, Transformers, vLLM, and vLLM-Omni backends, specify the backend via the --torch-backend flag:


# CUDA 13 driver

uv pip install --torch-backend=cu130

# CUDA 12 driver

uv pip install --torch-backend=cu128

Note: When --torch-backend=auto is used, uv attempts to infer the correct wheel from the driver, but this is not reliable for vLLM installations according to the notebook comments in run_with_vllm.ipynb (line 42). Explicitly specifying cu130 or cu128 avoids ambiguity.

Common Failure Symptoms and Root Causes

When troubleshooting CUDA driver compatibility with PyTorch backends for Cosmos 3, look for these specific symptoms:

Symptom Likely Cause
torch.cuda.is_available() returns False Mismatch between driver and torch wheel (wrong cuXXX tag).
ImportError: libcuda.so or "CUDA driver version is insufficient for CUDA runtime version" Using a CUDA 13 wheel (cu130) on a driver that only supports CUDA 12.
GPU not listed by torch.cuda.device_count() Driver/wheel mismatch or missing LD_LIBRARY_PATH when using the NGC PyTorch container.

Step-by-Step Troubleshooting Guide

Follow these steps to resolve CUDA compatibility issues in Cosmos 3:

  1. Check your driver version using nvidia-smi:

    nvidia-smi

    The "CUDA Version" field indicates the maximum CUDA runtime your driver supports.

  2. Select the matching backend tag based on the table above (cu128 for CUDA 12.x, cu130 for CUDA 13.x).

  3. Install dependencies with the correct group or flag:

    • For Cosmos Framework:

      export COSMOS3_UV_GROUP=cu128-train   # Adjust for your driver
      
      uv sync --all-extras --group=$COSMOS3_UV_GROUP
    • For Diffusers/Transformers/vLLM:

      uv pip install --torch-backend=cu128 ...
  4. Verify PyTorch detects the GPU using the verification script documented in the repository.

  5. If verification fails:

    • Clear the LD_LIBRARY_PATH environment variable after activating the virtual environment when using the NGC PyTorch container.
    • Re-run uv sync with the correct group; old wheels may be cached in ~/.cache/uv.
  6. Repeat verification until torch.cuda.is_available() reports True and the device name displays correctly.

Verification Scripts

Use these scripts to confirm your CUDA environment is configured correctly before running Cosmos 3 notebooks.

Complete Environment Check

Run this Python snippet from the repository's verification section to inspect torch and CUDA versions:

.venv/bin/python - <<'PY'
import torch

print("torch:", torch.__version__)
print("torch cuda:", torch.version.cuda)
print("cuda available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
if torch.cuda.is_available():
    print("device 0:", torch.cuda.get_device_name(0))
PY

Notebook GPU Inspection

When running Cosmos 3 notebooks (such as run_with_cosmos_framework.ipynb at line 250), use this code to list available devices:

print("torch cuda:", torch.version.cuda)
print("cuda available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
for i in range(torch.cuda.device_count()):
    print(f"device {i}:", torch.cuda.get_device_name(i))

Installing Diffusers with Correct Backend

For CUDA 12 hosts installing the Diffusers backend:

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate

uv pip install --torch-backend=cu128 \
    "diffusers @ git+https://github.com/huggingface/diffusers.git" \
    accelerate av cosmos_guardrail huggingface_hub imageio imageio-ffmpeg \
    torch torchvision transformers

This ensures the torch wheel matches your CUDA 12 driver capabilities.

Summary

  • Match the cuXXX tag to your NVIDIA driver version: use cu128 for CUDA 12.x drivers and cu130 for CUDA 13.x drivers.
  • Use uv sync with the appropriate group (cu128-train or cu130-train) for the Cosmos Framework backend.
  • Set --torch-backend explicitly when installing Diffusers, Transformers, or vLLM to avoid automatic detection failures.
  • Clear LD_LIBRARY_PATH after activating the virtual environment when using NGC containers to prevent library conflicts.
  • Verify with the provided Python scripts before executing notebooks to ensure torch.cuda.is_available() returns True.

Frequently Asked Questions

Why does torch.cuda.is_available() return False even with a GPU installed?

This occurs when the PyTorch wheel's CUDA version exceeds what your NVIDIA driver supports. For example, installing a cu130 wheel on a host with a CUDA 12 driver causes the CUDA runtime to fail initialization. Check your driver with nvidia-smi and reinstall using the matching cu128 or cu130 tag via uv sync or uv pip install.

What is the difference between cu128 and cu130 wheels?

The cu128 wheel is compiled against CUDA 12.8 libraries and requires a driver supporting CUDA 12.x, while cu130 requires CUDA 13.x driver support. The digits correspond to the CUDA toolkit version used during compilation. Using the wrong version results in "insufficient driver" errors when PyTorch attempts to initialize the CUDA context.

Can I use --torch-backend=auto for vLLM installations?

No, you should not rely on --torch-backend=auto for vLLM or vLLM-Omni backends in Cosmos 3. According to the source code comments in cookbooks/cosmos3/reasoner/run_with_vllm.ipynb (line 42), automatic detection is unreliable for these specific backends. Always specify --torch-backend=cu128 or --torch-backend=cu130 explicitly to ensure the correct wheel is installed.

How do I fix ImportError: libcuda.so errors in Cosmos 3?

This error typically appears when using the NGC PyTorch container alongside Cosmos 3's virtual environment. The solution is to clear the LD_LIBRARY_PATH environment variable after activating your .venv but before running uv sync. This prevents the container's system libraries from conflicting with the uv-installed PyTorch wheels. If the error persists, verify you are using the correct cuXXX tag for your driver version.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →