Troubleshooting vLLM DeepGEMM Errors with Cosmos 3: A Complete Guide

Set the environment variable VLLM_USE_DEEP_GEMM=0 before launching the vLLM server to disable DeepGEMM and resolve compatibility warnings while maintaining inference functionality.

When deploying the Cosmos 3 Omni model using vLLM, you may encounter warnings stating that DeepGEMM is unavailable. This issue occurs when CUDA drivers or PyTorch wheels mismatch the DeepGEMM kernel requirements, causing the server to fall back to slower cuBLAS kernels. Understanding how to troubleshoot and configure DeepGEMM in the NVIDIA Cosmos repository ensures optimal inference performance without startup failures.

Why DeepGEMM Becomes Unavailable in Cosmos 3

The Cosmos 3 vLLM-Omni server relies on DeepGEMM kernels for accelerated matrix multiplication on NVIDIA GPUs. According to the source code in cookbooks/cosmos3/README.md, the server loads both the reasoner and diffusion components, which require specific CUDA runtime support. When DeepGEMM cannot initialize, the system falls back to standard cuBLAS kernels, resulting in significantly slower inference speeds.

CUDA Driver Version Mismatch

If your CUDA driver is older than the CUDA 13 or 12.8 wheel that vLLM was compiled against, DeepGEMM initialization fails. The default behavior sets VLLM_USE_DEEP_GEMM=1, which triggers the startup warning: "DeepGEMM is unavailable". Verify your driver compatibility using:

nvidia-smi  # Displays driver version

python -c "import torch; print(torch.version.cuda)"  # Shows PyTorch CUDA version

PyTorch Wheel Compatibility

Generic PyTorch builds or CPU-only wheels lack the DeepGEMM kernels entirely. When the PyTorch + CUDA wheel used by vLLM does not contain these kernels, the initialization silently fails and logs the unavailability warning.

Container Image Limitations

Custom Dockerfiles that omit the DeepGEMM plugin will trigger the same warning. The official vllm/vllm-omni:cosmos3 image includes these components, but older or custom builds may exclude them.

Disabling DeepGEMM Safely

The NVIDIA Cosmos documentation explicitly recommends disabling DeepGEMM when encountering compatibility issues. As noted in README.md at lines 464-468, setting VLLM_USE_DEEP_GEMM=0 prevents the server from attempting to load unavailable kernels.

Shell Environment Configuration

For local development without Docker, export the variable before launching the server:

export VLLM_USE_DEEP_GEMM=0
vllm serve nvidia/Cosmos3-Nano \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --port 8000 \
    --init-timeout 1800

Docker Runtime Configuration

When using the official vLLM-Omni container, pass the environment variable via the -e flag:

docker run --runtime nvidia --gpus all \
  -e VLLM_USE_DEEP_GEMM=0 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v "$(pwd):/workspace" \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Nano \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

Persistent Environment Files

For repeated experiments, create a .env file in the repository root containing:

VLLM_USE_DEEP_GEMM=0

When using uv for dependency management, running uv sync automatically loads variables from .env, ensuring the flag persists across sessions.

When to Keep DeepGEMM Enabled

You should maintain DeepGEMM activation only when running compatible hardware and software configurations. The performance benchmarks in inference_benchmarks.md demonstrate significant speed improvements when DeepGEMM functions correctly.

CUDA 13 Driver Requirements

Enable DeepGEMM explicitly when using CUDA 13 drivers (or matching CUDA 12.8 drivers) with vLLM compiled against --torch-backend=cu130 or cu128:

export VLLM_USE_DEEP_GEMM=1

Before enabling, confirm that your driver version meets or exceeds the CUDA major version used in your PyTorch wheel. Mismatched configurations will continue to generate warnings and fallback to cuBLAS.

Environment Compatibility Checklist

Follow this verification sequence to resolve DeepGEMM errors in Cosmos 3 deployments:

  1. Verify driver-CUDA alignment – Ensure your NVIDIA driver supports the CUDA version packaged with your PyTorch installation
  2. Select appropriate torch backends – Install using uv pip install --torch-backend=cu130 torch for CUDA 13 compatibility
  3. Disable DeepGEMM if incompatible – Set VLLM_USE_DEEP_GEMM=0 via shell, Docker -e flag, or .env file when drivers mismatch
  4. Restart the vLLM server – Apply changes by restarting the process; the warning should disappear and inference will proceed using standard kernels

Summary

  • DeepGEMM errors in Cosmos 3 occur when CUDA drivers or PyTorch wheels lack the required kernel support, forcing fallbacks to slower cuBLAS operations
  • Immediate fix: Set VLLM_USE_DEEP_GEMM=0 before starting the server, as documented in README.md lines 464-468
  • Configuration methods: Export shell variables, pass Docker -e flags, or use persistent .env files with uv sync
  • Performance trade-off: Disabling DeepGEMM resolves startup warnings but reduces inference speed; enable only when running CUDA 13 compatible drivers with matching PyTorch wheels
  • Verification: Use nvidia-smi and torch.version.cuda to check compatibility before enabling DeepGEMM

Frequently Asked Questions

Why does my Cosmos 3 vLLM server show "DeepGEMM is unavailable"?

This warning appears when your CUDA driver version is older than the CUDA 13 or 12.8 wheel used to compile vLLM, or when your PyTorch installation lacks the DeepGEMM kernels. The server continues running but falls back to standard cuBLAS kernels, resulting in slower inference performance as documented in the inference_benchmarks.md file.

How do I permanently disable DeepGEMM for all Cosmos 3 runs?

Create a .env file in your repository root containing VLLM_USE_DEEP_GEMM=0. When using uv for environment management, running uv sync automatically loads these variables. Alternatively, add export VLLM_USE_DEEP_GEMM=0 to your shell profile for persistent terminal sessions.

Can I use DeepGEMM with older CUDA drivers?

No. DeepGEMM requires CUDA 13 drivers (or the matching CUDA 12.8 driver) and vLLM compiled with --torch-backend=cu130 or cu128. Attempting to enable DeepGEMM with incompatible drivers will trigger the unavailability warning. Verify compatibility using nvidia-smi to check your driver version against torch.version.cuda before enabling the feature.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →