Best Practices for Running Cosmos 3 in Docker Containers with NVIDIA NGC Images

Deploy Cosmos 3 in Docker containers using official NVIDIA NGC images by matching your workload to the correct backend (vLLM‑Omni for the Generator or NIM for the Reasoner), authenticating to nvcr.io with your NGC API key, and launching with --runtime=nvidia --gpus all --shm-size=32GB for optimal GPU performance.

The NVIDIA Cosmos repository provides production‑ready containerization strategies for deploying Cosmos 3 world models at scale. Whether you are serving the Generator for multimodal content creation or the Reasoner for physics‑aware inference, running Cosmos 3 in Docker containers with NVIDIA NGC images ensures consistent CUDA toolkit compatibility and optimized driver integration. The following guide references the exact source files and commands found in the NVIDIA/cosmos GitHub repository.

Choose the Appropriate NGC Base Image for Your Workload

Selecting the correct NGC image depends on whether you need the turn‑key Reasoner, the production Generator API, or the full research framework.

Backend Recommended NGC Image CUDA Version Use Case
vLLM‑Omni (Generator) nvcr.io/nvidia/vllm-omni:cosmos3 CUDA 13 (nvcr.io/nvidia/pytorch:25.09-py3) or CUDA 12.8 (nvcr.io/nvidia/pytorch:25.06-py3) Production‑grade OpenAI‑compatible API for image, video, sound, and action generation.
NIM Reasoner nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0 CUDA 13 (default) Turn‑key Reasoner container with no vLLM or CUDA‑driver setup required.
Cosmos Framework (research) Build from source (Dockerfile in the repo) Match your driver (use --torch-backend=cu130 for CUDA 13, --torch-backend=cu128 for CUDA 12.8) Full framework for training, evaluation, and custom inference pipelines.

See the base‑container recommendation in README.md lines 491‑495.

Authenticate to NGC and Configure Environment Variables

Before pulling private NGC images, you must generate an API key and authenticate Docker.

  1. Generate an NGC API key – Visit the NGC catalog page for the Cosmos 3 Reasoner container and click "Generate API Key."

  2. Authenticate Docker – Run a one‑time login using $oauthtoken as the username and your API key as the password:

docker login nvcr.io -u $oauthtoken -p <NGC_API_KEY>

Authentication steps are documented in cookbooks/cosmos3/README.md lines 29‑31.

  1. Export the environment variable before any docker run command:
export NGC_API_KEY=<your_key>

Deploy the Reasoner with NIM (Turn‑Key Container)

The NIM Reasoner provides the fastest path to production for physics‑aware reasoning tasks. The container requires no manual CUDA or vLLM configuration.

export CONTAINER_NAME="nvidia-cosmos3-reasoner"
export IMG_NAME="nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0"
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"

docker run -it --rm --name=$CONTAINER_NAME \
  --runtime=nvidia \
  --gpus all \
  --shm-size=32GB \
  -e NGC_API_KEY=$NGC_API_KEY \
  -e NIM_MODEL_SIZE=nano \
  -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
  -p 8000:8000 \
  $IMG_NAME

The full command block appears in README.md lines 496‑505.

Key configuration details:

  • Use --shm-size=32GB to accommodate large model weights in shared memory.
  • Set NIM_MODEL_SIZE to nano for the 8B parameter model or super for the 64B parameter model.
  • Mount a persistent cache (/opt/nim/.cache) to avoid re‑downloading weights on every restart.

Serve the Generator with vLLM‑Omni (Production API)

For multimodal generation (image, video, audio), use the vLLM‑Omni image with the OpenAI‑compatible API.

docker run --runtime=nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v "$(pwd):/workspace" \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Nano \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

See the vLLM‑Omni launch snippet in README.md lines 292‑306.

Performance optimization flags:

  • --tensor-parallel-size N – Split the 64B Super model across N GPUs.
  • --enable-layerwise-offload – Offload layers to CPU when GPU memory is constrained (increases CPU‑GPU traffic).
  • --cfg-parallel-size 2 – Run classifier‑free guidance branches on two GPUs to improve throughput.
  • --ulysses-degree 2 – Split the sequence dimension for large batch sizes.
  • --ipc=host – Reduces inter‑process communication overhead for high‑throughput serving.

Safety guardrails are enabled by default; disable per‑request via extra_params={"guardrails":false} as shown in README.md lines 94‑100.

Build Custom Images from the Cosmos Framework

When you need the full framework for training or custom inference pipelines, build from the Dockerfile in the repository.

git clone https://github.com/NVIDIA/cosmos.git
cd cosmos

docker build -t cosmos-framework:latest .

The build command appears in cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md lines 19‑25.

Best practices for custom builds:

  • Pass --group=cu130-train (CUDA 13) or --group=cu128-train when syncing dependencies to match your host driver.
  • Always mount the Hugging Face cache (-v $HOME/.cache/huggingface:/root/.cache/huggingface) to speed up subsequent runs.

Essential Docker Run Flags and Resource Management

The following flags are critical for running Cosmos 3 in Docker containers with NVIDIA NGC images:

Flag Purpose
--runtime=nvidia Enables GPU access via the NVIDIA container runtime.
--gpus all Exposes all GPUs; use --gpus '"device=0"' for single‑GPU isolation.
--shm-size=32GB Allocates sufficient shared memory for the NIM Reasoner’s memory‑mapped files.
--ipc=host Optimizes inter‑process communication for vLLM‑Omni high‑throughput serving.
-e NGC_API_KEY=$NGC_API_KEY Authenticates the container to NGC for model downloads.
--init-timeout 1800 Prevents server timeouts during large checkpoint loading (30 minutes).
--allowed-local-media-path / Permits the server to read media files from mounted host directories.

Docker flag explanations are derived from README.md lines 293‑301.

Summary

  • Select the correct image – Use nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0 for turn‑key reasoning and nvcr.io/nvidia/vllm-omni:cosmos3 for production generation APIs.
  • Authenticate once – Run docker login nvcr.io with your NGC API key before pulling private images.
  • Configure GPU and memory – Always include --runtime=nvidia, --gpus all, and --shm-size=32GB for NIM containers.
  • Optimize throughput – Use tensor parallelism (--tensor-parallel-size), layerwise offloading, and sequence parallelism (--ulysses-degree) for the 64B Super model.
  • Persist caches – Mount host directories for ~/.cache/nim and ~/.cache/huggingface to avoid redundant downloads.
  • Pin versions – Use exact image tags (1.7.0, cosmos3) rather than latest for reproducible deployments.

Frequently Asked Questions

How do I authenticate to NGC when running Cosmos 3 containers?

You must generate an NGC API key from the NVIDIA catalog, then run docker login nvcr.io -u $oauthtoken -p <NGC_API_KEY> once on your host. After authentication, export the key as NGC_API_KEY and pass it to the container via -e NGC_API_KEY=$NGC_API_KEY. This process is required for both the NIM Reasoner and vLLM‑Omni images.

What is the difference between the NIM Reasoner and vLLM‑Omni Generator containers?

The NIM Reasoner (nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0) is a turn‑key container optimized for physics‑aware reasoning with no vLLM setup required. The vLLM‑Omni image (nvcr.io/nvidia/vllm-omni:cosmos3) provides an OpenAI‑compatible API surface for multimodal generation (image, video, audio) and requires additional flags like --ipc=host and --allowed-local-media-path for full functionality.

How do I handle CUDA version mismatches when deploying Cosmos 3?

Ensure your host NVIDIA driver supports the CUDA version baked into the NGC image. For CUDA 13, use nvcr.io/nvidia/pytorch:25.09-py3-based images; for CUDA 12.8, use 25.06-py3 variants. When building custom images from the framework, specify --torch-backend=cu130 or --torch-backend=cu128 during dependency synchronization to match your driver capabilities.

Why does my Cosmos 3 container fail with shared memory errors?

The NIM Reasoner requires significant shared memory for memory‑mapped model weights. Add --shm-size=32GB (or larger) to your docker run command. For vLLM‑Omni, use --ipc=host to leverage the host’s shared memory namespace, which reduces overhead for high‑throughput inference but requires careful security considerations in multi‑tenant environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →