How to Configure Tensor Parallelism for Cosmos 3 Super Model
Configure tensor parallelism for Cosmos 3 Super by setting --tensor-parallel-size 4 (or higher) when launching the vLLM server to shard the 64B parameter model across multiple GPUs, optionally combining with --enable-layerwise-offload to reduce peak memory usage.
Cosmos 3 Super is a 64 billion parameter video world model developed by NVIDIA that exceeds the memory capacity of any single GPU. To deploy this model for inference, you must configure tensor parallelism to distribute weight tensors across multiple devices. This guide explains the exact flags, resource calculations, and deployment methods based on the official NVIDIA Cosmos repository.
Understanding Tensor Parallelism in Cosmos 3 Super
Tensor parallelism splits individual weight matrices across GPUs, allowing each device to store only a fraction of the model parameters. For Cosmos 3 Super, the --tensor-parallel-size parameter determines how many GPUs participate in sharding.
According to the repository's main documentation in README.md (line 311), the per-GPU memory footprint scales inversely with the tensor parallel size: model_size / tensor_parallel_size. When vLLM serves the model, it reconstructs full tensors on-the-fly during inference from these shards.
The implementation follows a strict parallelism product rule documented in README.md (lines 317-319). If you combine tensor parallelism with other strategies like CFG parallelism (--cfg-parallel-size) or Ulysses sequence parallelism (--ulysses-degree), the product of all parallel degrees must not exceed your total available GPU count.
GPU Requirements and Resource Planning
Cosmos 3 Super requires significant GPU memory that no single accelerator currently provides. The base configuration assumes:
- Minimum 4 GPUs for tensor parallelism alone (
--tensor-parallel-size 4) - 8 GPUs if combining tensor parallelism (
4) with CFG parallelism (--cfg-parallel-size 2) - 16 GPUs for full tensor + CFG + Ulysses (
4 × 2 × 2) configurations
Each GPU should have sufficient VRAM to hold its shard of the 64B parameters plus activation buffers. The --enable-layerwise-offload flag (defined in README.md, line 311) can extend this to workstations with less VRAM by offloading transformer blocks to CPU RAM between inference steps, trading latency for memory capacity.
Essential Configuration Flags
The following flags control tensor parallelism behavior in the Cosmos 3 Super deployment:
| Flag | Source Location | Purpose |
|---|---|---|
--tensor-parallel-size |
README.md (line 311) |
Sets the number of GPUs for splitting weight tensors |
--enable-layerwise-offload |
README.md (line 311) |
Offloads transformer blocks to CPU between steps to reduce peak GPU memory |
--cfg-parallel-size |
README.md (lines 317-319) |
Enables parallelization of Classifier-Free Guidance branches |
--ulysses-degree |
README.md (lines 317-319) |
Configures Ulysses sequence parallelism for long sequences |
--init-timeout |
cookbooks/cosmos3/README.md (lines 40-56) |
Sets server initialization timeout (required 1800s for Super model compilation) |
--mm-encoder-tp-mode data |
cookbooks/cosmos3/README.md (line 160) |
Uses data parallelism for the visual encoder (Reasoner cookbook) |
Deployment Methods
Native vLLM Installation
For bare-metal deployments with a Python virtual environment, specify the visible devices and tensor parallel size directly:
# Activate environment
source .venv/bin/activate
# Launch with tensor parallelism across 4 GPUs
CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve nvidia/Cosmos3-Super \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--init-timeout 1800 \
--port 8000
The --init-timeout 1800 flag (30 minutes) is mandatory for Cosmos 3 Super because the 64B checkpoint requires extensive CUDA graph compilation that exceeds default timeouts.
Docker Deployment (vLLM-Omni)
The vLLM-Omni Docker image provides a containerized environment with all dependencies pre-installed. The same tensor parallelism flags apply:
docker run --runtime nvidia --gpus all \
-v "${HF_HOME:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
-v "$(pwd):/workspace" \
-p 8000:8000 --ipc=host \
vllm/vllm-omni:cosmos3 \
vllm serve nvidia/Cosmos3-Super \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--port 8000 \
--init-timeout 1800
This configuration appears in the Docker recipe documented in cookbooks/cosmos3/README.md (lines 40-56).
Combining with CFG Parallelism
For advanced configurations utilizing Classifier-Free Guidance parallelism, adjust your GPU count accordingly:
vllm serve nvidia/Cosmos3-Super \
--tensor-parallel-size 4 \
--cfg-parallel-size 2 \
--enable-layerwise-offload \
--init-timeout 1800 \
--port 8000
Critical: This configuration requires at least 8 GPUs because tensor_parallel_size (4) × cfg_parallel_size (2) = 8.
Optimization Strategies
Layer-Wise Offloading: Enable --enable-layerwise-offload when GPU memory is constrained. As implemented in the Cosmos 3 source, this moves entire transformer blocks to host RAM when not actively computing, allowing the model to run on fewer or smaller GPUs at the cost of inference latency.
Timeout Configuration: Always include --init-timeout 1800 when loading Cosmos 3 Super. The model weight initialization and CUDA graph compilation for 64B parameters can take 15-30 minutes depending on storage I/O and GPU interconnect speed.
Device Visibility: Use CUDA_VISIBLE_DEVICES to restrict tensor parallelism to specific GPUs rather than defaulting to all available devices. This is essential when sharing a node with other workloads or when specific GPUs have different memory capacities.
Summary
- Tensor parallelism is mandatory for Cosmos 3 Super (64B parameters) because the model exceeds single-GPU memory limits
- Use
--tensor-parallel-size 4as the baseline configuration, requiring 4 GPUs minimum - The product of all parallelism degrees (tensor × CFG × Ulysses) must not exceed total available GPUs
- Add
--enable-layerwise-offloadto reduce peak GPU memory by utilizing CPU RAM for inactive layers - Always set
--init-timeout 1800to prevent server timeout during the lengthy model compilation phase - Configuration flags are defined in
README.md(lines 311, 317-319) and deployment examples are provided incookbooks/cosmos3/README.md
Frequently Asked Questions
How many GPUs do I need to run Cosmos 3 Super with tensor parallelism?
You need a minimum of 4 GPUs to run Cosmos 3 Super with --tensor-parallel-size 4. If you enable additional parallelism strategies like CFG parallelism (--cfg-parallel-size 2), you must multiply the degrees: 4 × 2 = 8 GPUs required. The model will not load if the GPU count is less than the product of your configured parallelism sizes.
What is the difference between tensor parallelism and layer-wise offloading?
Tensor parallelism splits individual weight matrices across multiple GPUs so each stores a fraction of every layer, enabling simultaneous computation across devices. Layer-wise offloading moves entire transformer blocks to CPU RAM between inference steps, keeping only the currently executing layer on GPU. Tensor parallelism improves throughput while layer-wise offloading reduces memory requirements at the cost of latency.
Can I combine tensor parallelism with other forms of parallelism?
Yes, Cosmos 3 Super supports combining tensor parallelism with CFG parallelism (--cfg-parallel-size) and Ulysses sequence parallelism (--ulysses-degree). However, the product of all parallel degrees must be less than or equal to your total GPU count. For example, --tensor-parallel-size 4 combined with --ulysses-degree 2 requires 8 GPUs as documented in README.md (lines 317-319).
Why does the vLLM server timeout when initializing Cosmos 3 Super?
The 64B parameter checkpoint requires extensive CUDA graph compilation and weight loading that exceeds default server timeouts. You must add --init-timeout 1800 (30 minutes) to your serve command, as shown in the cookbooks at cookbooks/cosmos3/README.md (lines 40-56). Without this flag, the server aborts during the initialization phase before inference can begin.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →