How to Configure Tensor Parallelism for Cosmos3-Super on Multiple GPUs

Set the --tensor-parallel-size flag equal to your GPU count when launching the vLLM server to shard the 64B Cosmos3-Super model across devices, reducing per-GPU memory usage proportionally.

The NVIDIA Cosmos repository provides Cosmos3-Super, a 64 billion parameter video world model that requires tensor parallelism to run efficiently on multi-GPU systems. By splitting weight tensors across devices, you can serve this model on hardware with limited VRAM while maintaining inference throughput.

Understanding Tensor Parallelism for Cosmos3-Super

Tensor parallelism distributes the model's weight matrices—specifically attention and feed-forward layers—across multiple GPUs. Each device holds only a fractional slice of the parameters, enabling the full 64B model to fit within available memory.

Weight Sharding Implementation

When configured with --tensor-parallel-size 4, the framework shards each transformer layer across four GPUs. Each device stores approximately 25% of the layer weights, reducing memory requirements per GPU by roughly 75%. This pattern is demonstrated in cookbooks/cosmos3/reasoner/run_with_vllm.ipynb (lines 145-169).

GPU Visibility Requirements

Before launching, restrict CUDA device visibility to your target GPUs using CUDA_VISIBLE_DEVICES. This ensures tensor parallelism allocates shards to the correct physical devices and prevents the process from attempting to use unavailable hardware.

Basic 4-GPU Configuration

The standard configuration for Cosmos3-Super uses four GPUs with tensor parallelism enabled. This setup represents the minimum recommended configuration for serving the 64B model with acceptable latency.

Launch Command

Execute the following to distribute the model across four GPUs:

export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype float16

The --tensor-parallel-size 4 parameter instructs vLLM to partition the model weights across the devices specified in CUDA_VISIBLE_DEVICES. According to the implementation in cookbooks/cosmos3/reasoner/run_with_vllm.ipynb, this configuration is sufficient to run the model on GPUs with 24GB+ VRAM when using float16 precision.

Advanced Memory Optimization with Layer-wise Offload

For hardware with limited GPU memory, combine tensor parallelism with layer-wise offloading. This technique moves entire transformer blocks to CPU RAM when not actively computing, further reducing peak GPU memory consumption.

Enabling CPU Offload

Add the --enable-layerwise-offload flag to your launch command, as shown in the audiovisual cookbook at cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb (lines 66-81):

export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype float16

Performance trade-offs: Layer-wise offload consumes CPU RAM equivalent to the model size and introduces latency due to CPU-GPU transfer overhead. Reserve this configuration for deployments where GPU memory constraints outweigh latency requirements.

Combining Parallelism Strategies

Cosmos3-Super supports multiple parallelism axes that multiply together for large-scale deployments. When combining tensor parallelism with Classifier-Free Guidance (CFG) parallelism and Ulysses attention parallelism, the total GPU requirement equals the product of all parallelism degrees.

Multi-Axis Configuration

As documented in README.md (lines 321), the following configuration requires 16 GPUs (4 tensor × 2 CFG × 2 Ulysses):

export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --cfg-parallel-size 2 \
    --ulysses-degree 2 \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype bfloat16

Initialization timeout: For large multi-GPU setups, add --init-timeout 1800 (seconds) to prevent timeout errors during model loading, as recommended in README.md (line 312).

Key Configuration Parameters

Parameter Effect Source Reference
--tensor-parallel-size N Shards model weights across N GPUs cookbooks/cosmos3/reasoner/run_with_vllm.ipynb (lines 145-169)
--enable-layerwise-offload Moves transformer layers to CPU RAM cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb (lines 66-81)
--dtype float16 or bfloat16 Reduces memory by 50% versus float32 README.md (line 474)
--cfg-parallel-size N Parallelizes Classifier-Free Guidance computation README.md (lines 321)
--ulysses-degree N Enables Ulysses attention parallelism README.md (lines 321)
--init-timeout <seconds> Extends model loading timeout for large deployments README.md (line 312)

Summary

  • Tensor parallelism splits the 64B Cosmos3-Super model across GPUs using --tensor-parallel-size, with each GPU holding a proportional fraction of the weights.
  • Four GPUs represent the standard configuration for tensor parallelism, as implemented in cookbooks/cosmos3/reasoner/run_with_vllm.ipynb.
  • Layer-wise offload (--enable-layerwise-offload) complements tensor parallelism by moving inactive layers to CPU, further reducing GPU memory at the cost of latency.
  • Combined parallelism requires multiplying the degrees: tensor × CFG × Ulysses equals total required GPUs.
  • Precision settings (--dtype float16 or bfloat16) are essential for reducing memory footprint on multi-GPU setups.

Frequently Asked Questions

How many GPUs are required to run Cosmos3-Super with tensor parallelism?

The minimum viable configuration typically requires four GPUs with tensor parallelism enabled (--tensor-parallel-size 4), though you can scale to eight or more for improved throughput. As documented in inference_benchmarks.md (lines 130, 231), the repository provides benchmarks for both 4× and 8× GPU tensor parallelism configurations.

Can I combine tensor parallelism with other memory optimization techniques?

Yes. You can simultaneously use --tensor-parallel-size with --enable-layerwise-offload to move transformer layers to CPU RAM when not in use, and --dtype float16 or bfloat16 to halve memory usage versus float32. These strategies are composable and demonstrated across the repository's cookbooks, including the audiovisual reasoning examples.

What is the difference between tensor parallelism and layer-wise offload?

Tensor parallelism splits individual weight tensors horizontally across GPUs, with all devices working simultaneously on different portions of each layer. Layer-wise offload moves entire transformer blocks to CPU memory, keeping only the actively computing layer on GPU. Tensor parallelism reduces per-GPU memory proportionally to the number of GPUs, while layer-wise offload reduces peak GPU memory by caching inactive layers in system RAM.

Where can I find benchmark results for different GPU configurations?

Performance benchmarks for tensor parallelism configurations are documented in inference_benchmarks.md (lines 130, 231). These results detail the throughput and latency characteristics when running Cosmos3-Super on 4× and 8× GPU setups, providing guidance for selecting the appropriate hardware configuration for your deployment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →