How to Configure Tensor Parallelism for Cosmos3-Super on Multiple GPUs to Reduce Memory Usage
Set the --tensor-parallel-size flag equal to the number of GPUs when you run vllm serve with Cosmos3-Super to shard the 64B model’s weights across devices and lower per-GPU VRAM by roughly that factor.
The NVIDIA/cosmos repository hosts Cosmos3-Super, a 64 billion parameter “Super” model built for efficient multi-GPU inference. To reduce per-GPU memory usage and serve this architecture on hardware with limited VRAM, you configure tensor parallelism for Cosmos3-Super through vLLM’s command-line interface, splitting weight tensors across multiple CUDA devices. The repository’s cookbooks and benchmarks provide tested launch commands and parallelism strategies.
Configure Tensor Parallelism Using --tensor-parallel-size
The primary way to configure tensor parallelism for Cosmos3-Super is the --tensor-parallel-size flag passed to vllm serve. This value sets how many GPUs will shard the model’s large weight matrices—such as attention and feed-forward layers—so each device stores only a fraction of the parameters.
As shown in the reasoner cookbook cookbooks/cosmos3/reasoner/run_with_vllm.ipynb (lines 145–169), a standard four-GPU launch looks like this:
export CUDA_VISIBLE_DEVICES=0,1,2,3
vllm serve nvidia/Cosmos3-Super \
--tensor-parallel-size 4 \
--model-id nvidia/Cosmos3-Super \
--port 8000 \
--dtype float16
By distributing weights across four GPUs, the memory requirement per device drops by roughly a factor of four, enabling the 64B model to run on cards with limited individual VRAM.
Restrict GPUs with CUDA_VISIBLE_DEVICES
Before launching the server, use the CUDA_VISIBLE_DEVICES environment variable to define exactly which physical GPUs participate in the tensor-parallel group. The number of devices you expose must be at least equal to the value passed to --tensor-parallel-size.
If you skip this step, vLLM may attempt to use all visible GPUs, which can clash with other workloads or exceed your intended topology.
Lower Memory Further with Layer-Wise Offload
For scenarios where tensor parallelism alone leaves insufficient headroom, the NVIDIA/cosmos codebase supports an optional --enable-layerwise-offload flag. This moves entire transformer blocks to CPU when they are not actively being computed, trading additional CPU RAM and a small latency penalty for lower peak GPU memory.
The audiovisual cookbook cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb (lines 66–81) demonstrates this combination:
export CUDA_VISIBLE_DEVICES=0,1,2,3
vllm serve nvidia/Cosmos3-Super \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--model-id nvidia/Cosmos3-Super \
--port 8000 \
--dtype float16
Add this flag when you need to minimize per-GPU allocation at the cost of inference speed.
Advanced: Combine Tensor Parallelism with CFG and Ulysses
The repository supports stacking multiple parallelism axes. When you combine tensor parallelism with classifier-free guidance (CFG) parallelism or Ulysses sequence parallelism, the total GPU count required is the product of the individual degrees.
According to the “Combining parallelism options” paragraph in [README.md](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 321), a configuration using --tensor-parallel-size 4, --cfg-parallel-size 2, and --ulysses-degree 2 demands 16 GPUs:
export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15
vllm serve nvidia/Cosmos3-Super \
--tensor-parallel-size 4 \
--cfg-parallel-size 2 \
--ulysses-degree 2 \
--model-id nvidia/Cosmos3-Super \
--port 8000 \
--dtype bfloat16
Consult [cookbooks/cosmos3/README.md](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) (lines 159–306) for the full list of command-line options.
Additional Flags for Large-Model Deployment
Several supporting vLLM flags cited in the Cosmos3-Super documentation help stabilize large deployments:
--dtype float16or--dtype bfloat16— Halves model memory compared to float32, as noted across the cookbooks.--init-timeout <seconds>— Prevents timeout during the long weight-loading phase for the 64B model. The main [README.md](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 312) shows an example using an 1800-second timeout.- Benchmark baselines — The [
inference_benchmarks.md](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) file (lines 130 and 231) validates that 4× and 8× GPU configurations rely on tensor parallelism for throughput and memory scaling.
Summary
- Use
--tensor-parallel-size Nwithvllm serveto shard Cosmos3-Super’s 64B parameters across N GPUs, reducing per-GPU VRAM by roughly a factor of N. - Set
CUDA_VISIBLE_DEVICESto expose exactly the GPUs you want in the tensor-parallel group. - Add
--enable-layerwise-offloadwhen you need to trade CPU RAM and latency for even lower peak GPU memory. - Stack parallelism axes with care: the total GPU count is the product of
--tensor-parallel-size,--cfg-parallel-size, and--ulysses-degree. - Refer to the official cookbooks in
cookbooks/cosmos3/reasoner/run_with_vllm.ipynbandcookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynbfor complete, tested launch commands.
Frequently Asked Questions
What is the minimum number of GPUs needed to run Cosmos3-Super with tensor parallelism?
There is no fixed minimum enforced by the codebase, but the repository’s examples and benchmarks typically demonstrate 4-GPU and 8-GPU tensor-parallel configurations. You must set --tensor-parallel-size to a value that divides evenly across the model’s layers, and expose at least that many devices via CUDA_VISIBLE_DEVICES.
Does --enable-layerwise-offload eliminate the need for multiple GPUs?
No. Layer-wise offload reduces peak GPU memory by moving transformer blocks to CPU, but it does not shard weights across devices. The omni cookbook shows both flags used together because offload is designed to complement tensor parallelism, not replace it.
How do I calculate the total GPUs when mixing tensor, CFG, and Ulysses parallelism?
Multiply the degrees of each parallelism axis to find the total GPU count. For example, --tensor-parallel-size 4, --cfg-parallel-size 2, and --ulysses-degree 2 require 4 × 2 × 2 = 16 GPUs. This rule is documented in the main README.md at line 321.
Which data type should I use for Cosmos3-Super inference?
Use --dtype float16 or --dtype bfloat16. Both precision formats halve memory usage compared to float32 and are supported by vLLM for the Cosmos3-Super architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →