How to Configure Tensor Parallelism for Cosmos3-Super on Multiple GPUs to Reduce Memory Usage

Set the --tensor-parallel-size flag equal to the number of GPUs when you run vllm serve with Cosmos3-Super to shard the 64B model’s weights across devices and lower per-GPU VRAM by roughly that factor.

The NVIDIA/cosmos repository hosts Cosmos3-Super, a 64 billion parameter “Super” model built for efficient multi-GPU inference. To reduce per-GPU memory usage and serve this architecture on hardware with limited VRAM, you configure tensor parallelism for Cosmos3-Super through vLLM’s command-line interface, splitting weight tensors across multiple CUDA devices. The repository’s cookbooks and benchmarks provide tested launch commands and parallelism strategies.

Configure Tensor Parallelism Using --tensor-parallel-size

The primary way to configure tensor parallelism for Cosmos3-Super is the --tensor-parallel-size flag passed to vllm serve. This value sets how many GPUs will shard the model’s large weight matrices—such as attention and feed-forward layers—so each device stores only a fraction of the parameters.

As shown in the reasoner cookbook cookbooks/cosmos3/reasoner/run_with_vllm.ipynb (lines 145–169), a standard four-GPU launch looks like this:

export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype float16

By distributing weights across four GPUs, the memory requirement per device drops by roughly a factor of four, enabling the 64B model to run on cards with limited individual VRAM.

Restrict GPUs with CUDA_VISIBLE_DEVICES

Before launching the server, use the CUDA_VISIBLE_DEVICES environment variable to define exactly which physical GPUs participate in the tensor-parallel group. The number of devices you expose must be at least equal to the value passed to --tensor-parallel-size.

If you skip this step, vLLM may attempt to use all visible GPUs, which can clash with other workloads or exceed your intended topology.

Lower Memory Further with Layer-Wise Offload

For scenarios where tensor parallelism alone leaves insufficient headroom, the NVIDIA/cosmos codebase supports an optional --enable-layerwise-offload flag. This moves entire transformer blocks to CPU when they are not actively being computed, trading additional CPU RAM and a small latency penalty for lower peak GPU memory.

The audiovisual cookbook cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb (lines 66–81) demonstrates this combination:

export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype float16

Add this flag when you need to minimize per-GPU allocation at the cost of inference speed.

Advanced: Combine Tensor Parallelism with CFG and Ulysses

The repository supports stacking multiple parallelism axes. When you combine tensor parallelism with classifier-free guidance (CFG) parallelism or Ulysses sequence parallelism, the total GPU count required is the product of the individual degrees.

According to the “Combining parallelism options” paragraph in [README.md](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 321), a configuration using --tensor-parallel-size 4, --cfg-parallel-size 2, and --ulysses-degree 2 demands 16 GPUs:

export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --cfg-parallel-size 2 \
    --ulysses-degree 2 \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype bfloat16

Consult [cookbooks/cosmos3/README.md](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) (lines 159–306) for the full list of command-line options.

Additional Flags for Large-Model Deployment

Several supporting vLLM flags cited in the Cosmos3-Super documentation help stabilize large deployments:

Summary

  • Use --tensor-parallel-size N with vllm serve to shard Cosmos3-Super’s 64B parameters across N GPUs, reducing per-GPU VRAM by roughly a factor of N.
  • Set CUDA_VISIBLE_DEVICES to expose exactly the GPUs you want in the tensor-parallel group.
  • Add --enable-layerwise-offload when you need to trade CPU RAM and latency for even lower peak GPU memory.
  • Stack parallelism axes with care: the total GPU count is the product of --tensor-parallel-size, --cfg-parallel-size, and --ulysses-degree.
  • Refer to the official cookbooks in cookbooks/cosmos3/reasoner/run_with_vllm.ipynb and cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb for complete, tested launch commands.

Frequently Asked Questions

What is the minimum number of GPUs needed to run Cosmos3-Super with tensor parallelism?

There is no fixed minimum enforced by the codebase, but the repository’s examples and benchmarks typically demonstrate 4-GPU and 8-GPU tensor-parallel configurations. You must set --tensor-parallel-size to a value that divides evenly across the model’s layers, and expose at least that many devices via CUDA_VISIBLE_DEVICES.

Does --enable-layerwise-offload eliminate the need for multiple GPUs?

No. Layer-wise offload reduces peak GPU memory by moving transformer blocks to CPU, but it does not shard weights across devices. The omni cookbook shows both flags used together because offload is designed to complement tensor parallelism, not replace it.

How do I calculate the total GPUs when mixing tensor, CFG, and Ulysses parallelism?

Multiply the degrees of each parallelism axis to find the total GPU count. For example, --tensor-parallel-size 4, --cfg-parallel-size 2, and --ulysses-degree 2 require 4 × 2 × 2 = 16 GPUs. This rule is documented in the main README.md at line 321.

Which data type should I use for Cosmos3-Super inference?

Use --dtype float16 or --dtype bfloat16. Both precision formats halve memory usage compared to float32 and are supported by vLLM for the Cosmos3-Super architecture.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →