# How to Configure Tensor Parallelism for Cosmos3-Super on Multiple GPUs to Reduce Memory Usage

> Configure tensor parallelism for Cosmos3-Super on multiple GPUs with vllm serve. Set the --tensor-parallel-size flag to shard weights and reduce VRAM usage. Optimize your LLM deployment now.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-06

---

**Set the `--tensor-parallel-size` flag equal to the number of GPUs when you run `vllm serve` with Cosmos3-Super to shard the 64B model’s weights across devices and lower per-GPU VRAM by roughly that factor.**

The NVIDIA/cosmos repository hosts Cosmos3-Super, a 64 billion parameter “Super” model built for efficient multi-GPU inference. To reduce per-GPU memory usage and serve this architecture on hardware with limited VRAM, you configure tensor parallelism for Cosmos3-Super through vLLM’s command-line interface, splitting weight tensors across multiple CUDA devices. The repository’s cookbooks and benchmarks provide tested launch commands and parallelism strategies.

## Configure Tensor Parallelism Using `--tensor-parallel-size`

The primary way to configure tensor parallelism for Cosmos3-Super is the `--tensor-parallel-size` flag passed to `vllm serve`. This value sets how many GPUs will shard the model’s large weight matrices—such as attention and feed-forward layers—so each device stores only a fraction of the parameters.

As shown in the reasoner cookbook [`cookbooks/cosmos3/reasoner/run_with_vllm.ipynb`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/run_with_vllm.ipynb) (lines 145–169), a standard four-GPU launch looks like this:

```bash
export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype float16

```

By distributing weights across four GPUs, the memory requirement per device drops by roughly a factor of four, enabling the 64B model to run on cards with limited individual VRAM.

## Restrict GPUs with `CUDA_VISIBLE_DEVICES`

Before launching the server, use the `CUDA_VISIBLE_DEVICES` environment variable to define exactly which physical GPUs participate in the tensor-parallel group. The number of devices you expose must be at least equal to the value passed to `--tensor-parallel-size`.

If you skip this step, vLLM may attempt to use all visible GPUs, which can clash with other workloads or exceed your intended topology.

## Lower Memory Further with Layer-Wise Offload

For scenarios where tensor parallelism alone leaves insufficient headroom, the NVIDIA/cosmos codebase supports an optional `--enable-layerwise-offload` flag. This moves entire transformer blocks to CPU when they are not actively being computed, trading additional CPU RAM and a small latency penalty for lower peak GPU memory.

The audiovisual cookbook [`cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb) (lines 66–81) demonstrates this combination:

```bash
export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype float16

```

Add this flag when you need to minimize per-GPU allocation at the cost of inference speed.

## Advanced: Combine Tensor Parallelism with CFG and Ulysses

The repository supports stacking multiple parallelism axes. When you combine tensor parallelism with classifier-free guidance (CFG) parallelism or Ulysses sequence parallelism, the total GPU count required is the **product** of the individual degrees.

According to the “Combining parallelism options” paragraph in [[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 321), a configuration using `--tensor-parallel-size 4`, `--cfg-parallel-size 2`, and `--ulysses-degree 2` demands 16 GPUs:

```bash
export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --cfg-parallel-size 2 \
    --ulysses-degree 2 \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype bfloat16

```

Consult [[`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md)](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) (lines 159–306) for the full list of command-line options.

## Additional Flags for Large-Model Deployment

Several supporting vLLM flags cited in the Cosmos3-Super documentation help stabilize large deployments:

- **`--dtype float16` or `--dtype bfloat16`** — Halves model memory compared to float32, as noted across the cookbooks.
- **`--init-timeout <seconds>`** — Prevents timeout during the long weight-loading phase for the 64B model. The main [[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 312) shows an example using an 1800-second timeout.
- **Benchmark baselines** — The [[`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md)](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) file (lines 130 and 231) validates that 4× and 8× GPU configurations rely on tensor parallelism for throughput and memory scaling.

## Summary

- **Use `--tensor-parallel-size N`** with `vllm serve` to shard Cosmos3-Super’s 64B parameters across N GPUs, reducing per-GPU VRAM by roughly a factor of N.
- **Set `CUDA_VISIBLE_DEVICES`** to expose exactly the GPUs you want in the tensor-parallel group.
- **Add `--enable-layerwise-offload`** when you need to trade CPU RAM and latency for even lower peak GPU memory.
- **Stack parallelism axes** with care: the total GPU count is the product of `--tensor-parallel-size`, `--cfg-parallel-size`, and `--ulysses-degree`.
- **Refer to the official cookbooks** in `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` and `cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb` for complete, tested launch commands.

## Frequently Asked Questions

### What is the minimum number of GPUs needed to run Cosmos3-Super with tensor parallelism?

There is no fixed minimum enforced by the codebase, but the repository’s examples and benchmarks typically demonstrate 4-GPU and 8-GPU tensor-parallel configurations. You must set `--tensor-parallel-size` to a value that divides evenly across the model’s layers, and expose at least that many devices via `CUDA_VISIBLE_DEVICES`.

### Does `--enable-layerwise-offload` eliminate the need for multiple GPUs?

No. Layer-wise offload reduces peak GPU memory by moving transformer blocks to CPU, but it does not shard weights across devices. The omni cookbook shows both flags used together because offload is designed to complement tensor parallelism, not replace it.

### How do I calculate the total GPUs when mixing tensor, CFG, and Ulysses parallelism?

Multiply the degrees of each parallelism axis to find the total GPU count. For example, `--tensor-parallel-size 4`, `--cfg-parallel-size 2`, and `--ulysses-degree 2` require 4 × 2 × 2 = 16 GPUs. This rule is documented in the main [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) at line 321.

### Which data type should I use for Cosmos3-Super inference?

Use `--dtype float16` or `--dtype bfloat16`. Both precision formats halve memory usage compared to float32 and are supported by vLLM for the Cosmos3-Super architecture.