# How to Configure Tensor Parallelism for Cosmos3-Super on Multiple GPUs

> Configure tensor parallelism for Cosmos3-Super across multiple GPUs. Set tensor parallel size to your GPU count to reduce memory usage on each device.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-12

---

**Set the `--tensor-parallel-size` flag equal to your GPU count when launching the vLLM server to shard the 64B Cosmos3-Super model across devices, reducing per-GPU memory usage proportionally.**

The NVIDIA Cosmos repository provides Cosmos3-Super, a 64 billion parameter video world model that requires tensor parallelism to run efficiently on multi-GPU systems. By splitting weight tensors across devices, you can serve this model on hardware with limited VRAM while maintaining inference throughput.

## Understanding Tensor Parallelism for Cosmos3-Super

Tensor parallelism distributes the model's weight matrices—specifically attention and feed-forward layers—across multiple GPUs. Each device holds only a fractional slice of the parameters, enabling the full 64B model to fit within available memory.

### Weight Sharding Implementation

When configured with `--tensor-parallel-size 4`, the framework shards each transformer layer across four GPUs. Each device stores approximately 25% of the layer weights, reducing memory requirements per GPU by roughly 75%. This pattern is demonstrated in `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` (lines 145-169).

### GPU Visibility Requirements

Before launching, restrict CUDA device visibility to your target GPUs using `CUDA_VISIBLE_DEVICES`. This ensures tensor parallelism allocates shards to the correct physical devices and prevents the process from attempting to use unavailable hardware.

## Basic 4-GPU Configuration

The standard configuration for Cosmos3-Super uses four GPUs with tensor parallelism enabled. This setup represents the minimum recommended configuration for serving the 64B model with acceptable latency.

### Launch Command

Execute the following to distribute the model across four GPUs:

```bash
export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype float16

```

The `--tensor-parallel-size 4` parameter instructs vLLM to partition the model weights across the devices specified in `CUDA_VISIBLE_DEVICES`. According to the implementation in `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb`, this configuration is sufficient to run the model on GPUs with 24GB+ VRAM when using float16 precision.

## Advanced Memory Optimization with Layer-wise Offload

For hardware with limited GPU memory, combine tensor parallelism with layer-wise offloading. This technique moves entire transformer blocks to CPU RAM when not actively computing, further reducing peak GPU memory consumption.

### Enabling CPU Offload

Add the `--enable-layerwise-offload` flag to your launch command, as shown in the audiovisual cookbook at `cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb` (lines 66-81):

```bash
export CUDA_VISIBLE_DEVICES=0,1,2,3

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype float16

```

**Performance trade-offs:** Layer-wise offload consumes CPU RAM equivalent to the model size and introduces latency due to CPU-GPU transfer overhead. Reserve this configuration for deployments where GPU memory constraints outweigh latency requirements.

## Combining Parallelism Strategies

Cosmos3-Super supports multiple parallelism axes that multiply together for large-scale deployments. When combining tensor parallelism with Classifier-Free Guidance (CFG) parallelism and Ulysses attention parallelism, the total GPU requirement equals the product of all parallelism degrees.

### Multi-Axis Configuration

As documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 321), the following configuration requires 16 GPUs (4 tensor × 2 CFG × 2 Ulysses):

```bash
export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15

vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --cfg-parallel-size 2 \
    --ulysses-degree 2 \
    --model-id nvidia/Cosmos3-Super \
    --port 8000 \
    --dtype bfloat16

```

**Initialization timeout:** For large multi-GPU setups, add `--init-timeout 1800` (seconds) to prevent timeout errors during model loading, as recommended in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 312).

## Key Configuration Parameters

| Parameter | Effect | Source Reference |
|-----------|--------|------------------|
| `--tensor-parallel-size N` | Shards model weights across N GPUs | `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` (lines 145-169) |
| `--enable-layerwise-offload` | Moves transformer layers to CPU RAM | `cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb` (lines 66-81) |
| `--dtype float16` or `bfloat16` | Reduces memory by 50% versus float32 | [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 474) |
| `--cfg-parallel-size N` | Parallelizes Classifier-Free Guidance computation | [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 321) |
| `--ulysses-degree N` | Enables Ulysses attention parallelism | [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 321) |
| `--init-timeout <seconds>` | Extends model loading timeout for large deployments | [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 312) |

## Summary

- **Tensor parallelism** splits the 64B Cosmos3-Super model across GPUs using `--tensor-parallel-size`, with each GPU holding a proportional fraction of the weights.
- **Four GPUs** represent the standard configuration for tensor parallelism, as implemented in `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb`.
- **Layer-wise offload** (`--enable-layerwise-offload`) complements tensor parallelism by moving inactive layers to CPU, further reducing GPU memory at the cost of latency.
- **Combined parallelism** requires multiplying the degrees: tensor × CFG × Ulysses equals total required GPUs.
- **Precision settings** (`--dtype float16` or `bfloat16`) are essential for reducing memory footprint on multi-GPU setups.

## Frequently Asked Questions

### How many GPUs are required to run Cosmos3-Super with tensor parallelism?

The minimum viable configuration typically requires **four GPUs** with tensor parallelism enabled (`--tensor-parallel-size 4`), though you can scale to eight or more for improved throughput. As documented in [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) (lines 130, 231), the repository provides benchmarks for both 4× and 8× GPU tensor parallelism configurations.

### Can I combine tensor parallelism with other memory optimization techniques?

Yes. You can simultaneously use `--tensor-parallel-size` with `--enable-layerwise-offload` to move transformer layers to CPU RAM when not in use, and `--dtype float16` or `bfloat16` to halve memory usage versus float32. These strategies are composable and demonstrated across the repository's cookbooks, including the audiovisual reasoning examples.

### What is the difference between tensor parallelism and layer-wise offload?

**Tensor parallelism** splits individual weight tensors horizontally across GPUs, with all devices working simultaneously on different portions of each layer. **Layer-wise offload** moves entire transformer blocks to CPU memory, keeping only the actively computing layer on GPU. Tensor parallelism reduces per-GPU memory proportionally to the number of GPUs, while layer-wise offload reduces peak GPU memory by caching inactive layers in system RAM.

### Where can I find benchmark results for different GPU configurations?

Performance benchmarks for tensor parallelism configurations are documented in [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) (lines 130, 231). These results detail the throughput and latency characteristics when running Cosmos3-Super on 4× and 8× GPU setups, providing guidance for selecting the appropriate hardware configuration for your deployment.