# How to Configure Tensor Parallelism for Cosmos 3 Super Model

> Learn how to configure tensor parallelism for the Cosmos 3 Super model. Boost performance by sharding the 64B parameter model across multiple GPUs using vLLM.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-05

---

**Configure tensor parallelism for Cosmos 3 Super by setting `--tensor-parallel-size 4` (or higher) when launching the vLLM server to shard the 64B parameter model across multiple GPUs, optionally combining with `--enable-layerwise-offload` to reduce peak memory usage.**

Cosmos 3 Super is a 64 billion parameter video world model developed by NVIDIA that exceeds the memory capacity of any single GPU. To deploy this model for inference, you must configure **tensor parallelism** to distribute weight tensors across multiple devices. This guide explains the exact flags, resource calculations, and deployment methods based on the official NVIDIA Cosmos repository.

## Understanding Tensor Parallelism in Cosmos 3 Super

Tensor parallelism splits individual weight matrices across GPUs, allowing each device to store only a fraction of the model parameters. For Cosmos 3 Super, the `--tensor-parallel-size` parameter determines how many GPUs participate in sharding.

According to the repository's main documentation in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 311), the per-GPU memory footprint scales inversely with the tensor parallel size: `model_size / tensor_parallel_size`. When vLLM serves the model, it reconstructs full tensors on-the-fly during inference from these shards.

The implementation follows a strict **parallelism product rule** documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 317-319). If you combine tensor parallelism with other strategies like CFG parallelism (`--cfg-parallel-size`) or Ulysses sequence parallelism (`--ulysses-degree`), the product of all parallel degrees must not exceed your total available GPU count.

## GPU Requirements and Resource Planning

Cosmos 3 Super requires significant GPU memory that no single accelerator currently provides. The base configuration assumes:

- **Minimum 4 GPUs** for tensor parallelism alone (`--tensor-parallel-size 4`)
- **8 GPUs** if combining tensor parallelism (`4`) with CFG parallelism (`--cfg-parallel-size 2`)
- **16 GPUs** for full tensor + CFG + Ulysses (`4 × 2 × 2`) configurations

Each GPU should have sufficient VRAM to hold its shard of the 64B parameters plus activation buffers. The `--enable-layerwise-offload` flag (defined in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), line 311) can extend this to workstations with less VRAM by offloading transformer blocks to CPU RAM between inference steps, trading latency for memory capacity.

## Essential Configuration Flags

The following flags control tensor parallelism behavior in the Cosmos 3 Super deployment:

| Flag | Source Location | Purpose |
|------|----------------|---------|
| `--tensor-parallel-size` | [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 311) | Sets the number of GPUs for splitting weight tensors |
| `--enable-layerwise-offload` | [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 311) | Offloads transformer blocks to CPU between steps to reduce peak GPU memory |
| `--cfg-parallel-size` | [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 317-319) | Enables parallelization of Classifier-Free Guidance branches |
| `--ulysses-degree` | [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 317-319) | Configures Ulysses sequence parallelism for long sequences |
| `--init-timeout` | [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) (lines 40-56) | Sets server initialization timeout (required 1800s for Super model compilation) |
| `--mm-encoder-tp-mode data` | [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) (line 160) | Uses data parallelism for the visual encoder (Reasoner cookbook) |

## Deployment Methods

### Native vLLM Installation

For bare-metal deployments with a Python virtual environment, specify the visible devices and tensor parallel size directly:

```bash

# Activate environment

source .venv/bin/activate

# Launch with tensor parallelism across 4 GPUs

CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve nvidia/Cosmos3-Super \
  --tensor-parallel-size 4 \
  --enable-layerwise-offload \
  --init-timeout 1800 \
  --port 8000

```

The `--init-timeout 1800` flag (30 minutes) is mandatory for Cosmos 3 Super because the 64B checkpoint requires extensive CUDA graph compilation that exceeds default timeouts.

### Docker Deployment (vLLM-Omni)

The vLLM-Omni Docker image provides a containerized environment with all dependencies pre-installed. The same tensor parallelism flags apply:

```bash
docker run --runtime nvidia --gpus all \
  -v "${HF_HOME:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
  -v "$(pwd):/workspace" \
  -p 8000:8000 --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Super \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --port 8000 \
    --init-timeout 1800

```

This configuration appears in the Docker recipe documented in [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) (lines 40-56).

### Combining with CFG Parallelism

For advanced configurations utilizing Classifier-Free Guidance parallelism, adjust your GPU count accordingly:

```bash
vllm serve nvidia/Cosmos3-Super \
  --tensor-parallel-size 4 \
  --cfg-parallel-size 2 \
  --enable-layerwise-offload \
  --init-timeout 1800 \
  --port 8000

```

**Critical:** This configuration requires at least 8 GPUs because `tensor_parallel_size (4) × cfg_parallel_size (2) = 8`.

## Optimization Strategies

**Layer-Wise Offloading**: Enable `--enable-layerwise-offload` when GPU memory is constrained. As implemented in the Cosmos 3 source, this moves entire transformer blocks to host RAM when not actively computing, allowing the model to run on fewer or smaller GPUs at the cost of inference latency.

**Timeout Configuration**: Always include `--init-timeout 1800` when loading Cosmos 3 Super. The model weight initialization and CUDA graph compilation for 64B parameters can take 15-30 minutes depending on storage I/O and GPU interconnect speed.

**Device Visibility**: Use `CUDA_VISIBLE_DEVICES` to restrict tensor parallelism to specific GPUs rather than defaulting to all available devices. This is essential when sharing a node with other workloads or when specific GPUs have different memory capacities.

## Summary

- **Tensor parallelism is mandatory** for Cosmos 3 Super (64B parameters) because the model exceeds single-GPU memory limits
- Use `--tensor-parallel-size 4` as the baseline configuration, requiring 4 GPUs minimum
- The product of all parallelism degrees (tensor × CFG × Ulysses) must not exceed total available GPUs
- Add `--enable-layerwise-offload` to reduce peak GPU memory by utilizing CPU RAM for inactive layers
- Always set `--init-timeout 1800` to prevent server timeout during the lengthy model compilation phase
- Configuration flags are defined in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 311, 317-319) and deployment examples are provided in [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md)

## Frequently Asked Questions

### How many GPUs do I need to run Cosmos 3 Super with tensor parallelism?

You need a minimum of 4 GPUs to run Cosmos 3 Super with `--tensor-parallel-size 4`. If you enable additional parallelism strategies like CFG parallelism (`--cfg-parallel-size 2`), you must multiply the degrees: 4 × 2 = 8 GPUs required. The model will not load if the GPU count is less than the product of your configured parallelism sizes.

### What is the difference between tensor parallelism and layer-wise offloading?

**Tensor parallelism** splits individual weight matrices across multiple GPUs so each stores a fraction of every layer, enabling simultaneous computation across devices. **Layer-wise offloading** moves entire transformer blocks to CPU RAM between inference steps, keeping only the currently executing layer on GPU. Tensor parallelism improves throughput while layer-wise offloading reduces memory requirements at the cost of latency.

### Can I combine tensor parallelism with other forms of parallelism?

Yes, Cosmos 3 Super supports combining tensor parallelism with CFG parallelism (`--cfg-parallel-size`) and Ulysses sequence parallelism (`--ulysses-degree`). However, the product of all parallel degrees must be less than or equal to your total GPU count. For example, `--tensor-parallel-size 4` combined with `--ulysses-degree 2` requires 8 GPUs as documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 317-319).

### Why does the vLLM server timeout when initializing Cosmos 3 Super?

The 64B parameter checkpoint requires extensive CUDA graph compilation and weight loading that exceeds default server timeouts. You must add `--init-timeout 1800` (30 minutes) to your serve command, as shown in the cookbooks at [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) (lines 40-56). Without this flag, the server aborts during the initialization phase before inference can begin.