# Configuring Tensor Parallelism for Cosmos 3 Super on Multi-GPU Setups

> Discover how to configure tensor parallelism for Cosmos 3 Super on multi-GPU setups. Reduce memory usage by sharding model weights and synchronizing activations with the tensor parallelism flag. Learn more.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-13

---

**Tensor parallelism shards the 64B parameter Cosmos 3 Super model across multiple GPUs using the `--tensor-parallel-size` flag, reducing per-GPU memory consumption by dividing transformer weights and synchronizing activations via All-Reduce.**

Cosmos 3 Super is the 64 billion parameter variant of NVIDIA's Cosmos 3 family, requiring distributed inference across multiple GPUs to fit in memory. According to the NVIDIA/cosmos repository, enabling tensor parallelism is mandatory for deploying this checkpoint, as the model weights exceed the capacity of any single consumer or datacenter GPU. This guide explains how to configure tensor parallelism for multi-GPU setups using both native vLLM-Omni and Docker-based deployments.

## What Tensor Parallelism Does for Cosmos 3 Super

Tensor parallelism splits transformer weights—including QKV projections, linear layers, and embeddings—across selected GPUs so each device holds only a fraction of the model. During forward passes, each GPU computes its slice of the tensor and collaborates via All-Reduce operations to produce final activations. This architecture reduces per-GPU memory consumption roughly by a factor of `1 / tensor_parallel_size`.

The Cosmos 3 Super implementation in [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) demonstrates that each GPU stores only a partitioned subset of the 64B parameters, enabling inference on hardware that would otherwise be insufficient.

## When to Enable Tensor Parallelism

Any deployment running the **Cosmos 3 Super** checkpoint on more than one GPU must enable tensor parallelism. The reference configuration specifically uses four GPUs, setting `--tensor-parallel-size 4` to distribute the model adequately. Without this configuration, the vLLM server will exhaust available memory during model loading.

As documented in the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 312-321), you should also add `--enable-layerwise-offload` when running the Super model to further reduce peak GPU memory by moving transformer blocks between CPU and GPU during computation.

## Configuring Tensor Parallelism in vLLM-Omni

### Native vLLM-Omni (Non-Docker) Setup

For bare-metal deployments with four GPUs, export the necessary environment variables and launch the server with tensor parallelism enabled. The command references specific model architecture overrides required for the Cosmos 3 Super checkpoint:

```bash

# From the repository root

export COSMOS3_MEDIA_ROOT="$(pwd)/cookbooks/cosmos3"
export VLLM_PORT="${VLLM_PORT:-8001}"

CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve nvidia/Cosmos3-Super \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --tensor-parallel-size 4 \
  --mm-encoder-tp-mode data \
  --async-scheduling \
  --allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
  --media-io-kwargs '{"video": {"num_frames": -1}}' \
  --port "$VLLM_PORT" \
  --enable-layerwise-offload \
  --init-timeout 1800

```

The `--tensor-parallel-size 4` parameter instructs vLLM to shard the model across four GPUs, while `--enable-layerwise-offload` activates CPU offloading for transformer blocks.

### Docker-Based Configuration

For containerized deployments, mount the necessary volumes and pass the same tensor parallelism arguments to the vLLM-Omni image:

```bash
docker run --runtime nvidia --gpus all \
  -v "${HF_HOME:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
  -v "$(pwd):/workspace" \
  -p "${COSMOS3_HOST_PORT:-8000}:8000" --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Super \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --port 8000 \
    --init-timeout 1800

```

The `--ipc=host` flag is essential for allowing NCCL communication between containers when utilizing multiple GPUs.

## Combining Tensor Parallelism with Other Strategies

Tensor parallelism can operate alongside **CFG-parallel** (classifier-free guidance parallel) and **Ulysses sequence parallelism** as long as the total product of parallel degrees does not exceed your available GPU count. As documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 49-57), the total GPU requirement equals `tensor_parallel_size × cfg_parallel_size × ulysses_degree`.

To add supplementary parallelism dimensions, append these flags to your launch command:

```bash
--cfg-parallel-size 2 \
--ulysses-degree 2 \

```

For example, combining `--tensor-parallel-size 4` with `--cfg-parallel-size 2` requires eight total GPUs. The [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) file contains benchmark tables explicitly validating 4-GPU tensor-parallel configurations for Cosmos 3 Super.

## Memory Optimization Through Layer-Wise Offload

While tensor parallelism distributes weights horizontally across GPUs, **layer-wise offload** (`--enable-layerwise-offload`) moves transformer blocks vertically between CPU DRAM and GPU VRAM. This technique is particularly critical for Cosmos 3 Super because it reduces peak memory usage during inference by keeping only active layers on the GPU.

The NVIDIA/cosmos documentation emphasizes that layer-wise offload works synergistically with tensor parallelism, allowing the 64B model to run on commodity hardware configurations that would otherwise require significantly more expensive GPU setups.

## Summary

- **Tensor parallelism is mandatory** for Cosmos 3 Super (64B parameters) and requires the `--tensor-parallel-size` flag set to the number of GPUs available.
- **Memory reduction scales linearly** with the tensor parallel size, roughly dividing per-GPU consumption by the number of participating devices.
- **Reference configuration uses 4 GPUs** as shown in [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) and [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md).
- **Layer-wise offload** (`--enable-layerwise-offload`) should always accompany tensor parallelism for Super deployments to minimize peak GPU memory.
- **Parallelism dimensions multiply**: ensure `tensor_parallel_size × cfg_parallel_size × ulysses_degree` does not exceed your physical GPU count.

## Frequently Asked Questions

### How many GPUs do I need for Cosmos 3 Super inference?

The reference implementation recommends **four GPUs** for tensor parallelism when running the Cosmos 3 Super checkpoint. While you could theoretically use fewer with aggressive offloading, the official cookbooks and benchmarks specifically validate the 4-GPU configuration as the optimal balance between throughput and memory efficiency.

### Can I combine tensor parallelism with other parallelism methods?

Yes. According to the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), you can combine tensor parallelism with CFG-parallel (`--cfg-parallel-size`) and Ulysses sequence parallelism (`--ulysses-degree`) provided that the **product of all degrees** does not exceed your available GPU count. For example, tensor parallelism size 4 combined with CFG-parallel size 2 requires eight total GPUs.

### What is the difference between tensor parallelism and layer-wise offload?

**Tensor parallelism** horizontally shards model weights across GPUs during inference, with each GPU computing partial results that are synchronized via All-Reduce. **Layer-wise offload** vertically moves entire transformer blocks between CPU and GPU memory, keeping only the currently executing layer resident in VRAM. For Cosmos 3 Super, you should enable both: tensor parallelism to distribute the 64B parameters, and layer-wise offload to further reduce peak memory consumption per device.

### Where are the configuration examples located in the repository?

The definitive configuration examples reside in three key files: [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) contains detailed launch commands for each backend including the exact `--tensor-parallel-size` flag usage; [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 312-321) explains the high-level interaction between tensor parallelism and layer-wise offload; and [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) provides validated performance tables for 4-GPU tensor-parallel setups.