# Configuring Tensor Parallelism for Cosmos3-Super Inference on Multiple GPUs

> Configure tensor parallelism for Cosmos3-Super inference on multiple GPUs. Use vLLM-Omni and set tensor parallel size to shard weights and reduce memory usage by 75%.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**To run Cosmos3-Super (64B parameters) across multiple GPUs, set `--tensor-parallel-size 4` in your vLLM-Omni server command to shard transformer weights and reduce per-GPU memory usage by roughly 75%.**

Cosmos3-Super is the 64 B parameter variant of NVIDIA’s Cosmos 3 family, requiring distributed inference to avoid out-of-memory errors on single devices. Because the model exceeds the capacity of individual GPUs, the NVIDIA/cosmos repository implements tensor parallelism to split computation across multiple accelerators.

## What is Tensor Parallelism for Cosmos3-Super?

**Tensor parallelism** shards transformer weights—including QKV projections, linear layers, and embeddings—across the selected GPUs so that each device holds only a fraction of the model parameters. During a forward pass, each GPU computes its local slice of the tensor and collaborates via All‑Reduce operations to produce the final activations.

This approach reduces per‑GPU memory consumption approximately by a factor of `1 / tensor_parallel_size`. For the Cosmos3-Super model, splitting across four GPUs cuts per-device memory usage to roughly 25% of the total model footprint, enabling inference on standard consumer or datacenter hardware.

## When to Enable Tensor Parallelism

Any deployment hosting the **Cosmos3-Super** checkpoint on more than one GPU must enable tensor parallelism. Without it, the server will exhaust available VRAM during model loading or initial inference. The flag is unnecessary only if you are running smaller variants (such as base instruct models) that fit within a single GPU’s memory constraints.

## Configuration Syntax and Parameters

The primary mechanism for enabling this feature is the `--tensor-parallel-size` flag passed to the vLLM-Omni server. According to the source code in [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) (lines 24–33), the reference configuration for Super uses four GPUs, setting the flag to `4`.

### Layer-Wise Offload Integration

For the Super model specifically, the top‑level [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 312–321) recommends pairing tensor parallelism with `--enable-layerwise-offload`. This technique moves transformer blocks between CPU and GPU dynamically, further reducing peak GPU memory pressure during long sequences or high batch sizes.

### Combining Parallelism Dimensions

Tensor parallelism can operate alongside other parallelism strategies as long as the total product of degrees does not exceed your available GPU count. As documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 49–57), you may combine:

- **CFG‑parallel** (`--cfg-parallel-size`): Splits positive and negative guidance across GPUs
- **Ulysses sequence parallelism** (`--ulysses-degree`): Distributes sequence chunks across devices

The total GPU requirement equals `tensor_parallel_size × cfg_parallel_size × ulysses_degree`.

## Implementation Examples

### Native vLLM-Omni (Non-Docker)

The following command launches Cosmos3-Super with four‑way tensor parallelism and layer‑wise offload, as specified in the cookbooks:

```bash

# From the repository root

export COSMOS3_MEDIA_ROOT="$(pwd)/cookbooks/cosmos3"
export VLLM_PORT="${VLLM_PORT:-8001}"

CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve nvidia/Cosmos3-Super \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --tensor-parallel-size 4 \
  --mm-encoder-tp-mode data \
  --async-scheduling \
  --allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
  --media-io-kwargs '{"video": {"num_frames": -1}}' \
  --port "$VLLM_PORT" \
  --enable-layerwise-offload \
  --init-timeout 1800

```

### Docker-Based Deployment

For containerized environments, mount your Hugging Face cache and workspace, then pass the same parallelism flags:

```bash
docker run --runtime nvidia --gpus all \
  -v "${HF_HOME:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
  -v "$(pwd):/workspace" \
  -p "${COSMOS3_HOST_PORT:-8000}:8000" --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Super \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --port 8000 \
    --init-timeout 1800

```

### Advanced Multi-Dimensional Parallelism

To add CFG‑parallel and Ulysses sequence parallelism alongside your tensor parallelism, append the additional flags:

```bash
--cfg-parallel-size 2 \
--ulysses-degree 2 \

```

Ensure your hardware provides at least `4 × 2 × 2 = 16` GPUs before enabling this configuration.

## Performance and Memory Considerations

The [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) file explicitly documents 4‑GPU tensor‑parallel configurations for Cosmos3-Super, establishing this topology as the reference standard for production deployments. While tensor parallelism introduces communication overhead via All‑Reduce, the memory savings allow the 64 B parameter model to run on commodity multi‑GPU servers rather than requiring a single massive accelerator.

## Summary

- **Tensor parallelism** is mandatory for Cosmos3-Super (64B) inference, splitting weights across GPUs to prevent OOM errors.
- Set `--tensor-parallel-size 4` to distribute the model across four GPUs, reducing per‑device memory by approximately 75%.
- Combine with `--enable-layerwise-offload` (as recommended in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)) to further optimize peak memory usage.
- You can stack additional parallelism methods (CFG, Ulysses) provided the product of all degrees fits your available GPU count.
- Reference implementations exist in [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) and [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md).

## Frequently Asked Questions

### What GPU count is required for Cosmos3-Super inference?

You need at least four GPUs for the standard tensor‑parallel configuration. The model’s 64 B parameters exceed the memory capacity of single‑GPU setups, making distributed inference essential according to the repository’s configuration examples.

### How does tensor parallelism differ from layer-wise offload?

**Tensor parallelism** splits individual weight matrices across GPUs (spatial sharding), while **layer-wise offload** moves entire transformer blocks between CPU and GPU memory (temporal offloading). The former reduces per‑GPU memory during computation; the latter reduces peak memory during idle phases. The Super model benefits from using both simultaneously.

### Can I combine tensor parallelism with other parallelism methods?

Yes. You can enable CFG‑parallel (`--cfg-parallel-size`) and Ulysses sequence parallelism (`--ulysses-degree`) alongside tensor parallelism. Ensure the product of `tensor_parallel_size`, `cfg_parallel_size`, and `ulysses_degree` does not exceed your physical GPU count, as documented in the repository’s parallelism guidelines.

### Why does my server fail to start with a single GPU?

Cosmos3-Super contains 64 B parameters, which cannot fit into the VRAM of a single consumer or datacenter GPU. Without `--tensor-parallel-size` set to divide the model weights across multiple devices, the vLLM‑Omni server will raise an out‑of‑memory error during checkpoint loading.