# How to Configure Tensor Parallelism and CFG Parallelism for Cosmos 3 Super Model Inference

> Learn to configure tensor parallelism and CFG parallelism for Cosmos 3 Super model inference. Optimize inference speed by sharding models across GPUs and accelerating model guided diffusion.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-14

---

**Use the `--tensor-parallel-size` and `--cfg-parallel-size` flags when launching the vLLM‑Omni server to shard the 64 B Cosmos 3 Super model across GPUs and accelerate classifier‑free guidance by running dual forward passes in parallel.**

The **NVIDIA Cosmos** repository provides a vLLM‑Omni inference engine for the Cosmos 3 Super text‑to‑image/video generation model. Because the Super variant contains 64 billion parameters, single‑GPU inference is impossible; you must enable **tensor parallelism** to distribute weights across devices and optionally enable **CFG parallelism** to accelerate the positive‑ and negative‑prompt branches simultaneously.

## Understanding Tensor Parallelism for Cosmos 3 Super

Tensor parallelism splits the transformer layers of the Cosmos 3 Super model across multiple GPUs, with each device holding only a slice of the parameters. This is essential for fitting the 64 B model into GPU memory.

According to the repository’s [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), you activate tensor parallelism by passing the `--tensor-parallel-size N` argument to the `vllm serve` command. The value `N` represents the number of GPUs across which the model weights are sharded via NCCL communication. When `N=4`, each GPU loads approximately one‑fourth of the model parameters, reducing the per‑GPU memory footprint from impossible levels to roughly **40 GB per slice**.

## Understanding CFG Parallelism for Cosmos 3 Super

**CFG parallelism** (Classifier‑Free Guidance parallelism) accelerates inference by executing the positive‑prompt and negative‑prompt forward passes on separate GPU groups simultaneously. Without this optimization, the two passes run sequentially, doubling latency.

To enable CFG parallelism, set `--cfg-parallel-size N` when launching the server. When `N=2`, the vLLM‑Omni engine assigns distinct GPU sets to each CFG branch, combines the results on the host, and then proceeds to the diffusion step. This configuration roughly halves the latency of the guidance computation compared to sequential execution.

## Combining Parallelism Strategies

When you enable multiple parallelism modes, the total GPU requirement is the product of the individual degrees:

```

Required GPUs ≥ tensor_parallel_size × cfg_parallel_size × ulysses_degree

```

For example, a configuration with `--tensor-parallel-size 4` and `--cfg-parallel-size 2` demands **8 GPUs** (4 × 2). You can also add sequence parallelism via `--ulysses-degree N` (often called Ulysses parallelism) for extremely long sequences, though this is optional for most workloads.

The [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) file in the NVIDIA Cosmos repository provides detailed resource‑size guidance and recommends starting with tensor and CFG parallelism before experimenting with Ulysses degrees.

## Launch Commands and Configuration Examples

### Basic Tensor Parallelism Only

Deploy the Cosmos 3 Super model across 4 GPUs with pure tensor sharding and no CFG acceleration:

```bash
vllm serve nvidia/Cosmos3-Super \
    --omni \
    --tensor-parallel-size 4 \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

```

### Tensor and CFG Parallelism Combined

Run the model on 4 GPUs total, splitting the model weights across 2 GPUs while running the two CFG branches on separate pairs:

```bash
vllm serve nvidia/Cosmos3-Super \
    --omni \
    --tensor-parallel-size 2 \
    --cfg-parallel-size 2 \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

```

### Full Parallelism Configuration

For an 8‑GPU node, maximize throughput by combining tensor parallelism (4) with CFG parallelism (2):

```bash
vllm serve nvidia/Cosmos3-Super \
    --omni \
    --tensor-parallel-size 4 \
    --cfg-parallel-size 2 \
    --ulysses-degree 1 \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

```

### Python Client Integration

After launching the server, send generation requests using the OpenAI‑compatible client. Specify the `guidance_scale` to control CFG strength (do not use `true_cfg_scale`):

```python
import openai

client = openai.OpenAI(base_url="http://localhost:8000/v1")

resp = client.images.generate(
    model="nvidia/Cosmos3-Super",
    prompt="A futuristic cityscape at sunset",
    negative_prompt="low‑resolution, blurry",
    guidance_scale=7.5,
    size="1024x1024",
)

# Access the generated image via resp.data[0].b64_json

```

### Memory Optimization with Layerwise Offload

If GPU memory is still constrained even with tensor parallelism, enable `--enable-layerwise-offload` to move activations to CPU. This reduces per‑GPU memory usage at the cost of increased latency:

```bash
vllm serve nvidia/Cosmos3-Super \
    --omni \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

```

## Resource Requirements and Best Practices

Always verify that your hardware meets the product of the parallelism degrees before launching. The server will abort with a "not enough GPUs" error if the available device count is insufficient.

Memory budgeting requires careful attention. Even with tensor parallelism, each GPU in a Cosmos 3 Super deployment typically requires approximately **40 GB of VRAM** per slice. If your GPUs have less memory, use `--enable-layerwise-offload` to compensate.

CFG parallelism only accelerates the guidance computation phase. The diffusion steps themselves continue to run on the tensor‑parallel group, so the overall speedup follows the formula roughly **1 + 1/`cfg_parallel_size`**. When mixing parallelism strategies, start with tensor and CFG, then experiment with Ulysses degrees only if you have sufficient hardware and need to process extremely long sequences.

## Summary

- **Tensor parallelism** (`--tensor-parallel-size`) shards the 64 B Cosmos 3 Super model across GPUs to fit within memory constraints.
- **CFG parallelism** (`--cfg-parallel-size`) runs positive and negative prompt branches simultaneously, reducing guidance latency.
- **GPU requirements** multiply: ensure you have at least `tensor_parallel_size × cfg_parallel_size × ulysses_degree` GPUs available.
- **Memory optimization** via `--enable-layerwise-offload` allows running on smaller GPUs by offloading layers to CPU.
- **Client configuration** uses `guidance_scale` for CFG strength in OpenAI‑compatible API calls.

## Frequently Asked Questions

### How many GPUs do I need to run Cosmos 3 Super with tensor and CFG parallelism?

You need at least the product of the two parallelism degrees. For example, `--tensor-parallel-size 4` combined with `--cfg-parallel-size 2` requires **8 GPUs** total. The server validates this at startup and will fail if the available GPU count is insufficient.

### Can I run Cosmos 3 Super on fewer than 8 GPUs?

Yes, if you only use tensor parallelism. You can run the model on 4 GPUs using `--tensor-parallel-size 4` without CFG parallelism, or on 2 GPUs using `--tensor-parallel-size 2` and `--cfg-parallel-size 1`. However, you cannot combine high degrees of both parallelism types without meeting the multiplicative GPU requirement.

### What is the difference between `guidance_scale` and `true_cfg_scale` in the client?

Use `guidance_scale` in your API requests to control CFG strength. Do not use `true_cfg_scale`, as the vLLM‑Omni engine handles the parallel CFG execution internally when `--cfg-parallel-size` is configured. The `guidance_scale` parameter determines how strongly the negative prompt influences the final output.

### Why should I enable layerwise offload instead of increasing tensor parallelism?

Layerwise offload (`--enable-layerwise-offload`) is useful when you have limited GPU memory but cannot add more GPUs. Moving layers to CPU reduces VRAM usage below the typical 40 GB per‑GPU requirement, allowing the model to run on smaller devices. Increasing tensor parallelism requires additional GPUs but maintains native GPU speed, whereas offload introduces CPU‑GPU transfer latency.