# How to Configure CFG Parallelism and Ulysses Sequence Parallelism in NVIDIA Cosmos

> Learn how to configure CFG parallelism and Ulysses sequence parallelism in NVIDIA Cosmos for efficient model serving. Optimize performance with vLLM Omni.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-12

---

**Set `--cfg-parallel-size 2` to run positive and negative CFG branches on separate GPUs, and `--ulysses-degree 2` to split the sequence dimension across GPUs when serving Cosmos 3 models with vLLM‑Omni.**

The NVIDIA Cosmos repository provides production-ready inference optimizations for the Cosmos 3 family of generative AI models. To maximize throughput and minimize latency on multi-GPU deployments, you can configure CFG parallelism and Ulysses sequence parallelism alongside traditional tensor parallelism. These orthogonal strategies distribute computation across different dimensions of the model, allowing you to scale performance without altering the underlying numerics.

## Understanding CFG Parallelism

**CFG parallelism** eliminates the latency penalty of classifier-free guidance (CFG) by executing the positive and negative branches simultaneously on separate GPUs. Normally, CFG requires two sequential forward passes per token: one with the conditioned prompt (positive branch) and one with the unconditioned prompt (negative branch). This process doubles the generation time when run on a single device.

To enable CFG parallelism, set the `--cfg-parallel-size` flag to `2` when launching the vLLM‑Omni server. According to the "Additional parallelism options" table in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 16-20), this configuration launches the two passes on different GPUs, effectively halving the latency of CFG-guided generation. The actual guidance strength is controlled at request time via the `guidance_scale` field in the JSON payload—do not use the deprecated `true_cfg_scale` flag.

## Understanding Ulysses Sequence Parallelism

**Ulysses sequence parallelism** splits the sequence dimension of the transformer across multiple GPUs, allowing a single forward pass to be processed in parallel. For a batch size of 1 and sequence length of 1024, a Ulysses degree of 2 allocates half of the tokens to each GPU, reducing per-layer compute time without changing the model's output.

Enable this feature by setting `--ulysses-degree 2` in your launch command. The same README table that defines CFG parallelism notes that this flag "enables Ulysses sequence parallelism, splitting the sequence dimension across GPUs." The degree can be any integer that divides the sequence length, though 2 is the most common configuration for dual-GPU setups.

## Calculating Total GPU Requirements

When mixing tensor parallelism, CFG parallelism, and Ulysses sequence parallelism, the total number of GPUs required is the product of the three degrees:

```

total_gpus = tensor_parallel_size × cfg_parallel_size × ulysses_degree

```

For example, a configuration with `--tensor-parallel-size 2`, `--cfg-parallel-size 2`, and `--ulysses-degree 2` requires eight GPUs. The server will fail to start if your hardware cannot satisfy this product, so verify your GPU count before launching.

## Implementation Examples

The following command launches a 4-GPU server with tensor parallelism set to 2, CFG parallelism set to 2, and Ulysses degree set to 2:

```bash
vllm serve nvidia/Cosmos3-Nano \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --tensor-parallel-size 2 \
    --cfg-parallel-size 2 \
    --ulysses-degree 2 \
    --port 8000 \
    --init-timeout 1800

```

For a minimal single-GPU deployment without parallelism, omit the multiplicity flags:

```bash
vllm serve nvidia/Cosmos3-Nano \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --port 8000

```

When sending requests to the server, specify the CFG strength using `guidance_scale` in the JSON payload:

```json
{
  "model": "nvidia/Cosmos3-Nano",
  "prompt": "A futuristic city at sunset",
  "guidance_scale": 7.5,
  "max_new_tokens": 128
}

```

## Key Source Files

- [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (root): Contains the authoritative table of parallelism flags and the GPU product calculation rule.
- [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md): Mirrors the root documentation for the Cosmos 3 cookbook, providing quick reference material.
- `cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb`: Demonstrates the `--cfg-parallel-size` flag in a runnable Jupyter notebook example.

## Summary

- **CFG parallelism** (`--cfg-parallel-size 2`) runs positive and negative branches on separate GPUs to halve CFG latency.
- **Ulysses sequence parallelism** (`--ulysses-degree 2`) distributes the sequence dimension across GPUs to accelerate transformer layers.
- **Total GPU count** equals the product of tensor-parallel, CFG-parallel, and Ulysses degrees.
- **Control guidance scale** via the `guidance_scale` JSON field, not legacy flags.
- **Reference files** include the root [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) and the `run_with_vllm_omni.ipynb` notebook.

## Frequently Asked Questions

### What is the minimum GPU requirement for CFG parallelism?

CFG parallelism requires at least 2 GPUs because it splits the positive and negative branches across separate devices. Setting `--cfg-parallel-size 2` on a single GPU will cause the server to fail during initialization.

### Can I use Ulysses sequence parallelism without tensor parallelism?

Yes, Ulysses sequence parallelism operates independently of tensor parallelism. You can set `--ulysses-degree 2` while keeping `--tensor-parallel-size 1`, which is useful for long-sequence workloads where the sequence dimension is the primary bottleneck rather than the model weights.

### How does Ulysses sequence parallelism affect batch processing?

Ulysses splits the sequence dimension across GPUs, so it applies to the tokens within a single sequence. For batch size greater than 1, each sequence in the batch is split across the Ulysses group, meaning all GPUs still participate in processing every batch item. The degree must divide the sequence length evenly to avoid padding inefficiencies.

### Where is the guidance scale configured in Cosmos 3?

The guidance scale is configured client-side in the JSON request payload using the `guidance_scale` field. As implemented in the Cosmos source code, you should not use the older `true_cfg_scale` server flag, which has been deprecated in favor of per-request control.