# How to Enable CFG Parallelism in Cosmos 3: Run Positive and Negative Branches on Separate GPUs

> Enable CFG Parallelism in Cosmos 3 to run positive and negative branches on separate GPUs. Slash generation latency by half using --cfg-parallel-size 2 with no extra memory costs.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-06

---

**Set `--cfg-parallel-size 2` (or higher) when launching the vLLM-Omni server or Cosmos Framework CLI to distribute the classifier-free guidance negative and positive passes across distinct GPUs, cutting generation latency roughly in half without increasing per-device memory consumption.**

NVIDIA Cosmos 3 implements *classifier-free guidance* (CFG) via two separate inference passes: an unconditioned (negative) pass and a conditioned (positive) pass. By default, the `Cosmos3OmniDiffusersPipeline` executes these passes sequentially on the same GPU, doubling end-to-end latency. Enabling CFG parallelism allows each branch to run simultaneously on dedicated hardware, and this guide details the exact flags, resource formulas, and configuration files required to activate it.

## Understanding CFG Parallelism Architecture

CFG parallelism splits the standard guidance computation across multiple workers. In the default sequential mode, the model first computes the negative logits, then the positive logits, and finally aggregates them using the formula `logits = negative + guidance_scale * (positive – negative)`. When `--cfg-parallel-size` is set to a value greater than 1, the vLLM-Omni server (or Cosmos Framework runtime) spawns independent inference workers—each bound to a specific GPU via `CUDA_VISIBLE_DEVICES`—and executes both passes in parallel. The first GPU typically handles the negative pass while the second handles the positive pass, though this ordering is configurable. The server then aggregates the results before returning the final output to the client.

## Enabling CFG Parallelism via vLLM-Omni

The vLLM-Omni server exposes parallelism controls through dedicated CLI flags. To run the positive and negative branches on separate GPUs, launch the server with `--cfg-parallel-size 2` and ensure `--tensor-parallel-size` is set to 1 (unless combining with tensor parallelism, which requires additional GPUs).

```bash

# Launch Cosmos 3-Nano with 2-GPU CFG parallelism

vllm serve nvidia/Cosmos3-Nano \
  --omni \
  --model-class-name Cosmos3OmniDiffusersPipeline \
  --allowed-local-media-path / \
  --cfg-parallel-size 2 \
  --tensor-parallel-size 1 \
  --port 8000 \
  --init-timeout 1800

```

In this configuration, the server creates two distinct workers: one bound to `CUDA_VISIBLE_DEVICES=0` for the negative pass and one to `CUDA_VISIBLE_DEVICES=1` for the positive pass. For a complete working example, refer to the notebook at `cosmos/cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb`.

## Launching with the Cosmos Framework CLI

If you are using the native Cosmos Framework instead of vLLM-Omni, the same flag syntax applies. The CLI automatically handles GPU allocation and worker spawning based on the `--cfg-parallel-size` argument.

```bash
cosmos serve \
  --model nvidia/Cosmos3-Nano \
  --cfg-parallel-size 2 \
  --tensor-parallel-size 1 \
  --port 8000

```

Both methods reference the architecture documentation in [`cosmos/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cosmos/README.md), which explains how `--cfg-parallel-size` composes with other parallelism dimensions.

## GPU Allocation and Worker Binding

Under the hood, the vLLM-Omni server manages a pool of worker processes equal to the value of `--cfg-parallel-size`. Each worker initializes its own copy of the model weights and is isolated to a specific GPU index. The default assignment maps the negative (unconditioned) branch to the first available GPU and the positive (conditioned) branch to the second. Because each branch holds a full replica of the model parameters, the per-GPU memory footprint remains identical to a non-parallel deployment; however, the total memory consumption across the node scales linearly with `cfg_parallel_size`.

## Resource Requirements and Sizing Formulas

CFG parallelism multiplies with other distributed strategies. The total number of required GPUs must satisfy the following formula documented in the repository root [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md):

```

tensor_parallel_size × cfg_parallel_size × ulysses_degree

```

For example, a configuration with `--tensor-parallel-size 2`, `--cfg-parallel-size 2`, and a Ulysses sequence parallelism degree of 1 requires exactly four GPUs. Attempting to run with fewer devices will result in initialization errors or CUDA out-of-memory failures. Because each CFG branch loads a complete model instance, CFG parallelism is most efficient with smaller model variants (e.g., *Cosmos3-Nano*) or when tensor parallelism is disabled.

## Configuring Client Requests

Once the server is running with CFG parallelism enabled, clients interact with it identically to sequential mode. The `guidance_scale` parameter controls the strength of the classifier-free guidance and is supplied per request; no additional flags are required to utilize the parallel backend.

```json
{
  "model": "nvidia/Cosmos3-Nano",
  "messages": [
    {"role": "system", "content": [{"type":"text","text":"You are a helpful assistant."}]},
    {"role": "user",   "content": [{"type":"text","text":"A red robot is pouring water into a glass."}]}
  ],
  "guidance_scale": 7.5
}

```

The server routes the request to both workers, aggregates the logits using the standard CFG formula, and returns the guided output without exposing the parallelism mechanism to the client.

## Summary

- **Enable parallel branches** by setting `--cfg-parallel-size 2` (or higher) in either the vLLM-Omni server or Cosmos Framework CLI.
- **Calculate total GPUs** using `tensor_parallel_size × cfg_parallel_size × ulysses_degree` to avoid resource conflicts.
- **Understand worker mapping**: The first GPU typically runs the negative pass, the second the positive pass, each holding a full model replica.
- **Maintain client compatibility**: Use the standard `guidance_scale` field in JSON requests; parallelism is transparent to the API consumer.
- **Consult reference files**: [`cosmos/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cosmos/README.md) and [`cosmos/cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cosmos/cookbooks/cosmos3/README.md) provide the authoritative parallelism formulas, while `cosmos/cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb` offers a runnable implementation.

## Frequently Asked Questions

### What is the minimum GPU requirement for CFG parallelism?

You need at least two GPUs to benefit from CFG parallelism, as the feature requires one device for the negative pass and one for the positive pass when `--cfg-parallel-size 2` is set. If you also enable tensor parallelism, multiply the base requirement accordingly using the formula `tensor_parallel_size × cfg_parallel_size × ulysses_degree`.

### Does CFG parallelism increase total VRAM consumption?

Yes, total cluster VRAM usage increases linearly with `cfg_parallel_size` because each worker loads a full copy of the model weights into its assigned GPU. However, the per-GPU memory footprint remains identical to a single-device deployment, meaning you trade total memory capacity for reduced latency.

### Can I combine CFG parallelism with tensor parallelism?

Yes, the flags are orthogonal. You can set `--tensor-parallel-size 2` and `--cfg-parallel-size 2` simultaneously, which shards each CFG branch across multiple GPUs. This configuration requires four GPUs total and is useful for running larger models like *Cosmos3-Small* or *Cosmos3-Large* with guidance enabled.

### How do I verify that CFG parallelism is active?

Check the server logs during initialization for worker spawning messages that reference distinct `CUDA_VISIBLE_DEVICES` assignments. Additionally, inspect `cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb` for validation scripts that assert the negative and positive logits are computed on separate devices before aggregation.