How to Configure CFG Parallelism and Ulysses Sequence Parallelism in NVIDIA Cosmos

Set --cfg-parallel-size 2 to run positive and negative CFG branches on separate GPUs, and --ulysses-degree 2 to split the sequence dimension across GPUs when serving Cosmos 3 models with vLLM‑Omni.

The NVIDIA Cosmos repository provides production-ready inference optimizations for the Cosmos 3 family of generative AI models. To maximize throughput and minimize latency on multi-GPU deployments, you can configure CFG parallelism and Ulysses sequence parallelism alongside traditional tensor parallelism. These orthogonal strategies distribute computation across different dimensions of the model, allowing you to scale performance without altering the underlying numerics.

Understanding CFG Parallelism

CFG parallelism eliminates the latency penalty of classifier-free guidance (CFG) by executing the positive and negative branches simultaneously on separate GPUs. Normally, CFG requires two sequential forward passes per token: one with the conditioned prompt (positive branch) and one with the unconditioned prompt (negative branch). This process doubles the generation time when run on a single device.

To enable CFG parallelism, set the --cfg-parallel-size flag to 2 when launching the vLLM‑Omni server. According to the "Additional parallelism options" table in README.md (lines 16-20), this configuration launches the two passes on different GPUs, effectively halving the latency of CFG-guided generation. The actual guidance strength is controlled at request time via the guidance_scale field in the JSON payload—do not use the deprecated true_cfg_scale flag.

Understanding Ulysses Sequence Parallelism

Ulysses sequence parallelism splits the sequence dimension of the transformer across multiple GPUs, allowing a single forward pass to be processed in parallel. For a batch size of 1 and sequence length of 1024, a Ulysses degree of 2 allocates half of the tokens to each GPU, reducing per-layer compute time without changing the model's output.

Enable this feature by setting --ulysses-degree 2 in your launch command. The same README table that defines CFG parallelism notes that this flag "enables Ulysses sequence parallelism, splitting the sequence dimension across GPUs." The degree can be any integer that divides the sequence length, though 2 is the most common configuration for dual-GPU setups.

Calculating Total GPU Requirements

When mixing tensor parallelism, CFG parallelism, and Ulysses sequence parallelism, the total number of GPUs required is the product of the three degrees:


total_gpus = tensor_parallel_size × cfg_parallel_size × ulysses_degree

For example, a configuration with --tensor-parallel-size 2, --cfg-parallel-size 2, and --ulysses-degree 2 requires eight GPUs. The server will fail to start if your hardware cannot satisfy this product, so verify your GPU count before launching.

Implementation Examples

The following command launches a 4-GPU server with tensor parallelism set to 2, CFG parallelism set to 2, and Ulysses degree set to 2:

vllm serve nvidia/Cosmos3-Nano \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --tensor-parallel-size 2 \
    --cfg-parallel-size 2 \
    --ulysses-degree 2 \
    --port 8000 \
    --init-timeout 1800

For a minimal single-GPU deployment without parallelism, omit the multiplicity flags:

vllm serve nvidia/Cosmos3-Nano \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --port 8000

When sending requests to the server, specify the CFG strength using guidance_scale in the JSON payload:

{
  "model": "nvidia/Cosmos3-Nano",
  "prompt": "A futuristic city at sunset",
  "guidance_scale": 7.5,
  "max_new_tokens": 128
}

Key Source Files

  • README.md (root): Contains the authoritative table of parallelism flags and the GPU product calculation rule.
  • cookbooks/cosmos3/README.md: Mirrors the root documentation for the Cosmos 3 cookbook, providing quick reference material.
  • cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb: Demonstrates the --cfg-parallel-size flag in a runnable Jupyter notebook example.

Summary

  • CFG parallelism (--cfg-parallel-size 2) runs positive and negative branches on separate GPUs to halve CFG latency.
  • Ulysses sequence parallelism (--ulysses-degree 2) distributes the sequence dimension across GPUs to accelerate transformer layers.
  • Total GPU count equals the product of tensor-parallel, CFG-parallel, and Ulysses degrees.
  • Control guidance scale via the guidance_scale JSON field, not legacy flags.
  • Reference files include the root README.md and the run_with_vllm_omni.ipynb notebook.

Frequently Asked Questions

What is the minimum GPU requirement for CFG parallelism?

CFG parallelism requires at least 2 GPUs because it splits the positive and negative branches across separate devices. Setting --cfg-parallel-size 2 on a single GPU will cause the server to fail during initialization.

Can I use Ulysses sequence parallelism without tensor parallelism?

Yes, Ulysses sequence parallelism operates independently of tensor parallelism. You can set --ulysses-degree 2 while keeping --tensor-parallel-size 1, which is useful for long-sequence workloads where the sequence dimension is the primary bottleneck rather than the model weights.

How does Ulysses sequence parallelism affect batch processing?

Ulysses splits the sequence dimension across GPUs, so it applies to the tokens within a single sequence. For batch size greater than 1, each sequence in the batch is split across the Ulysses group, meaning all GPUs still participate in processing every batch item. The degree must divide the sequence length evenly to avoid padding inefficiencies.

Where is the guidance scale configured in Cosmos 3?

The guidance scale is configured client-side in the JSON request payload using the guidance_scale field. As implemented in the Cosmos source code, you should not use the older true_cfg_scale server flag, which has been deprecated in favor of per-request control.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →