How to Set Up Multi-GPU Inference with Tensor-Parallel-Size, CFG-Parallel-Size, and Ulysses-Degree in Cosmos 3

To run NVIDIA Cosmos 3 on multiple GPUs, multiply --tensor-parallel-size, --cfg-parallel-size, and --ulysses-degree to determine the total GPU count required, then pass these flags to vllm serve.

NVIDIA Cosmos 3 relies on vLLM-Omni to serve diffusion transformer models across multiple GPUs. The framework supports three independent axes of parallelism—tensor parallelism, classifier-free guidance (CFG) parallelism, and Ulysses sequence parallelism—that can be combined to scale inference across large GPU clusters. Understanding how these flags interact is essential for maximizing throughput while avoiding resource allocation errors.

Understanding the Three Parallelism Axes

Cosmos 3 implements three distinct parallelism strategies that address different memory and compute bottlenecks. Each axis is controlled by a specific command-line flag passed to the vLLM-Omni server.

Tensor Parallelism (--tensor-parallel-size)

The --tensor-parallel-size parameter splits the model’s weight tensors across N GPUs. This is the standard data-parallel approach for transformer layers, sharding the model weights so that large models (such as Cosmos-3 Super) can fit in GPU memory even when they exceed the capacity of a single device.

CFG Parallelism (--cfg-parallel-size)

The --cfg-parallel-size parameter executes the positive and negative classifier-free guidance branches in parallel across M GPUs. Instead of running the conditional and unconditional forward passes sequentially, this flag assigns separate GPUs to each branch, reducing the latency of CFG inference when spare GPUs are available.

Ulysses Sequence Parallelism (--ulysses-degree)

The --ulysses-degree parameter enables Ulysses sequence parallelism, cutting the sequence dimension across K GPUs. This strategy is critical for very long prompts or video sequences where the token length would otherwise exceed GPU memory limits, allowing you to process sequences that are too long for standard tensor parallelism alone.

The Product Rule for GPU Allocation

When you enable more than one parallelism axis, the total number of GPUs required equals the product of all enabled degrees. According to the source code in README.md (lines 14-21), the calculation is:


tensor_parallel_size × cfg_parallel_size × ulysses_degree = total_gpus_required

If the product exceeds the number of GPUs exposed to the server, vLLM-Omni will abort with a resource-allocation error. For example, setting --tensor-parallel-size 4, --cfg-parallel-size 2, and --ulysses-degree 2 requires 16 GPUs total (4 × 2 × 2).

Configuration Examples

The following examples demonstrate how to combine these flags in practice. Each configuration assumes the use of Cosmos3OmniDiffusersPipeline as the model class, which is the standard pipeline for Cosmos 3 inference.

Basic Weight Sharding Only

Use this configuration when you need to shard a large model across GPUs but do not require CFG or sequence parallelism:

vllm serve nvidia/Cosmos3-Nano \
  --tensor-parallel-size 4 \
  --model-class-name Cosmos3OmniDiffusersPipeline \
  --port 8000

Requires 4 GPUs.

Adding CFG Parallelism

This configuration adds CFG parallelism to the weight sharding, running the positive and negative guidance branches on separate GPU sets:

vllm serve nvidia/Cosmos3-Nano \
  --tensor-parallel-size 4 \
  --cfg-parallel-size 2 \
  --model-class-name Cosmos3OmniDiffusersPipeline \
  --port 8000

Requires 8 GPUs (4 × 2).

Full Three-Axis Configuration

For maximum parallelism with large sequences, combine all three axes. This example uses smaller individual degrees to fit on an 8-GPU server:

vllm serve nvidia/Cosmos3-Nano \
  --tensor-parallel-size 2 \
  --cfg-parallel-size 2 \
  --ulysses-degree 2 \
  --model-class-name Cosmos3OmniDiffusersPipeline \
  --port 8000

Requires 8 GPUs (2 × 2 × 2). Adjust the numbers so the product matches your available hardware.

Docker Deployment

The same flags work with the official Docker wrapper. Ensure the host has enough GPUs for your chosen configuration:

docker run --gpus all -p 8000:8000 \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Nano \
    --tensor-parallel-size 4 \
    --cfg-parallel-size 2 \
    --ulysses-degree 2 \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --port 8000

Verifying Your Multi-GPU Setup

After launching the server, verify the configuration by checking the health endpoint. A successful multi-GPU launch will log:


Application startup complete.

Query the server to confirm it is serving requests:

curl http://localhost:8000/v1/health

A JSON response containing "status":"healthy" confirms that the multi-GPU configuration is active and the parallelism groups are correctly initialized.

Summary

  • Multi-GPU inference in Cosmos 3 requires passing --tensor-parallel-size, --cfg-parallel-size, and --ulysses-degree to the vllm serve command.
  • Total GPU requirement is the product of all enabled parallelism degrees: tensor_parallel_size × cfg_parallel_size × ulysses_degree.
  • Tensor parallelism shards model weights, CFG parallelism splits guidance branches, and Ulysses parallelism distributes sequence lengths across GPUs.
  • Configuration validation occurs in README.md (lines 14-21) and cookbooks/cosmos3/README.md (lines 354-357), which document the product rule and provide command-line examples.

Frequently Asked Questions

What happens if the product of parallelism degrees exceeds available GPUs?

vLLM-Omni will abort during initialization with a resource-allocation error. The server checks that the product of --tensor-parallel-size, --cfg-parallel-size, and --ulysses-degree does not exceed the number of GPUs exposed to the container or process, as enforced by the validation logic in the top-level README.md.

Can I use Ulysses parallelism without tensor parallelism?

Yes, the three axes are independent. You can set --tensor-parallel-size 1 while using --ulysses-degree 4 to distribute a long sequence across GPUs without sharding the model weights. This is useful when the model fits on a single GPU but the input sequence length exceeds memory capacity.

Where are the parallelism flags documented in the source code?

The primary documentation appears in the top-level README.md (lines 14-21), which includes the parallelism table and the product rule. Additional examples and implementation details are provided in cookbooks/cosmos3/README.md (lines 354-357), and practical usage patterns appear in the generator notebooks under cookbooks/cosmos3/generator/.

How do I know which parallelism strategy to use for my workload?

Use tensor parallelism when the model weights exceed single-GPU memory. Add CFG parallelism when you have spare GPUs and need lower latency for classifier-free guidance sampling. Enable Ulysses parallelism when processing video sequences or long prompts that exceed the context window capacity of individual GPUs, regardless of model size.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →