How to Configure Tensor Parallelism and CFG Parallelism for Cosmos 3 Super Model Inference

Use the --tensor-parallel-size and --cfg-parallel-size flags when launching the vLLM‑Omni server to shard the 64 B Cosmos 3 Super model across GPUs and accelerate classifier‑free guidance by running dual forward passes in parallel.

The NVIDIA Cosmos repository provides a vLLM‑Omni inference engine for the Cosmos 3 Super text‑to‑image/video generation model. Because the Super variant contains 64 billion parameters, single‑GPU inference is impossible; you must enable tensor parallelism to distribute weights across devices and optionally enable CFG parallelism to accelerate the positive‑ and negative‑prompt branches simultaneously.

Understanding Tensor Parallelism for Cosmos 3 Super

Tensor parallelism splits the transformer layers of the Cosmos 3 Super model across multiple GPUs, with each device holding only a slice of the parameters. This is essential for fitting the 64 B model into GPU memory.

According to the repository’s README.md, you activate tensor parallelism by passing the --tensor-parallel-size N argument to the vllm serve command. The value N represents the number of GPUs across which the model weights are sharded via NCCL communication. When N=4, each GPU loads approximately one‑fourth of the model parameters, reducing the per‑GPU memory footprint from impossible levels to roughly 40 GB per slice.

Understanding CFG Parallelism for Cosmos 3 Super

CFG parallelism (Classifier‑Free Guidance parallelism) accelerates inference by executing the positive‑prompt and negative‑prompt forward passes on separate GPU groups simultaneously. Without this optimization, the two passes run sequentially, doubling latency.

To enable CFG parallelism, set --cfg-parallel-size N when launching the server. When N=2, the vLLM‑Omni engine assigns distinct GPU sets to each CFG branch, combines the results on the host, and then proceeds to the diffusion step. This configuration roughly halves the latency of the guidance computation compared to sequential execution.

Combining Parallelism Strategies

When you enable multiple parallelism modes, the total GPU requirement is the product of the individual degrees:


Required GPUs ≥ tensor_parallel_size × cfg_parallel_size × ulysses_degree

For example, a configuration with --tensor-parallel-size 4 and --cfg-parallel-size 2 demands 8 GPUs (4 × 2). You can also add sequence parallelism via --ulysses-degree N (often called Ulysses parallelism) for extremely long sequences, though this is optional for most workloads.

The cookbooks/cosmos3/README.md file in the NVIDIA Cosmos repository provides detailed resource‑size guidance and recommends starting with tensor and CFG parallelism before experimenting with Ulysses degrees.

Launch Commands and Configuration Examples

Basic Tensor Parallelism Only

Deploy the Cosmos 3 Super model across 4 GPUs with pure tensor sharding and no CFG acceleration:

vllm serve nvidia/Cosmos3-Super \
    --omni \
    --tensor-parallel-size 4 \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

Tensor and CFG Parallelism Combined

Run the model on 4 GPUs total, splitting the model weights across 2 GPUs while running the two CFG branches on separate pairs:

vllm serve nvidia/Cosmos3-Super \
    --omni \
    --tensor-parallel-size 2 \
    --cfg-parallel-size 2 \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

Full Parallelism Configuration

For an 8‑GPU node, maximize throughput by combining tensor parallelism (4) with CFG parallelism (2):

vllm serve nvidia/Cosmos3-Super \
    --omni \
    --tensor-parallel-size 4 \
    --cfg-parallel-size 2 \
    --ulysses-degree 1 \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

Python Client Integration

After launching the server, send generation requests using the OpenAI‑compatible client. Specify the guidance_scale to control CFG strength (do not use true_cfg_scale):

import openai

client = openai.OpenAI(base_url="http://localhost:8000/v1")

resp = client.images.generate(
    model="nvidia/Cosmos3-Super",
    prompt="A futuristic cityscape at sunset",
    negative_prompt="low‑resolution, blurry",
    guidance_scale=7.5,
    size="1024x1024",
)

# Access the generated image via resp.data[0].b64_json

Memory Optimization with Layerwise Offload

If GPU memory is still constrained even with tensor parallelism, enable --enable-layerwise-offload to move activations to CPU. This reduces per‑GPU memory usage at the cost of increased latency:

vllm serve nvidia/Cosmos3-Super \
    --omni \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --port 8000 \
    --init-timeout 1800

Resource Requirements and Best Practices

Always verify that your hardware meets the product of the parallelism degrees before launching. The server will abort with a "not enough GPUs" error if the available device count is insufficient.

Memory budgeting requires careful attention. Even with tensor parallelism, each GPU in a Cosmos 3 Super deployment typically requires approximately 40 GB of VRAM per slice. If your GPUs have less memory, use --enable-layerwise-offload to compensate.

CFG parallelism only accelerates the guidance computation phase. The diffusion steps themselves continue to run on the tensor‑parallel group, so the overall speedup follows the formula roughly 1 + 1/cfg_parallel_size. When mixing parallelism strategies, start with tensor and CFG, then experiment with Ulysses degrees only if you have sufficient hardware and need to process extremely long sequences.

Summary

  • Tensor parallelism (--tensor-parallel-size) shards the 64 B Cosmos 3 Super model across GPUs to fit within memory constraints.
  • CFG parallelism (--cfg-parallel-size) runs positive and negative prompt branches simultaneously, reducing guidance latency.
  • GPU requirements multiply: ensure you have at least tensor_parallel_size × cfg_parallel_size × ulysses_degree GPUs available.
  • Memory optimization via --enable-layerwise-offload allows running on smaller GPUs by offloading layers to CPU.
  • Client configuration uses guidance_scale for CFG strength in OpenAI‑compatible API calls.

Frequently Asked Questions

How many GPUs do I need to run Cosmos 3 Super with tensor and CFG parallelism?

You need at least the product of the two parallelism degrees. For example, --tensor-parallel-size 4 combined with --cfg-parallel-size 2 requires 8 GPUs total. The server validates this at startup and will fail if the available GPU count is insufficient.

Can I run Cosmos 3 Super on fewer than 8 GPUs?

Yes, if you only use tensor parallelism. You can run the model on 4 GPUs using --tensor-parallel-size 4 without CFG parallelism, or on 2 GPUs using --tensor-parallel-size 2 and --cfg-parallel-size 1. However, you cannot combine high degrees of both parallelism types without meeting the multiplicative GPU requirement.

What is the difference between guidance_scale and true_cfg_scale in the client?

Use guidance_scale in your API requests to control CFG strength. Do not use true_cfg_scale, as the vLLM‑Omni engine handles the parallel CFG execution internally when --cfg-parallel-size is configured. The guidance_scale parameter determines how strongly the negative prompt influences the final output.

Why should I enable layerwise offload instead of increasing tensor parallelism?

Layerwise offload (--enable-layerwise-offload) is useful when you have limited GPU memory but cannot add more GPUs. Moving layers to CPU reduces VRAM usage below the typical 40 GB per‑GPU requirement, allowing the model to run on smaller devices. Increasing tensor parallelism requires additional GPUs but maintains native GPU speed, whereas offload introduces CPU‑GPU transfer latency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →