How to Configure Tensor Parallelism and CFG Parallelism for Cosmos 3 Super Model Inference
Use the --tensor-parallel-size and --cfg-parallel-size flags when launching the vLLM‑Omni server to shard the 64 B Cosmos 3 Super model across GPUs and accelerate classifier‑free guidance by running dual forward passes in parallel.
The NVIDIA Cosmos repository provides a vLLM‑Omni inference engine for the Cosmos 3 Super text‑to‑image/video generation model. Because the Super variant contains 64 billion parameters, single‑GPU inference is impossible; you must enable tensor parallelism to distribute weights across devices and optionally enable CFG parallelism to accelerate the positive‑ and negative‑prompt branches simultaneously.
Understanding Tensor Parallelism for Cosmos 3 Super
Tensor parallelism splits the transformer layers of the Cosmos 3 Super model across multiple GPUs, with each device holding only a slice of the parameters. This is essential for fitting the 64 B model into GPU memory.
According to the repository’s README.md, you activate tensor parallelism by passing the --tensor-parallel-size N argument to the vllm serve command. The value N represents the number of GPUs across which the model weights are sharded via NCCL communication. When N=4, each GPU loads approximately one‑fourth of the model parameters, reducing the per‑GPU memory footprint from impossible levels to roughly 40 GB per slice.
Understanding CFG Parallelism for Cosmos 3 Super
CFG parallelism (Classifier‑Free Guidance parallelism) accelerates inference by executing the positive‑prompt and negative‑prompt forward passes on separate GPU groups simultaneously. Without this optimization, the two passes run sequentially, doubling latency.
To enable CFG parallelism, set --cfg-parallel-size N when launching the server. When N=2, the vLLM‑Omni engine assigns distinct GPU sets to each CFG branch, combines the results on the host, and then proceeds to the diffusion step. This configuration roughly halves the latency of the guidance computation compared to sequential execution.
Combining Parallelism Strategies
When you enable multiple parallelism modes, the total GPU requirement is the product of the individual degrees:
Required GPUs ≥ tensor_parallel_size × cfg_parallel_size × ulysses_degree
For example, a configuration with --tensor-parallel-size 4 and --cfg-parallel-size 2 demands 8 GPUs (4 × 2). You can also add sequence parallelism via --ulysses-degree N (often called Ulysses parallelism) for extremely long sequences, though this is optional for most workloads.
The cookbooks/cosmos3/README.md file in the NVIDIA Cosmos repository provides detailed resource‑size guidance and recommends starting with tensor and CFG parallelism before experimenting with Ulysses degrees.
Launch Commands and Configuration Examples
Basic Tensor Parallelism Only
Deploy the Cosmos 3 Super model across 4 GPUs with pure tensor sharding and no CFG acceleration:
vllm serve nvidia/Cosmos3-Super \
--omni \
--tensor-parallel-size 4 \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--port 8000 \
--init-timeout 1800
Tensor and CFG Parallelism Combined
Run the model on 4 GPUs total, splitting the model weights across 2 GPUs while running the two CFG branches on separate pairs:
vllm serve nvidia/Cosmos3-Super \
--omni \
--tensor-parallel-size 2 \
--cfg-parallel-size 2 \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--port 8000 \
--init-timeout 1800
Full Parallelism Configuration
For an 8‑GPU node, maximize throughput by combining tensor parallelism (4) with CFG parallelism (2):
vllm serve nvidia/Cosmos3-Super \
--omni \
--tensor-parallel-size 4 \
--cfg-parallel-size 2 \
--ulysses-degree 1 \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--port 8000 \
--init-timeout 1800
Python Client Integration
After launching the server, send generation requests using the OpenAI‑compatible client. Specify the guidance_scale to control CFG strength (do not use true_cfg_scale):
import openai
client = openai.OpenAI(base_url="http://localhost:8000/v1")
resp = client.images.generate(
model="nvidia/Cosmos3-Super",
prompt="A futuristic cityscape at sunset",
negative_prompt="low‑resolution, blurry",
guidance_scale=7.5,
size="1024x1024",
)
# Access the generated image via resp.data[0].b64_json
Memory Optimization with Layerwise Offload
If GPU memory is still constrained even with tensor parallelism, enable --enable-layerwise-offload to move activations to CPU. This reduces per‑GPU memory usage at the cost of increased latency:
vllm serve nvidia/Cosmos3-Super \
--omni \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--model-class-name Cosmos3OmniDiffusersPipeline \
--allowed-local-media-path / \
--port 8000 \
--init-timeout 1800
Resource Requirements and Best Practices
Always verify that your hardware meets the product of the parallelism degrees before launching. The server will abort with a "not enough GPUs" error if the available device count is insufficient.
Memory budgeting requires careful attention. Even with tensor parallelism, each GPU in a Cosmos 3 Super deployment typically requires approximately 40 GB of VRAM per slice. If your GPUs have less memory, use --enable-layerwise-offload to compensate.
CFG parallelism only accelerates the guidance computation phase. The diffusion steps themselves continue to run on the tensor‑parallel group, so the overall speedup follows the formula roughly 1 + 1/cfg_parallel_size. When mixing parallelism strategies, start with tensor and CFG, then experiment with Ulysses degrees only if you have sufficient hardware and need to process extremely long sequences.
Summary
- Tensor parallelism (
--tensor-parallel-size) shards the 64 B Cosmos 3 Super model across GPUs to fit within memory constraints. - CFG parallelism (
--cfg-parallel-size) runs positive and negative prompt branches simultaneously, reducing guidance latency. - GPU requirements multiply: ensure you have at least
tensor_parallel_size × cfg_parallel_size × ulysses_degreeGPUs available. - Memory optimization via
--enable-layerwise-offloadallows running on smaller GPUs by offloading layers to CPU. - Client configuration uses
guidance_scalefor CFG strength in OpenAI‑compatible API calls.
Frequently Asked Questions
How many GPUs do I need to run Cosmos 3 Super with tensor and CFG parallelism?
You need at least the product of the two parallelism degrees. For example, --tensor-parallel-size 4 combined with --cfg-parallel-size 2 requires 8 GPUs total. The server validates this at startup and will fail if the available GPU count is insufficient.
Can I run Cosmos 3 Super on fewer than 8 GPUs?
Yes, if you only use tensor parallelism. You can run the model on 4 GPUs using --tensor-parallel-size 4 without CFG parallelism, or on 2 GPUs using --tensor-parallel-size 2 and --cfg-parallel-size 1. However, you cannot combine high degrees of both parallelism types without meeting the multiplicative GPU requirement.
What is the difference between guidance_scale and true_cfg_scale in the client?
Use guidance_scale in your API requests to control CFG strength. Do not use true_cfg_scale, as the vLLM‑Omni engine handles the parallel CFG execution internally when --cfg-parallel-size is configured. The guidance_scale parameter determines how strongly the negative prompt influences the final output.
Why should I enable layerwise offload instead of increasing tensor parallelism?
Layerwise offload (--enable-layerwise-offload) is useful when you have limited GPU memory but cannot add more GPUs. Moving layers to CPU reduces VRAM usage below the typical 40 GB per‑GPU requirement, allowing the model to run on smaller devices. Increasing tensor parallelism requires additional GPUs but maintains native GPU speed, whereas offload introduces CPU‑GPU transfer latency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →