How to Enable CFG Parallelism in Cosmos 3: Run Positive and Negative Branches on Separate GPUs

Set --cfg-parallel-size 2 (or higher) when launching the vLLM-Omni server or Cosmos Framework CLI to distribute the classifier-free guidance negative and positive passes across distinct GPUs, cutting generation latency roughly in half without increasing per-device memory consumption.

NVIDIA Cosmos 3 implements classifier-free guidance (CFG) via two separate inference passes: an unconditioned (negative) pass and a conditioned (positive) pass. By default, the Cosmos3OmniDiffusersPipeline executes these passes sequentially on the same GPU, doubling end-to-end latency. Enabling CFG parallelism allows each branch to run simultaneously on dedicated hardware, and this guide details the exact flags, resource formulas, and configuration files required to activate it.

Understanding CFG Parallelism Architecture

CFG parallelism splits the standard guidance computation across multiple workers. In the default sequential mode, the model first computes the negative logits, then the positive logits, and finally aggregates them using the formula logits = negative + guidance_scale * (positive – negative). When --cfg-parallel-size is set to a value greater than 1, the vLLM-Omni server (or Cosmos Framework runtime) spawns independent inference workers—each bound to a specific GPU via CUDA_VISIBLE_DEVICES—and executes both passes in parallel. The first GPU typically handles the negative pass while the second handles the positive pass, though this ordering is configurable. The server then aggregates the results before returning the final output to the client.

Enabling CFG Parallelism via vLLM-Omni

The vLLM-Omni server exposes parallelism controls through dedicated CLI flags. To run the positive and negative branches on separate GPUs, launch the server with --cfg-parallel-size 2 and ensure --tensor-parallel-size is set to 1 (unless combining with tensor parallelism, which requires additional GPUs).


# Launch Cosmos 3-Nano with 2-GPU CFG parallelism

vllm serve nvidia/Cosmos3-Nano \
  --omni \
  --model-class-name Cosmos3OmniDiffusersPipeline \
  --allowed-local-media-path / \
  --cfg-parallel-size 2 \
  --tensor-parallel-size 1 \
  --port 8000 \
  --init-timeout 1800

In this configuration, the server creates two distinct workers: one bound to CUDA_VISIBLE_DEVICES=0 for the negative pass and one to CUDA_VISIBLE_DEVICES=1 for the positive pass. For a complete working example, refer to the notebook at cosmos/cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb.

Launching with the Cosmos Framework CLI

If you are using the native Cosmos Framework instead of vLLM-Omni, the same flag syntax applies. The CLI automatically handles GPU allocation and worker spawning based on the --cfg-parallel-size argument.

cosmos serve \
  --model nvidia/Cosmos3-Nano \
  --cfg-parallel-size 2 \
  --tensor-parallel-size 1 \
  --port 8000

Both methods reference the architecture documentation in cosmos/README.md, which explains how --cfg-parallel-size composes with other parallelism dimensions.

GPU Allocation and Worker Binding

Under the hood, the vLLM-Omni server manages a pool of worker processes equal to the value of --cfg-parallel-size. Each worker initializes its own copy of the model weights and is isolated to a specific GPU index. The default assignment maps the negative (unconditioned) branch to the first available GPU and the positive (conditioned) branch to the second. Because each branch holds a full replica of the model parameters, the per-GPU memory footprint remains identical to a non-parallel deployment; however, the total memory consumption across the node scales linearly with cfg_parallel_size.

Resource Requirements and Sizing Formulas

CFG parallelism multiplies with other distributed strategies. The total number of required GPUs must satisfy the following formula documented in the repository root README.md:


tensor_parallel_size × cfg_parallel_size × ulysses_degree

For example, a configuration with --tensor-parallel-size 2, --cfg-parallel-size 2, and a Ulysses sequence parallelism degree of 1 requires exactly four GPUs. Attempting to run with fewer devices will result in initialization errors or CUDA out-of-memory failures. Because each CFG branch loads a complete model instance, CFG parallelism is most efficient with smaller model variants (e.g., Cosmos3-Nano) or when tensor parallelism is disabled.

Configuring Client Requests

Once the server is running with CFG parallelism enabled, clients interact with it identically to sequential mode. The guidance_scale parameter controls the strength of the classifier-free guidance and is supplied per request; no additional flags are required to utilize the parallel backend.

{
  "model": "nvidia/Cosmos3-Nano",
  "messages": [
    {"role": "system", "content": [{"type":"text","text":"You are a helpful assistant."}]},
    {"role": "user",   "content": [{"type":"text","text":"A red robot is pouring water into a glass."}]}
  ],
  "guidance_scale": 7.5
}

The server routes the request to both workers, aggregates the logits using the standard CFG formula, and returns the guided output without exposing the parallelism mechanism to the client.

Summary

  • Enable parallel branches by setting --cfg-parallel-size 2 (or higher) in either the vLLM-Omni server or Cosmos Framework CLI.
  • Calculate total GPUs using tensor_parallel_size × cfg_parallel_size × ulysses_degree to avoid resource conflicts.
  • Understand worker mapping: The first GPU typically runs the negative pass, the second the positive pass, each holding a full model replica.
  • Maintain client compatibility: Use the standard guidance_scale field in JSON requests; parallelism is transparent to the API consumer.
  • Consult reference files: cosmos/README.md and cosmos/cookbooks/cosmos3/README.md provide the authoritative parallelism formulas, while cosmos/cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb offers a runnable implementation.

Frequently Asked Questions

What is the minimum GPU requirement for CFG parallelism?

You need at least two GPUs to benefit from CFG parallelism, as the feature requires one device for the negative pass and one for the positive pass when --cfg-parallel-size 2 is set. If you also enable tensor parallelism, multiply the base requirement accordingly using the formula tensor_parallel_size × cfg_parallel_size × ulysses_degree.

Does CFG parallelism increase total VRAM consumption?

Yes, total cluster VRAM usage increases linearly with cfg_parallel_size because each worker loads a full copy of the model weights into its assigned GPU. However, the per-GPU memory footprint remains identical to a single-device deployment, meaning you trade total memory capacity for reduced latency.

Can I combine CFG parallelism with tensor parallelism?

Yes, the flags are orthogonal. You can set --tensor-parallel-size 2 and --cfg-parallel-size 2 simultaneously, which shards each CFG branch across multiple GPUs. This configuration requires four GPUs total and is useful for running larger models like Cosmos3-Small or Cosmos3-Large with guidance enabled.

How do I verify that CFG parallelism is active?

Check the server logs during initialization for worker spawning messages that reference distinct CUDA_VISIBLE_DEVICES assignments. Additionally, inspect cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb for validation scripts that assert the negative and positive logits are computed on separate devices before aggregation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →