Configuring Tensor Parallelism for Cosmos 3 Super on Multi-GPU Setups

Tensor parallelism shards the 64B parameter Cosmos 3 Super model across multiple GPUs using the --tensor-parallel-size flag, reducing per-GPU memory consumption by dividing transformer weights and synchronizing activations via All-Reduce.

Cosmos 3 Super is the 64 billion parameter variant of NVIDIA's Cosmos 3 family, requiring distributed inference across multiple GPUs to fit in memory. According to the NVIDIA/cosmos repository, enabling tensor parallelism is mandatory for deploying this checkpoint, as the model weights exceed the capacity of any single consumer or datacenter GPU. This guide explains how to configure tensor parallelism for multi-GPU setups using both native vLLM-Omni and Docker-based deployments.

What Tensor Parallelism Does for Cosmos 3 Super

Tensor parallelism splits transformer weights—including QKV projections, linear layers, and embeddings—across selected GPUs so each device holds only a fraction of the model. During forward passes, each GPU computes its slice of the tensor and collaborates via All-Reduce operations to produce final activations. This architecture reduces per-GPU memory consumption roughly by a factor of 1 / tensor_parallel_size.

The Cosmos 3 Super implementation in cookbooks/cosmos3/README.md demonstrates that each GPU stores only a partitioned subset of the 64B parameters, enabling inference on hardware that would otherwise be insufficient.

When to Enable Tensor Parallelism

Any deployment running the Cosmos 3 Super checkpoint on more than one GPU must enable tensor parallelism. The reference configuration specifically uses four GPUs, setting --tensor-parallel-size 4 to distribute the model adequately. Without this configuration, the vLLM server will exhaust available memory during model loading.

As documented in the repository's README.md (lines 312-321), you should also add --enable-layerwise-offload when running the Super model to further reduce peak GPU memory by moving transformer blocks between CPU and GPU during computation.

Configuring Tensor Parallelism in vLLM-Omni

Native vLLM-Omni (Non-Docker) Setup

For bare-metal deployments with four GPUs, export the necessary environment variables and launch the server with tensor parallelism enabled. The command references specific model architecture overrides required for the Cosmos 3 Super checkpoint:


# From the repository root

export COSMOS3_MEDIA_ROOT="$(pwd)/cookbooks/cosmos3"
export VLLM_PORT="${VLLM_PORT:-8001}"

CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve nvidia/Cosmos3-Super \
  --hf-overrides '{"architectures": ["Cosmos3ReasonerForConditionalGeneration"]}' \
  --tensor-parallel-size 4 \
  --mm-encoder-tp-mode data \
  --async-scheduling \
  --allowed-local-media-path "$COSMOS3_MEDIA_ROOT" \
  --media-io-kwargs '{"video": {"num_frames": -1}}' \
  --port "$VLLM_PORT" \
  --enable-layerwise-offload \
  --init-timeout 1800

The --tensor-parallel-size 4 parameter instructs vLLM to shard the model across four GPUs, while --enable-layerwise-offload activates CPU offloading for transformer blocks.

Docker-Based Configuration

For containerized deployments, mount the necessary volumes and pass the same tensor parallelism arguments to the vLLM-Omni image:

docker run --runtime nvidia --gpus all \
  -v "${HF_HOME:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
  -v "$(pwd):/workspace" \
  -p "${COSMOS3_HOST_PORT:-8000}:8000" --ipc=host \
  vllm/vllm-omni:cosmos3 \
  vllm serve nvidia/Cosmos3-Super \
    --omni \
    --model-class-name Cosmos3OmniDiffusersPipeline \
    --allowed-local-media-path / \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --port 8000 \
    --init-timeout 1800

The --ipc=host flag is essential for allowing NCCL communication between containers when utilizing multiple GPUs.

Combining Tensor Parallelism with Other Strategies

Tensor parallelism can operate alongside CFG-parallel (classifier-free guidance parallel) and Ulysses sequence parallelism as long as the total product of parallel degrees does not exceed your available GPU count. As documented in README.md (lines 49-57), the total GPU requirement equals tensor_parallel_size × cfg_parallel_size × ulysses_degree.

To add supplementary parallelism dimensions, append these flags to your launch command:

--cfg-parallel-size 2 \
--ulysses-degree 2 \

For example, combining --tensor-parallel-size 4 with --cfg-parallel-size 2 requires eight total GPUs. The inference_benchmarks.md file contains benchmark tables explicitly validating 4-GPU tensor-parallel configurations for Cosmos 3 Super.

Memory Optimization Through Layer-Wise Offload

While tensor parallelism distributes weights horizontally across GPUs, layer-wise offload (--enable-layerwise-offload) moves transformer blocks vertically between CPU DRAM and GPU VRAM. This technique is particularly critical for Cosmos 3 Super because it reduces peak memory usage during inference by keeping only active layers on the GPU.

The NVIDIA/cosmos documentation emphasizes that layer-wise offload works synergistically with tensor parallelism, allowing the 64B model to run on commodity hardware configurations that would otherwise require significantly more expensive GPU setups.

Summary

  • Tensor parallelism is mandatory for Cosmos 3 Super (64B parameters) and requires the --tensor-parallel-size flag set to the number of GPUs available.
  • Memory reduction scales linearly with the tensor parallel size, roughly dividing per-GPU consumption by the number of participating devices.
  • Reference configuration uses 4 GPUs as shown in cookbooks/cosmos3/README.md and inference_benchmarks.md.
  • Layer-wise offload (--enable-layerwise-offload) should always accompany tensor parallelism for Super deployments to minimize peak GPU memory.
  • Parallelism dimensions multiply: ensure tensor_parallel_size × cfg_parallel_size × ulysses_degree does not exceed your physical GPU count.

Frequently Asked Questions

How many GPUs do I need for Cosmos 3 Super inference?

The reference implementation recommends four GPUs for tensor parallelism when running the Cosmos 3 Super checkpoint. While you could theoretically use fewer with aggressive offloading, the official cookbooks and benchmarks specifically validate the 4-GPU configuration as the optimal balance between throughput and memory efficiency.

Can I combine tensor parallelism with other parallelism methods?

Yes. According to the repository's README.md, you can combine tensor parallelism with CFG-parallel (--cfg-parallel-size) and Ulysses sequence parallelism (--ulysses-degree) provided that the product of all degrees does not exceed your available GPU count. For example, tensor parallelism size 4 combined with CFG-parallel size 2 requires eight total GPUs.

What is the difference between tensor parallelism and layer-wise offload?

Tensor parallelism horizontally shards model weights across GPUs during inference, with each GPU computing partial results that are synchronized via All-Reduce. Layer-wise offload vertically moves entire transformer blocks between CPU and GPU memory, keeping only the currently executing layer resident in VRAM. For Cosmos 3 Super, you should enable both: tensor parallelism to distribute the 64B parameters, and layer-wise offload to further reduce peak memory consumption per device.

Where are the configuration examples located in the repository?

The definitive configuration examples reside in three key files: cookbooks/cosmos3/README.md contains detailed launch commands for each backend including the exact --tensor-parallel-size flag usage; README.md (lines 312-321) explains the high-level interaction between tensor parallelism and layer-wise offload; and inference_benchmarks.md provides validated performance tables for 4-GPU tensor-parallel setups.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →