# How to Optimize Memory Usage for Cosmos 3 Super on Multiple GPUs with Layer Offloading

> Optimize Cosmos 3 Super GPU memory with layer offloading and tensor parallelism. Stream transformer blocks from CPU RAM to reduce per-GPU needs from 90GB.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: performance
- Published: 2026-06-14

---

**Use tensor parallelism combined with layer-wise offloading to distribute the 64B parameter model across multiple GPUs while streaming transformer blocks from CPU RAM, reducing per-GPU memory from 90GB to manageable slices.**

NVIDIA Cosmos 3 Super is a 64-billion parameter video generation model that requires approximately 90GB of memory for its weights alone. To optimize memory usage for Cosmos 3 Super on multiple GPUs with layer offloading, you must combine tensor parallelism with CPU-GPU layer streaming, as implemented in the NVIDIA Cosmos repository and supported through the vLLM-Omni integration.

## Understanding the Memory Challenge

Cosmos 3 Super represents the largest variant in the Cosmos 3 family, with weight tensors occupying roughly 90GB of storage. A single GPU cannot hold the full checkpoint in memory, necessitating distributed inference strategies. According to the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (line 312), the model relies on two complementary techniques to achieve feasible inference: tensor parallelism for horizontal weight distribution and layer-wise offloading for vertical memory management.

## Tensor Parallelism for Distributed Processing

**Tensor parallelism** splits the model's weight tensors evenly across multiple GPUs, reducing the per-device memory footprint proportionally.

When you specify `--tensor-parallel-size 4`, the framework divides each weight matrix across four GPUs, cutting per-GPU memory usage by approximately 75%. As noted in [`cookbooks/cosmos3/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/README.md) (line 291), this approach requires inter-GPU communication during forward passes, which introduces latency but enables the model to run on hardware that would otherwise be insufficient.

The trade-off involves increased communication overhead between GPUs during tensor operations, though this is typically acceptable for throughput gains.

## Layer-Wise Offloading Architecture

**Layer-wise offloading** complements tensor parallelism by changing the residency pattern of transformer blocks. Instead of keeping all layers in GPU memory simultaneously, the system streams individual transformer blocks from CPU RAM to GPU only when actively computing.

As documented in [`cookbooks/cosmos3/generator/audiovisual/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/README.md) (line 66), this technique maintains only the currently executing layer (plus a small activation cache) on the GPU, while storing the remaining layers in host memory. This requires approximately 30–40GB of free CPU RAM to serve as the offload buffer for the 64B model.

The scheduler in vLLM manages these transfers on-the-fly, dramatically shrinking the maximum GPU footprint while utilizing CPU memory as a backing store.

## Implementation Guide

### Multi-GPU Launch with vLLM

The recommended configuration combines both flags with an extended initialization timeout to accommodate the offloading setup process:

```bash
vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --enable-layerwise-offload \
    --init-timeout 1800 \
    --port 8001

```

Set `CUDA_VISIBLE_DEVICES=0,1,2,3` before execution to specify which GPUs participate in the tensor parallel group.

### Docker Deployment

For containerized deployments using the vLLM-Omni image:

```bash
docker run --gpus all -p 8001:8001 \
    -e CUDA_VISIBLE_DEVICES=0,1,2,3 \
    vllm/vllm-omni:cosmos3 \
    vllm serve nvidia/Cosmos3-Super \
        --tensor-parallel-size 4 \
        --enable-layerwise-offload \
        --init-timeout 1800

```

### Python API Integration

When using the Python API with `Cosmos3OmniPipeline.from_pretrained`, the same flags apply:

```python
from vllm import LLM, SamplingParams

llm = LLM(
    model="nvidia/Cosmos3-Super",
    tensor_parallel_size=4,
    enable_layerwise_offload=True,
    gpu_memory_utilization=0.85,
)

sampling_params = SamplingParams(temperature=0.7, max_new_tokens=256)
prompt = "A futuristic cityscape at sunset."
outputs = llm.generate([prompt], sampling_params)

print(outputs[0].outputs[0].text)

```

### Two-GPU Configuration

If you have only two GPUs, increase the reliance on offloading:

```bash
vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 2 \
    --enable-layerwise-offload \
    --init-timeout 1800 \
    --port 8002

```

With fewer GPUs, each device holds a larger slice of the model, making the layer-wise offloading flag even more critical for memory management.

## Configuration Requirements

### CPU Memory Allocation

Ensure the host system has sufficient free RAM. The offload buffer requires approximately 30–40GB for the 64B parameter model, depending on activation caching and overhead.

### Batch Size Constraints

Keep **batch size = 1** for maximal memory savings. Larger batches force the system to keep multiple layers active simultaneously, increasing GPU memory usage and reducing the effectiveness of offloading.

### Initialization Timeout

The `--init-timeout 1800` parameter (1800 seconds) accounts for the several minutes required to initialize the layer offloading buffers and distribute weights across the tensor parallel group. Without this extension, the system may timeout during startup.

## Summary

- **Tensor parallelism** (`--tensor-parallel-size`) distributes weight tensors across multiple GPUs, reducing per-GPU memory by roughly 1/N where N is the GPU count.
- **Layer-wise offloading** (`--enable-layerwise-offload`) streams transformer blocks from CPU RAM to GPU on-demand, requiring ~30-40GB of host memory but minimizing GPU residency.
- The combination enables Cosmos 3 Super inference on commodity multi-GPU setups by reducing the 90GB weight requirement to manageable slices per device.
- Always set `--init-timeout 1800` to accommodate offloading initialization overhead.
- Maintain batch size of 1 to maximize memory savings and minimize active layer count.

## Frequently Asked Questions

### What is the minimum GPU memory required per GPU when using layer offloading?

With layer offloading enabled, each GPU only needs to hold the active transformer layer plus activation caches rather than the full model. While the exact requirement depends on the layer size and precision, the technique reduces GPU memory requirements from 90GB to approximately 20-25GB per GPU when combined with tensor parallelism across four devices, though specific hardware requirements vary by configuration.

### Can I use layer offloading with only 2 GPUs?

Yes, layer offloading becomes even more critical with fewer GPUs. When using `--tensor-parallel-size 2`, each GPU must store a larger slice of the model weights, making the `--enable-layerwise-offload` flag essential to keep memory usage within device limits. Ensure your CPU has sufficient RAM (30-40GB) to serve as the offload buffer.

### How does the `--init-timeout` parameter affect startup?

The `--init-timeout 1800` flag extends the initialization window from the default to 1800 seconds (30 minutes). Layer offloading requires several minutes to allocate CPU buffers and distribute weights across the tensor parallel group. Without this extension, vLLM may terminate with a timeout error before the Cosmos 3 Super model finishes loading into the distributed memory architecture.

### Does batch size impact memory usage with layer offloading?

Yes, batch size directly impacts memory consumption. Keep **batch size = 1** for maximum memory savings. Larger batches require keeping multiple transformer layers active simultaneously on the GPU, which increases memory usage and reduces the effectiveness of the offloading strategy. For the 64B Super model, batch size 1 is recommended to maintain optimal memory efficiency.