How to Handle Out-of-Memory Errors When Running Large Models Like Cosmos3-Super
Use tensor parallelism across multiple GPUs, enable layer-wise CPU offloading, and extend the initialization timeout to 1800 seconds to successfully serve the 64-billion-parameter Cosmos3-Super model without OOM crashes.
The NVIDIA/cosmos repository hosts Cosmos3-Super, an omnimodal transformer that presents unique memory challenges due to its 64B parameter count and 30GB checkpoint size. When loaded on a single GPU, the model’s weight matrices and activation caches quickly exceed the 24GB capacity of typical devices. The repository implements three complementary mechanisms—tensor parallelism, layer-wise offload, and extended init timeouts—to handle out-of-memory errors when running large models like Cosmos3-Super.
Why Cosmos3-Super Causes OOM Errors
Cosmos3-Super stores its transformer weights in large contiguous matrices such as W_qkv. In a standard single-GPI load, this matrix alone consumes approximately 35GB, which immediately saturates consumer GPU memory before accounting for activation caches. The model architecture requires sharding these parameters across devices or temporarily moving them to host RAM to remain within hardware limits.
Three Memory Management Strategies
Tensor-Parallel Inference
Tensor parallelism splits the model’s weight matrix across N GPUs using the --tensor-parallel-size flag. Each GPU holds only a slice of the parameters (e.g., splitting 35GB across 4 GPUs yields ~8.75GB per device), with collective communication reconstructing the output via all-reduce operations. According to the README.md at line 312, this is the primary mechanism for fitting Cosmos3-Super into GPU memory.
Layer-Wise Offloading
Layer-wise offload moves entire transformer blocks between GPU and CPU on-the-fly, keeping only the active block resident on the device. This trades latency for a substantial reduction in peak GPU memory, requiring approximately 2× GPU memory capacity in available CPU RAM. Enable this with the --enable-layerwise-offload flag as documented in the cookbooks at cookbooks/cosmos3/README.md line 252.
Extended Initialization Timeout
Extended init timeout allows vLLM up to 30 minutes (1800 seconds) to perform checkpoint sharding before marking the server ready. This is required because the Cosmos3-Super checkpoint exceeds 30GB on disk, and the sharding step can take over 10 minutes. Set this with --init-timeout 1800.
Launch Configurations for Different Environments
Docker and vLLM-Omni (Recommended)
The quickest method uses the pre-built vLLM-Omni image with all three memory optimization flags:
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-v "$(pwd):/workspace" \
-p 8001:8001 \
--ipc=host \
vllm/vllm-omni:cosmos3 \
vllm serve nvidia/Cosmos3-Super \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--init-timeout 1800 \
--port 8001
This configuration splits the model across four GPUs, offloads inactive layers to CPU, and allows 30 minutes for initialization.
Native Python Installation
For virtual environment deployments, install the vLLM-Cosmos3 package and apply the same flags:
# Create a fresh venv with the matching CUDA build (e.g., cu130)
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
uv pip install --torch-backend=cu130 \
"vllm==0.21.0" \
"vllm-cosmos3 @ git+https://github.com/NVIDIA/cosmos-framework.git#subdirectory=packages/vllm-cosmos3"
# Launch the server
vllm serve nvidia/Cosmos3-Super \
--omni \
--model-class-name Cosmos3OmniDiffusersPipeline \
--tensor-parallel-size 4 \
--enable-layerwise-offload \
--init-timeout 1800 \
--port 8001
The cookbooks/cosmos3/reasoner/README.md at lines 71-73 confirms these flags are required for both Generator and Reasoner pipelines.
Single GPU with Precision Reduction
If limited to one GPU, reduce memory usage by running at BF16 precision (halving activation memory versus FP32), limiting resolution to 256p instead of 720p, and shortening video length. The cookbooks/cosmos3/generator/audiovisual/README.md at lines 34-42 demonstrates this approach:
pipe = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Super",
torch_dtype=torch.bfloat16,
device_map="cuda"
)
result = pipe(
prompt="A futuristic city skyline at dusk.",
num_frames=100, # smaller video
height=256, width=448, # lower resolution
num_inference_steps=30,
guidance_scale=6.0,
)
Debugging OOM Errors
If you encounter CUDA out-of-memory errors after applying the above flags:
- Verify GPU allocation with
nvidia-smibefore and after launch to confirm memory distribution. - Inspect vLLM logs for sharding timeouts—these indicate you need to increase
--init-timeout. - Reduce batch complexity by lowering
num_inference_stepsor reducing frame counts further. - Confirm CPU RAM availability for layer-wise offload; the system needs approximately 2× the GPU memory capacity in host RAM.
The inference_benchmarks.md at line 136 indicates that Cosmos3-Super officially requires four GPUs with tensor parallelism for standard operation.
Summary
- Tensor parallelism is mandatory for Cosmos3-Super, requiring
--tensor-parallel-size 4to split the 35GB weight matrix across multiple GPUs. - Layer-wise offload (
--enable-layerwise-offload) moves inactive transformer blocks to CPU, reducing peak GPU memory at the cost of inference speed. - Extended initialization timeout (
--init-timeout 1800) prevents premature timeout errors during the 30GB checkpoint loading process. - BF16 precision and reduced resolution (256p) enable limited inference on single-GPU setups when parallelism is unavailable.
Frequently Asked Questions
Can I run Cosmos3-Super on a single 24GB GPU?
No, the model's weight matrices alone require approximately 35GB of memory, exceeding typical consumer GPU capacity. You must use either tensor parallelism across multiple GPUs or aggressive layer-wise offloading with substantial CPU RAM (64GB+ recommended).
What does layer-wise offloading actually do?
Layer-wise offloading moves entire transformer blocks from GPU to host RAM when they are not actively processing, keeping only the current layer resident on the device. This reduces peak GPU memory to roughly the size of one layer plus activations, though it increases inference latency due to PCIe transfer overhead.
Why is the initialization timeout set to 1800 seconds?
The checkpoint file exceeds 30GB on disk, and vLLM requires significant time to shard and distribute these weights across GPUs in a tensor-parallel configuration. The default timeout is insufficient for this operation; 1800 seconds (30 minutes) ensures the server does not mark itself ready before sharding completes.
Does tensor parallelism require specific hardware interconnects?
While tensor parallelism works across any CUDA-capable GPUs, performance degrades significantly without high-bandwidth interconnects like NVLink or PCIe 4.0/5.0 x16. For optimal throughput with Cosmos3-Super, use GPUs with NVLink bridges or on the same NUMA node to minimize all-reduce communication overhead during matrix recombination.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →