Best Practices for Fine-Tuning Cosmos 3 Models Using the Cosmos Framework
Fine-tuning Cosmos 3 models requires starting from official NVIDIA checkpoints, configuring BF16 mixed precision with gradient accumulation, and following the Cosmos Framework's end-to-end workflow from data curation through distributed training to safety validation.
The NVIDIA Cosmos repository implements a unified Mixture-of-Transformers architecture designed for multimodal video, image, and audio understanding. When adapting these models to new domains, the Cosmos Framework provides the complete infrastructure for fine-tuning, including specialized tools for data preparation, distributed training orchestration, and post-training evaluation. Following the official fine-tuning practices ensures that your adapted model maintains the safety guardrails and physical plausibility inherent in the base Cosmos 3 knowledge.
Core Fine-Tuning Principles
Initialize From Official Checkpoints
Always begin fine-tuning from the official Cosmos 3 checkpoints to inherit the full multimodal knowledge base. The repository provides two primary model sizes: nvidia/Cosmos3-Nano for resource-constrained environments and nvidia/Cosmos3-Super (64B parameters) for maximum fidelity. According to the repository's README.md, these checkpoints include vision, video, audio, and action understanding weights that are essential for downstream tasks.
Configure Mixed Precision and Memory
Use BF16 mixed precision exclusively for fine-tuning on Ampere-generation GPUs and newer. This configuration reduces memory footprint while preserving training stability. Set --mixed_precision bf16 or export TORCH_PRECISION=bf16 to enable this mode. For limited GPU memory, implement gradient accumulation by setting --gradient_accumulation_steps N where the product of N and your per-device batch size equals your target global batch size.
Implement Distributed Training Strategies
Cosmos 3-Super requires multi-GPU or multi-node scaling due to its 64B parameter count. The framework supports distributed training via torchrun and DeepSpeed integration. As documented in cosmos_framework/scripts/train.py, the entry point handles both single-device and distributed configurations through the same YAML configuration interface.
Step-by-Step Fine-Tuning Workflow
Curate Multimodal Training Data
Convert raw videos, images, audio, and action logs into the Cosmos Curator format before training. This process creates a *.jsonl manifest paired with media files, handling de-duplication, modality alignment, and optional annotation. The Cosmos Curator submodule ensures your dataset conforms to the expected schema for the Cosmos3OmniModel training loop.
Define Training Configuration
Create a YAML configuration file (e.g., finetune_config.yaml) specifying model architecture, optimizer settings, and learning rate schedules. The configuration must reference the checkpoint directory, set the diffusion steps, and configure modality-specific tokenizers. The Cosmos Framework documentation at docs/training.md provides starter templates that include sensible defaults for learning rate schedules and warmup steps.
Launch Distributed Training
Execute training through the cosmos_framework.scripts.train module. For single-GPU training with Cosmos 3-Nano, run the module directly with Python. For multi-GPU training with Cosmos 3-Super, use torchrun with the --nnodes and --nproc_per_node arguments to scale across devices. The framework automatically handles checkpoint sharding and saves model states every save_checkpoint_steps to the specified output directory.
Validate and Deploy
After training, run the Cosmos-Evaluator on a held-out validation set to verify that safety guardrails and physical plausibility remain intact. The evaluator checks caption accuracy, spatial grounding, and temporal localization. Once validated, export checkpoints to Hugging Face using the --push_to_hub flag for integration with Diffusers, vLLM-Omni, or NVIDIA NIM for containerized inference.
Production Code Examples
Environment Setup and Installation
Install the Cosmos Framework with CUDA backend auto-detection to avoid torch.cuda.is_available() failures caused by mismatched torch wheels.
# Create Python 3.13 environment
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
# Install with automatic CUDA backend detection
uv pip install --torch-backend=auto \
"cosmos-framework @ git+https://github.com/NVIDIA/cosmos-framework.git" \
deepspeed \
tensorboard \
wandb
Single-GPU Fine-Tuning
Create the configuration file for Nano model fine-tuning:
# finetune_config.yaml
model:
name: "nvidia/Cosmos3-Nano"
dtype: "bfloat16"
datasets:
train: "./data/train_manifest.jsonl"
validation: "./data/val_manifest.jsonl"
training:
epochs: 3
batch_size: 2
gradient_accumulation_steps: 8
learning_rate: 1e-5
lr_schedule: "cosine"
max_steps: 100_000
mixed_precision: "bf16"
logging:
tensorboard_dir: "./logs"
report_to: "wandb"
wandb_project: "cosmos3-finetune"
output:
checkpoint_dir: "./outputs/checkpoints"
push_to_hub: true
Launch single-GPU training:
HF_TOKEN=<your_hf_token> \
COSMOS3_DATA_ROOT=./data \
python -m cosmos_framework.scripts.train \
--config finetune_config.yaml
Multi-GPU Training with DeepSpeed
For Cosmos 3-Super, distribute training across 4 GPUs:
torchrun --nnodes=1 --nproc_per_node=4 \
-m cosmos_framework.scripts.train \
--config finetune_config.yaml \
--deepspeed deepspeed_config.json
The DeepSpeed configuration enables ZeRO-3 partitioning and optimizer offloading, essential for fitting the 64B parameter model into VRAM.
Inference and Evaluation
Load a fine-tuned checkpoint using Diffusers for immediate validation:
import torch
from diffusers import Cosmos3OmniPipeline
pipe = Cosmos3OmniPipeline.from_pretrained(
"./outputs/checkpoints/latest",
torch_dtype=torch.bfloat16,
device_map="auto",
)
result = pipe(
prompt="A warehouse robot now carries a red box.",
num_frames=120,
height=720,
width=1280,
fps=24,
num_inference_steps=30,
guidance_scale=5.0,
)
result.video.save("robot_finetuned.mp4")
Run automated evaluation to ensure model integrity:
cosmos_evaluator \
--model_dir ./outputs/checkpoints/latest \
--eval_data ./data/val_manifest.jsonl \
--metrics caption,grounding,temporal_localization \
--output ./eval_report.json
Monitor training progress in real-time:
tensorboard --logdir ./logs
Summary
- Start from official checkpoints: Use
nvidia/Cosmos3-Nanoornvidia/Cosmos3-Superto preserve multimodal knowledge. - Configure BF16 precision: Set
--mixed_precision bf16and use gradient accumulation for larger effective batch sizes. - Prepare data with Cosmos Curator: Convert raw media to JSONL manifests with proper modality alignment.
- Use distributed training: Launch
cosmos_framework.scripts.trainviatorchrunfor multi-GPU scaling with DeepSpeed. - Validate with Cosmos-Evaluator: Check safety guardrails and physical plausibility before deployment.
- Deploy to Hugging Face: Use
--push_to_hubfor integration with Diffusers, vLLM-Omni, or NIM.
Frequently Asked Questions
What is the minimum GPU requirement for fine-tuning Cosmos 3 models?
Cosmos 3-Nano can be fine-tuned on single GPUs with sufficient VRAM, while Cosmos 3-Super requires multiple GPUs or nodes due to its 64 billion parameters. The framework supports BF16 mixed precision to reduce memory requirements, and DeepSpeed ZeRO-3 offloading can help fit larger models on limited hardware.
How do I prepare custom datasets for fine-tuning?
Use the Cosmos Curator submodule to process raw videos, images, audio, and action logs into the required JSONL manifest format. The curator handles de-duplication, modality alignment, and annotation, ensuring compatibility with the Cosmos3OmniModel training pipeline.
Where can I find the training configuration templates?
The Cosmos Framework repository contains reference configurations in docs/training.md and the cosmos_framework/scripts/train.py module accepts YAML files specifying model architecture, optimizer settings, and learning rate schedules. These templates include sensible defaults for diffusion steps and tokenization.
Can I deploy fine-tuned models to production APIs?
Yes, fine-tuned checkpoints can be exported to Hugging Face using the --push_to_hub flag, then served via vLLM-Omni for production APIs, integrated into Diffusers for generation tasks, or deployed through NVIDIA NIM for containerized inference. The Cosmos-Evaluator ensures your model meets safety standards before deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →