# Best Practices for Fine-Tuning Cosmos 3 Models Using the Cosmos Framework

> Master fine-tuning Cosmos 3 models with the NVIDIA Cosmos Framework. Learn best practices for data curation, mixed precision training, and safety validation to optimize your models.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: best-practices
- Published: 2026-06-14

---

**Fine-tuning Cosmos 3 models requires starting from official NVIDIA checkpoints, configuring BF16 mixed precision with gradient accumulation, and following the Cosmos Framework's end-to-end workflow from data curation through distributed training to safety validation.**

The NVIDIA Cosmos repository implements a unified Mixture-of-Transformers architecture designed for multimodal video, image, and audio understanding. When adapting these models to new domains, the Cosmos Framework provides the complete infrastructure for fine-tuning, including specialized tools for data preparation, distributed training orchestration, and post-training evaluation. Following the official fine-tuning practices ensures that your adapted model maintains the safety guardrails and physical plausibility inherent in the base Cosmos 3 knowledge.

## Core Fine-Tuning Principles

### Initialize From Official Checkpoints

Always begin fine-tuning from the official Cosmos 3 checkpoints to inherit the full multimodal knowledge base. The repository provides two primary model sizes: `nvidia/Cosmos3-Nano` for resource-constrained environments and `nvidia/Cosmos3-Super` (64B parameters) for maximum fidelity. According to the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), these checkpoints include vision, video, audio, and action understanding weights that are essential for downstream tasks.

### Configure Mixed Precision and Memory

Use **BF16 mixed precision** exclusively for fine-tuning on Ampere-generation GPUs and newer. This configuration reduces memory footprint while preserving training stability. Set `--mixed_precision bf16` or export `TORCH_PRECISION=bf16` to enable this mode. For limited GPU memory, implement **gradient accumulation** by setting `--gradient_accumulation_steps N` where the product of N and your per-device batch size equals your target global batch size.

### Implement Distributed Training Strategies

Cosmos 3-Super requires multi-GPU or multi-node scaling due to its 64B parameter count. The framework supports distributed training via `torchrun` and DeepSpeed integration. As documented in [`cosmos_framework/scripts/train.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/train.py), the entry point handles both single-device and distributed configurations through the same YAML configuration interface.

## Step-by-Step Fine-Tuning Workflow

### Curate Multimodal Training Data

Convert raw videos, images, audio, and action logs into the **Cosmos Curator** format before training. This process creates a `*.jsonl` manifest paired with media files, handling de-duplication, modality alignment, and optional annotation. The Cosmos Curator submodule ensures your dataset conforms to the expected schema for the `Cosmos3OmniModel` training loop.

### Define Training Configuration

Create a YAML configuration file (e.g., [`finetune_config.yaml`](https://github.com/NVIDIA/cosmos/blob/main/finetune_config.yaml)) specifying model architecture, optimizer settings, and learning rate schedules. The configuration must reference the checkpoint directory, set the diffusion steps, and configure modality-specific tokenizers. The Cosmos Framework documentation at [`docs/training.md`](https://github.com/NVIDIA/cosmos/blob/main/docs/training.md) provides starter templates that include sensible defaults for learning rate schedules and warmup steps.

### Launch Distributed Training

Execute training through the `cosmos_framework.scripts.train` module. For single-GPU training with Cosmos 3-Nano, run the module directly with Python. For multi-GPU training with Cosmos 3-Super, use `torchrun` with the `--nnodes` and `--nproc_per_node` arguments to scale across devices. The framework automatically handles checkpoint sharding and saves model states every `save_checkpoint_steps` to the specified output directory.

### Validate and Deploy

After training, run the **Cosmos-Evaluator** on a held-out validation set to verify that safety guardrails and physical plausibility remain intact. The evaluator checks caption accuracy, spatial grounding, and temporal localization. Once validated, export checkpoints to Hugging Face using the `--push_to_hub` flag for integration with Diffusers, vLLM-Omni, or NVIDIA NIM for containerized inference.

## Production Code Examples

### Environment Setup and Installation

Install the Cosmos Framework with CUDA backend auto-detection to avoid `torch.cuda.is_available()` failures caused by mismatched torch wheels.

```bash

# Create Python 3.13 environment

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate

# Install with automatic CUDA backend detection

uv pip install --torch-backend=auto \
  "cosmos-framework @ git+https://github.com/NVIDIA/cosmos-framework.git" \
  deepspeed \
  tensorboard \
  wandb

```

### Single-GPU Fine-Tuning

Create the configuration file for Nano model fine-tuning:

```yaml

# finetune_config.yaml

model:
  name: "nvidia/Cosmos3-Nano"
  dtype: "bfloat16"

datasets:
  train: "./data/train_manifest.jsonl"
  validation: "./data/val_manifest.jsonl"

training:
  epochs: 3
  batch_size: 2
  gradient_accumulation_steps: 8
  learning_rate: 1e-5
  lr_schedule: "cosine"
  max_steps: 100_000
  mixed_precision: "bf16"

logging:
  tensorboard_dir: "./logs"
  report_to: "wandb"
  wandb_project: "cosmos3-finetune"

output:
  checkpoint_dir: "./outputs/checkpoints"
  push_to_hub: true

```

Launch single-GPU training:

```bash
HF_TOKEN=<your_hf_token> \
COSMOS3_DATA_ROOT=./data \
python -m cosmos_framework.scripts.train \
  --config finetune_config.yaml

```

### Multi-GPU Training with DeepSpeed

For Cosmos 3-Super, distribute training across 4 GPUs:

```bash
torchrun --nnodes=1 --nproc_per_node=4 \
  -m cosmos_framework.scripts.train \
  --config finetune_config.yaml \
  --deepspeed deepspeed_config.json

```

The DeepSpeed configuration enables ZeRO-3 partitioning and optimizer offloading, essential for fitting the 64B parameter model into VRAM.

### Inference and Evaluation

Load a fine-tuned checkpoint using Diffusers for immediate validation:

```python
import torch
from diffusers import Cosmos3OmniPipeline

pipe = Cosmos3OmniPipeline.from_pretrained(
    "./outputs/checkpoints/latest",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

result = pipe(
    prompt="A warehouse robot now carries a red box.",
    num_frames=120,
    height=720,
    width=1280,
    fps=24,
    num_inference_steps=30,
    guidance_scale=5.0,
)
result.video.save("robot_finetuned.mp4")

```

Run automated evaluation to ensure model integrity:

```bash
cosmos_evaluator \
  --model_dir ./outputs/checkpoints/latest \
  --eval_data ./data/val_manifest.jsonl \
  --metrics caption,grounding,temporal_localization \
  --output ./eval_report.json

```

Monitor training progress in real-time:

```bash
tensorboard --logdir ./logs

```

## Summary

- **Start from official checkpoints**: Use `nvidia/Cosmos3-Nano` or `nvidia/Cosmos3-Super` to preserve multimodal knowledge.
- **Configure BF16 precision**: Set `--mixed_precision bf16` and use gradient accumulation for larger effective batch sizes.
- **Prepare data with Cosmos Curator**: Convert raw media to JSONL manifests with proper modality alignment.
- **Use distributed training**: Launch `cosmos_framework.scripts.train` via `torchrun` for multi-GPU scaling with DeepSpeed.
- **Validate with Cosmos-Evaluator**: Check safety guardrails and physical plausibility before deployment.
- **Deploy to Hugging Face**: Use `--push_to_hub` for integration with Diffusers, vLLM-Omni, or NIM.

## Frequently Asked Questions

### What is the minimum GPU requirement for fine-tuning Cosmos 3 models?

Cosmos 3-Nano can be fine-tuned on single GPUs with sufficient VRAM, while Cosmos 3-Super requires multiple GPUs or nodes due to its 64 billion parameters. The framework supports BF16 mixed precision to reduce memory requirements, and DeepSpeed ZeRO-3 offloading can help fit larger models on limited hardware.

### How do I prepare custom datasets for fine-tuning?

Use the **Cosmos Curator** submodule to process raw videos, images, audio, and action logs into the required JSONL manifest format. The curator handles de-duplication, modality alignment, and annotation, ensuring compatibility with the `Cosmos3OmniModel` training pipeline.

### Where can I find the training configuration templates?

The Cosmos Framework repository contains reference configurations in [`docs/training.md`](https://github.com/NVIDIA/cosmos/blob/main/docs/training.md) and the [`cosmos_framework/scripts/train.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/train.py) module accepts YAML files specifying model architecture, optimizer settings, and learning rate schedules. These templates include sensible defaults for diffusion steps and tokenization.

### Can I deploy fine-tuned models to production APIs?

Yes, fine-tuned checkpoints can be exported to Hugging Face using the `--push_to_hub` flag, then served via vLLM-Omni for production APIs, integrated into Diffusers for generation tasks, or deployed through NVIDIA NIM for containerized inference. The Cosmos-Evaluator ensures your model meets safety standards before deployment.