# Running Supervised Fine-Tuning on Cosmos 3 Vision Generators: A Complete Guide

> Learn to run supervised fine-tuning on Cosmos 3 vision generators with NVIDIA's cookbook scripts. Automate dataset prep, multi-GPU training, and checkpoint export for your custom video data.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Run supervised fine-tuning on Cosmos 3 vision generators using the provided cookbook scripts that automate dataset preparation, multi-GPU training, and checkpoint export for custom captioned video data.**

NVIDIA Cosmos 3 provides a unified omnimodal model family capable of generating video from text prompts. The repository ships a production-ready cookbook located in `cookbooks/cosmos3/generator/audiovisual/finetune/` that streamlines supervised fine-tuning (SFT) of the Vision generator on custom captioned-video datasets. This guide covers the complete workflow from environment setup to model export, enabling domain adaptation without manual training pipeline construction.

## Prerequisites and Environment Setup

Before executing the training scripts, configure your environment with the Cosmos Framework and necessary dependencies.

### Install the Cosmos Framework

The SFT recipes invoke `cosmos_framework.scripts.train`, requiring the framework source code to be placed under `packages/cosmos3`. Clone the framework repository and synchronize with CUDA-compatible wheels.

```bash
git clone https://github.com/NVIDIA/cosmos-framework.git packages/cosmos3
cd packages/cosmos3
uv sync --all-extras --group=cu130-train   # Use cu128-train for CUDA 12.x

source .venv/bin/activate

```

### Container Configuration

NVIDIA recommends using the NGC PyTorch container to ensure CUDA compatibility. Use `nvcr.io/nvidia/pytorch:25.09-py3` for CUDA 13 or `nvcr.io/nvidia/pytorch:25.06-py3` for CUDA 12.8. These containers include pre-built PyTorch and torchvision binaries matched to the specific CUDA versions required by the Cosmos Framework.

### Authenticate with Hugging Face

The base model, Wan 2.2 VAE checkpoint, and the training dataset are hosted on Hugging Face and require authentication. Run `uvx hf@latest auth login` to cache credentials, or export `HF_TOKEN` as an environment variable for programmatic access.

## Preparing the Training Data

The default training dataset for Vision SFT is **BridgeData2-Subset-Synthetic-Captions**. The launch scripts automatically download this dataset into `data/BridgeData2-Subset-Synthetic-Captions` if the folder is absent upon execution.

To manually download the dataset beforehand:

```bash
uvx hf@latest download \
    nvidia/BridgeData2-Subset-Synthetic-Captions \
    --repo-type dataset \
    --local-dir data/BridgeData2-Subset-Synthetic-Captions

```

## Running Supervised Fine-Tuning

The cookbook provides two distinct launch scripts in `cookbooks/cosmos3/generator/audiovisual/finetune/` that automate the entire training workflow.

### Full-Parameter Fine-Tuning on Cosmos3-Nano

Use [`launch_sft_vision_nano.sh`](https://github.com/NVIDIA/cosmos/blob/main/launch_sft_vision_nano.sh) for full-parameter fine-tuning of the Cosmos3-Nano model. This script defaults to multi-GPU training using `torchrun` with the TOML configuration file [`vision_sft_nano.toml`](https://github.com/NVIDIA/cosmos/blob/main/vision_sft_nano.toml).

```bash
cd $COSMOS_ROOT/cookbooks/cosmos3/generator/audiovisual/finetune
bash launch_sft_vision_nano.sh

```

### LoRA Fine-Tuning on Cosmos3-Super

For lighter-weight adaptation with reduced memory requirements, use [`launch_sft_vision_super.sh`](https://github.com/NVIDIA/cosmos/blob/main/launch_sft_vision_super.sh) to perform LoRA-style fine-tuning on Cosmos3-Super. This script utilizes [`vision_sft_super.toml`](https://github.com/NVIDIA/cosmos/blob/main/vision_sft_super.toml) for configuration.

```bash
bash launch_sft_vision_super.sh

```

Each script performs the following steps automatically:

1. Verifies the dataset is present in the expected location
2. Downloads the `Wan2.2_VAE.pth` checkpoint if not already cached
3. Converts the base model checkpoint to a local DCP (Distributed Checkpoint) format
4. Launches `torchrun` for multi-GPU distributed training

### Quick Smoke Testing

Validate your setup on a single GPU with minimal iterations before committing to full training:

```bash
export EXTRA_TAIL_OVERRIDES="job.wandb_mode=disabled trainer.max_iter=10 checkpoint.save_iter=10 dataloader_train.max_samples_per_batch=32"
bash launch_sft_vision_nano.sh

```

## Managing Training Outputs and Checkpoints

Training artifacts are stored under `outputs/train/<project>/<group>/<name>/`. The directory structure includes:

- `checkpoints/iter_<N>/` – Full DCP checkpoints containing model weights, optimizer states, and scheduler states
- [`config.yaml`](https://github.com/NVIDIA/cosmos/blob/main/config.yaml) – The exact TOML configuration used for the run, ensuring reproducibility
- Log files and any registered callbacks

The DCP format preserves complete training state, enabling seamless resumption from any saved iteration.

## Exporting Fine-Tuned Models for Inference

After SFT completes, convert the DCP checkpoints to Hugging Face-compatible `safetensors` format using the export utility. This enables integration with the audiovisual inference pipeline.

```bash
RUN_DIR=outputs/train/<project>/<group>/<name>
CKPT=$RUN_DIR/checkpoints/$(cat "$RUN_DIR/checkpoints/latest_checkpoint.txt")
python -m cosmos_framework.scripts.export_model \
    --checkpoint-path "$CKPT" \
    --config-file "$RUN_DIR/config.yaml" \
    -o "$RUN_DIR/model"

```

The exported model in `$RUN_DIR/model` can be used with the audiovisual inference cookbook referenced in the parent directory.

## Summary

- Install the Cosmos Framework under `packages/cosmos3` and use NGC PyTorch containers (`nvcr.io/nvidia/pytorch:25.09-py3` or `25.06-py3`) for CUDA compatibility
- Authenticate with Hugging Face to access base models, the `Wan2.2_VAE.pth` checkpoint, and the BridgeData2 training dataset
- Execute [`launch_sft_vision_nano.sh`](https://github.com/NVIDIA/cosmos/blob/main/launch_sft_vision_nano.sh) for full-parameter fine-tuning or [`launch_sft_vision_super.sh`](https://github.com/NVIDIA/cosmos/blob/main/launch_sft_vision_super.sh) for LoRA-based adaptation
- Training outputs use the DCP format in `outputs/train/<project>/<group>/<name>/checkpoints/`
- Export fine-tuned models to `safetensors` using `cosmos_framework.scripts.export_model` for downstream inference workloads

## Frequently Asked Questions

### What hardware configuration is required for fine-tuning Cosmos 3 vision generators?

Fine-tuning requires multi-GPU setups for full-parameter training (Cosmos3-Nano), while LoRA fine-tuning (Cosmos3-Super) offers a lighter-weight alternative suitable for reduced VRAM. The scripts utilize `torchrun` for distributed training across available GPUs. Ensure sufficient GPU memory and use the recommended NGC PyTorch containers for optimal CUDA compatibility.

### How do I resume training from an existing checkpoint?

The training infrastructure preserves full DCP checkpoints including optimizer and scheduler states under `checkpoints/iter_<N>/`. To resume, specify the checkpoint path in your TOML configuration or modify the launch script to point to the desired iteration directory. The [`latest_checkpoint.txt`](https://github.com/NVIDIA/cosmos/blob/main/latest_checkpoint.txt) file in the checkpoints directory tracks the most recent save for automatic resumption.

### Can I use custom video datasets instead of BridgeData2?

Yes. While the cookbook defaults to BridgeData2-Subset-Synthetic-Captions, you can substitute any captioned video dataset by organizing it in the expected format and updating the dataloader configuration in the TOML config files. Ensure your dataset follows the same structure as the BridgeData2 format for compatibility with the existing data loading pipeline in `cosmos_framework`.

### What is the difference between Nano and Super model fine-tuning?

Cosmos3-Nano fine-tuning performs full-parameter updates using [`launch_sft_vision_nano.sh`](https://github.com/NVIDIA/cosmos/blob/main/launch_sft_vision_nano.sh) and [`vision_sft_nano.toml`](https://github.com/NVIDIA/cosmos/blob/main/vision_sft_nano.toml), suitable when you have sufficient compute resources and want maximum model adaptability. Cosmos3-Super fine-tuning uses LoRA techniques via [`launch_sft_vision_super.sh`](https://github.com/NVIDIA/cosmos/blob/main/launch_sft_vision_super.sh) and [`vision_sft_super.toml`](https://github.com/NVIDIA/cosmos/blob/main/vision_sft_super.toml), updating only a subset of parameters for faster training and reduced memory requirements while maintaining generation quality.