Running Supervised Fine-Tuning on Cosmos 3 Vision Generators: A Complete Guide

Run supervised fine-tuning on Cosmos 3 vision generators using the provided cookbook scripts that automate dataset preparation, multi-GPU training, and checkpoint export for custom captioned video data.

NVIDIA Cosmos 3 provides a unified omnimodal model family capable of generating video from text prompts. The repository ships a production-ready cookbook located in cookbooks/cosmos3/generator/audiovisual/finetune/ that streamlines supervised fine-tuning (SFT) of the Vision generator on custom captioned-video datasets. This guide covers the complete workflow from environment setup to model export, enabling domain adaptation without manual training pipeline construction.

Prerequisites and Environment Setup

Before executing the training scripts, configure your environment with the Cosmos Framework and necessary dependencies.

Install the Cosmos Framework

The SFT recipes invoke cosmos_framework.scripts.train, requiring the framework source code to be placed under packages/cosmos3. Clone the framework repository and synchronize with CUDA-compatible wheels.

git clone https://github.com/NVIDIA/cosmos-framework.git packages/cosmos3
cd packages/cosmos3
uv sync --all-extras --group=cu130-train   # Use cu128-train for CUDA 12.x

source .venv/bin/activate

Container Configuration

NVIDIA recommends using the NGC PyTorch container to ensure CUDA compatibility. Use nvcr.io/nvidia/pytorch:25.09-py3 for CUDA 13 or nvcr.io/nvidia/pytorch:25.06-py3 for CUDA 12.8. These containers include pre-built PyTorch and torchvision binaries matched to the specific CUDA versions required by the Cosmos Framework.

Authenticate with Hugging Face

The base model, Wan 2.2 VAE checkpoint, and the training dataset are hosted on Hugging Face and require authentication. Run uvx hf@latest auth login to cache credentials, or export HF_TOKEN as an environment variable for programmatic access.

Preparing the Training Data

The default training dataset for Vision SFT is BridgeData2-Subset-Synthetic-Captions. The launch scripts automatically download this dataset into data/BridgeData2-Subset-Synthetic-Captions if the folder is absent upon execution.

To manually download the dataset beforehand:

uvx hf@latest download \
    nvidia/BridgeData2-Subset-Synthetic-Captions \
    --repo-type dataset \
    --local-dir data/BridgeData2-Subset-Synthetic-Captions

Running Supervised Fine-Tuning

The cookbook provides two distinct launch scripts in cookbooks/cosmos3/generator/audiovisual/finetune/ that automate the entire training workflow.

Full-Parameter Fine-Tuning on Cosmos3-Nano

Use launch_sft_vision_nano.sh for full-parameter fine-tuning of the Cosmos3-Nano model. This script defaults to multi-GPU training using torchrun with the TOML configuration file vision_sft_nano.toml.

cd $COSMOS_ROOT/cookbooks/cosmos3/generator/audiovisual/finetune
bash launch_sft_vision_nano.sh

LoRA Fine-Tuning on Cosmos3-Super

For lighter-weight adaptation with reduced memory requirements, use launch_sft_vision_super.sh to perform LoRA-style fine-tuning on Cosmos3-Super. This script utilizes vision_sft_super.toml for configuration.

bash launch_sft_vision_super.sh

Each script performs the following steps automatically:

  1. Verifies the dataset is present in the expected location
  2. Downloads the Wan2.2_VAE.pth checkpoint if not already cached
  3. Converts the base model checkpoint to a local DCP (Distributed Checkpoint) format
  4. Launches torchrun for multi-GPU distributed training

Quick Smoke Testing

Validate your setup on a single GPU with minimal iterations before committing to full training:

export EXTRA_TAIL_OVERRIDES="job.wandb_mode=disabled trainer.max_iter=10 checkpoint.save_iter=10 dataloader_train.max_samples_per_batch=32"
bash launch_sft_vision_nano.sh

Managing Training Outputs and Checkpoints

Training artifacts are stored under outputs/train/<project>/<group>/<name>/. The directory structure includes:

  • checkpoints/iter_<N>/ – Full DCP checkpoints containing model weights, optimizer states, and scheduler states
  • config.yaml – The exact TOML configuration used for the run, ensuring reproducibility
  • Log files and any registered callbacks

The DCP format preserves complete training state, enabling seamless resumption from any saved iteration.

Exporting Fine-Tuned Models for Inference

After SFT completes, convert the DCP checkpoints to Hugging Face-compatible safetensors format using the export utility. This enables integration with the audiovisual inference pipeline.

RUN_DIR=outputs/train/<project>/<group>/<name>
CKPT=$RUN_DIR/checkpoints/$(cat "$RUN_DIR/checkpoints/latest_checkpoint.txt")
python -m cosmos_framework.scripts.export_model \
    --checkpoint-path "$CKPT" \
    --config-file "$RUN_DIR/config.yaml" \
    -o "$RUN_DIR/model"

The exported model in $RUN_DIR/model can be used with the audiovisual inference cookbook referenced in the parent directory.

Summary

  • Install the Cosmos Framework under packages/cosmos3 and use NGC PyTorch containers (nvcr.io/nvidia/pytorch:25.09-py3 or 25.06-py3) for CUDA compatibility
  • Authenticate with Hugging Face to access base models, the Wan2.2_VAE.pth checkpoint, and the BridgeData2 training dataset
  • Execute launch_sft_vision_nano.sh for full-parameter fine-tuning or launch_sft_vision_super.sh for LoRA-based adaptation
  • Training outputs use the DCP format in outputs/train/<project>/<group>/<name>/checkpoints/
  • Export fine-tuned models to safetensors using cosmos_framework.scripts.export_model for downstream inference workloads

Frequently Asked Questions

What hardware configuration is required for fine-tuning Cosmos 3 vision generators?

Fine-tuning requires multi-GPU setups for full-parameter training (Cosmos3-Nano), while LoRA fine-tuning (Cosmos3-Super) offers a lighter-weight alternative suitable for reduced VRAM. The scripts utilize torchrun for distributed training across available GPUs. Ensure sufficient GPU memory and use the recommended NGC PyTorch containers for optimal CUDA compatibility.

How do I resume training from an existing checkpoint?

The training infrastructure preserves full DCP checkpoints including optimizer and scheduler states under checkpoints/iter_<N>/. To resume, specify the checkpoint path in your TOML configuration or modify the launch script to point to the desired iteration directory. The latest_checkpoint.txt file in the checkpoints directory tracks the most recent save for automatic resumption.

Can I use custom video datasets instead of BridgeData2?

Yes. While the cookbook defaults to BridgeData2-Subset-Synthetic-Captions, you can substitute any captioned video dataset by organizing it in the expected format and updating the dataloader configuration in the TOML config files. Ensure your dataset follows the same structure as the BridgeData2 format for compatibility with the existing data loading pipeline in cosmos_framework.

What is the difference between Nano and Super model fine-tuning?

Cosmos3-Nano fine-tuning performs full-parameter updates using launch_sft_vision_nano.sh and vision_sft_nano.toml, suitable when you have sufficient compute resources and want maximum model adaptability. Cosmos3-Super fine-tuning uses LoRA techniques via launch_sft_vision_super.sh and vision_sft_super.toml, updating only a subset of parameters for faster training and reduced memory requirements while maintaining generation quality.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →