Running Supervised Fine-Tuning on Cosmos 3 Vision Generators: A Complete Guide
Run supervised fine-tuning on Cosmos 3 vision generators using the provided cookbook scripts that automate dataset preparation, multi-GPU training, and checkpoint export for custom captioned video data.
NVIDIA Cosmos 3 provides a unified omnimodal model family capable of generating video from text prompts. The repository ships a production-ready cookbook located in cookbooks/cosmos3/generator/audiovisual/finetune/ that streamlines supervised fine-tuning (SFT) of the Vision generator on custom captioned-video datasets. This guide covers the complete workflow from environment setup to model export, enabling domain adaptation without manual training pipeline construction.
Prerequisites and Environment Setup
Before executing the training scripts, configure your environment with the Cosmos Framework and necessary dependencies.
Install the Cosmos Framework
The SFT recipes invoke cosmos_framework.scripts.train, requiring the framework source code to be placed under packages/cosmos3. Clone the framework repository and synchronize with CUDA-compatible wheels.
git clone https://github.com/NVIDIA/cosmos-framework.git packages/cosmos3
cd packages/cosmos3
uv sync --all-extras --group=cu130-train # Use cu128-train for CUDA 12.x
source .venv/bin/activate
Container Configuration
NVIDIA recommends using the NGC PyTorch container to ensure CUDA compatibility. Use nvcr.io/nvidia/pytorch:25.09-py3 for CUDA 13 or nvcr.io/nvidia/pytorch:25.06-py3 for CUDA 12.8. These containers include pre-built PyTorch and torchvision binaries matched to the specific CUDA versions required by the Cosmos Framework.
Authenticate with Hugging Face
The base model, Wan 2.2 VAE checkpoint, and the training dataset are hosted on Hugging Face and require authentication. Run uvx hf@latest auth login to cache credentials, or export HF_TOKEN as an environment variable for programmatic access.
Preparing the Training Data
The default training dataset for Vision SFT is BridgeData2-Subset-Synthetic-Captions. The launch scripts automatically download this dataset into data/BridgeData2-Subset-Synthetic-Captions if the folder is absent upon execution.
To manually download the dataset beforehand:
uvx hf@latest download \
nvidia/BridgeData2-Subset-Synthetic-Captions \
--repo-type dataset \
--local-dir data/BridgeData2-Subset-Synthetic-Captions
Running Supervised Fine-Tuning
The cookbook provides two distinct launch scripts in cookbooks/cosmos3/generator/audiovisual/finetune/ that automate the entire training workflow.
Full-Parameter Fine-Tuning on Cosmos3-Nano
Use launch_sft_vision_nano.sh for full-parameter fine-tuning of the Cosmos3-Nano model. This script defaults to multi-GPU training using torchrun with the TOML configuration file vision_sft_nano.toml.
cd $COSMOS_ROOT/cookbooks/cosmos3/generator/audiovisual/finetune
bash launch_sft_vision_nano.sh
LoRA Fine-Tuning on Cosmos3-Super
For lighter-weight adaptation with reduced memory requirements, use launch_sft_vision_super.sh to perform LoRA-style fine-tuning on Cosmos3-Super. This script utilizes vision_sft_super.toml for configuration.
bash launch_sft_vision_super.sh
Each script performs the following steps automatically:
- Verifies the dataset is present in the expected location
- Downloads the
Wan2.2_VAE.pthcheckpoint if not already cached - Converts the base model checkpoint to a local DCP (Distributed Checkpoint) format
- Launches
torchrunfor multi-GPU distributed training
Quick Smoke Testing
Validate your setup on a single GPU with minimal iterations before committing to full training:
export EXTRA_TAIL_OVERRIDES="job.wandb_mode=disabled trainer.max_iter=10 checkpoint.save_iter=10 dataloader_train.max_samples_per_batch=32"
bash launch_sft_vision_nano.sh
Managing Training Outputs and Checkpoints
Training artifacts are stored under outputs/train/<project>/<group>/<name>/. The directory structure includes:
checkpoints/iter_<N>/– Full DCP checkpoints containing model weights, optimizer states, and scheduler statesconfig.yaml– The exact TOML configuration used for the run, ensuring reproducibility- Log files and any registered callbacks
The DCP format preserves complete training state, enabling seamless resumption from any saved iteration.
Exporting Fine-Tuned Models for Inference
After SFT completes, convert the DCP checkpoints to Hugging Face-compatible safetensors format using the export utility. This enables integration with the audiovisual inference pipeline.
RUN_DIR=outputs/train/<project>/<group>/<name>
CKPT=$RUN_DIR/checkpoints/$(cat "$RUN_DIR/checkpoints/latest_checkpoint.txt")
python -m cosmos_framework.scripts.export_model \
--checkpoint-path "$CKPT" \
--config-file "$RUN_DIR/config.yaml" \
-o "$RUN_DIR/model"
The exported model in $RUN_DIR/model can be used with the audiovisual inference cookbook referenced in the parent directory.
Summary
- Install the Cosmos Framework under
packages/cosmos3and use NGC PyTorch containers (nvcr.io/nvidia/pytorch:25.09-py3or25.06-py3) for CUDA compatibility - Authenticate with Hugging Face to access base models, the
Wan2.2_VAE.pthcheckpoint, and the BridgeData2 training dataset - Execute
launch_sft_vision_nano.shfor full-parameter fine-tuning orlaunch_sft_vision_super.shfor LoRA-based adaptation - Training outputs use the DCP format in
outputs/train/<project>/<group>/<name>/checkpoints/ - Export fine-tuned models to
safetensorsusingcosmos_framework.scripts.export_modelfor downstream inference workloads
Frequently Asked Questions
What hardware configuration is required for fine-tuning Cosmos 3 vision generators?
Fine-tuning requires multi-GPU setups for full-parameter training (Cosmos3-Nano), while LoRA fine-tuning (Cosmos3-Super) offers a lighter-weight alternative suitable for reduced VRAM. The scripts utilize torchrun for distributed training across available GPUs. Ensure sufficient GPU memory and use the recommended NGC PyTorch containers for optimal CUDA compatibility.
How do I resume training from an existing checkpoint?
The training infrastructure preserves full DCP checkpoints including optimizer and scheduler states under checkpoints/iter_<N>/. To resume, specify the checkpoint path in your TOML configuration or modify the launch script to point to the desired iteration directory. The latest_checkpoint.txt file in the checkpoints directory tracks the most recent save for automatic resumption.
Can I use custom video datasets instead of BridgeData2?
Yes. While the cookbook defaults to BridgeData2-Subset-Synthetic-Captions, you can substitute any captioned video dataset by organizing it in the expected format and updating the dataloader configuration in the TOML config files. Ensure your dataset follows the same structure as the BridgeData2 format for compatibility with the existing data loading pipeline in cosmos_framework.
What is the difference between Nano and Super model fine-tuning?
Cosmos3-Nano fine-tuning performs full-parameter updates using launch_sft_vision_nano.sh and vision_sft_nano.toml, suitable when you have sufficient compute resources and want maximum model adaptability. Cosmos3-Super fine-tuning uses LoRA techniques via launch_sft_vision_super.sh and vision_sft_super.toml, updating only a subset of parameters for faster training and reduced memory requirements while maintaining generation quality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →