Best Practices for Fine-Tuning Cosmos 3 Models Using the Cosmos Framework

Fine-tuning Cosmos 3 models requires starting from official NVIDIA checkpoints, configuring BF16 mixed precision with gradient accumulation, and following the Cosmos Framework's end-to-end workflow from data curation through distributed training to safety validation.

The NVIDIA Cosmos repository implements a unified Mixture-of-Transformers architecture designed for multimodal video, image, and audio understanding. When adapting these models to new domains, the Cosmos Framework provides the complete infrastructure for fine-tuning, including specialized tools for data preparation, distributed training orchestration, and post-training evaluation. Following the official fine-tuning practices ensures that your adapted model maintains the safety guardrails and physical plausibility inherent in the base Cosmos 3 knowledge.

Core Fine-Tuning Principles

Initialize From Official Checkpoints

Always begin fine-tuning from the official Cosmos 3 checkpoints to inherit the full multimodal knowledge base. The repository provides two primary model sizes: nvidia/Cosmos3-Nano for resource-constrained environments and nvidia/Cosmos3-Super (64B parameters) for maximum fidelity. According to the repository's README.md, these checkpoints include vision, video, audio, and action understanding weights that are essential for downstream tasks.

Configure Mixed Precision and Memory

Use BF16 mixed precision exclusively for fine-tuning on Ampere-generation GPUs and newer. This configuration reduces memory footprint while preserving training stability. Set --mixed_precision bf16 or export TORCH_PRECISION=bf16 to enable this mode. For limited GPU memory, implement gradient accumulation by setting --gradient_accumulation_steps N where the product of N and your per-device batch size equals your target global batch size.

Implement Distributed Training Strategies

Cosmos 3-Super requires multi-GPU or multi-node scaling due to its 64B parameter count. The framework supports distributed training via torchrun and DeepSpeed integration. As documented in cosmos_framework/scripts/train.py, the entry point handles both single-device and distributed configurations through the same YAML configuration interface.

Step-by-Step Fine-Tuning Workflow

Curate Multimodal Training Data

Convert raw videos, images, audio, and action logs into the Cosmos Curator format before training. This process creates a *.jsonl manifest paired with media files, handling de-duplication, modality alignment, and optional annotation. The Cosmos Curator submodule ensures your dataset conforms to the expected schema for the Cosmos3OmniModel training loop.

Define Training Configuration

Create a YAML configuration file (e.g., finetune_config.yaml) specifying model architecture, optimizer settings, and learning rate schedules. The configuration must reference the checkpoint directory, set the diffusion steps, and configure modality-specific tokenizers. The Cosmos Framework documentation at docs/training.md provides starter templates that include sensible defaults for learning rate schedules and warmup steps.

Launch Distributed Training

Execute training through the cosmos_framework.scripts.train module. For single-GPU training with Cosmos 3-Nano, run the module directly with Python. For multi-GPU training with Cosmos 3-Super, use torchrun with the --nnodes and --nproc_per_node arguments to scale across devices. The framework automatically handles checkpoint sharding and saves model states every save_checkpoint_steps to the specified output directory.

Validate and Deploy

After training, run the Cosmos-Evaluator on a held-out validation set to verify that safety guardrails and physical plausibility remain intact. The evaluator checks caption accuracy, spatial grounding, and temporal localization. Once validated, export checkpoints to Hugging Face using the --push_to_hub flag for integration with Diffusers, vLLM-Omni, or NVIDIA NIM for containerized inference.

Production Code Examples

Environment Setup and Installation

Install the Cosmos Framework with CUDA backend auto-detection to avoid torch.cuda.is_available() failures caused by mismatched torch wheels.


# Create Python 3.13 environment

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate

# Install with automatic CUDA backend detection

uv pip install --torch-backend=auto \
  "cosmos-framework @ git+https://github.com/NVIDIA/cosmos-framework.git" \
  deepspeed \
  tensorboard \
  wandb

Single-GPU Fine-Tuning

Create the configuration file for Nano model fine-tuning:


# finetune_config.yaml

model:
  name: "nvidia/Cosmos3-Nano"
  dtype: "bfloat16"

datasets:
  train: "./data/train_manifest.jsonl"
  validation: "./data/val_manifest.jsonl"

training:
  epochs: 3
  batch_size: 2
  gradient_accumulation_steps: 8
  learning_rate: 1e-5
  lr_schedule: "cosine"
  max_steps: 100_000
  mixed_precision: "bf16"

logging:
  tensorboard_dir: "./logs"
  report_to: "wandb"
  wandb_project: "cosmos3-finetune"

output:
  checkpoint_dir: "./outputs/checkpoints"
  push_to_hub: true

Launch single-GPU training:

HF_TOKEN=<your_hf_token> \
COSMOS3_DATA_ROOT=./data \
python -m cosmos_framework.scripts.train \
  --config finetune_config.yaml

Multi-GPU Training with DeepSpeed

For Cosmos 3-Super, distribute training across 4 GPUs:

torchrun --nnodes=1 --nproc_per_node=4 \
  -m cosmos_framework.scripts.train \
  --config finetune_config.yaml \
  --deepspeed deepspeed_config.json

The DeepSpeed configuration enables ZeRO-3 partitioning and optimizer offloading, essential for fitting the 64B parameter model into VRAM.

Inference and Evaluation

Load a fine-tuned checkpoint using Diffusers for immediate validation:

import torch
from diffusers import Cosmos3OmniPipeline

pipe = Cosmos3OmniPipeline.from_pretrained(
    "./outputs/checkpoints/latest",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

result = pipe(
    prompt="A warehouse robot now carries a red box.",
    num_frames=120,
    height=720,
    width=1280,
    fps=24,
    num_inference_steps=30,
    guidance_scale=5.0,
)
result.video.save("robot_finetuned.mp4")

Run automated evaluation to ensure model integrity:

cosmos_evaluator \
  --model_dir ./outputs/checkpoints/latest \
  --eval_data ./data/val_manifest.jsonl \
  --metrics caption,grounding,temporal_localization \
  --output ./eval_report.json

Monitor training progress in real-time:

tensorboard --logdir ./logs

Summary

  • Start from official checkpoints: Use nvidia/Cosmos3-Nano or nvidia/Cosmos3-Super to preserve multimodal knowledge.
  • Configure BF16 precision: Set --mixed_precision bf16 and use gradient accumulation for larger effective batch sizes.
  • Prepare data with Cosmos Curator: Convert raw media to JSONL manifests with proper modality alignment.
  • Use distributed training: Launch cosmos_framework.scripts.train via torchrun for multi-GPU scaling with DeepSpeed.
  • Validate with Cosmos-Evaluator: Check safety guardrails and physical plausibility before deployment.
  • Deploy to Hugging Face: Use --push_to_hub for integration with Diffusers, vLLM-Omni, or NIM.

Frequently Asked Questions

What is the minimum GPU requirement for fine-tuning Cosmos 3 models?

Cosmos 3-Nano can be fine-tuned on single GPUs with sufficient VRAM, while Cosmos 3-Super requires multiple GPUs or nodes due to its 64 billion parameters. The framework supports BF16 mixed precision to reduce memory requirements, and DeepSpeed ZeRO-3 offloading can help fit larger models on limited hardware.

How do I prepare custom datasets for fine-tuning?

Use the Cosmos Curator submodule to process raw videos, images, audio, and action logs into the required JSONL manifest format. The curator handles de-duplication, modality alignment, and annotation, ensuring compatibility with the Cosmos3OmniModel training pipeline.

Where can I find the training configuration templates?

The Cosmos Framework repository contains reference configurations in docs/training.md and the cosmos_framework/scripts/train.py module accepts YAML files specifying model architecture, optimizer settings, and learning rate schedules. These templates include sensible defaults for diffusion steps and tokenization.

Can I deploy fine-tuned models to production APIs?

Yes, fine-tuned checkpoints can be exported to Hugging Face using the --push_to_hub flag, then served via vLLM-Omni for production APIs, integrated into Diffusers for generation tasks, or deployed through NVIDIA NIM for containerized inference. The Cosmos-Evaluator ensures your model meets safety standards before deployment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →