Using Cosmos 3 for Physical Plausibility Analysis and Classification: A Complete Guide
Cosmos 3 is an omni-model that couples a generator tower for video synthesis with a reasoner tower for large-language-vision classification, enabling end-to-end physical plausibility analysis through a unified two-step workflow.
The NVIDIA Cosmos repository provides a unified framework for evaluating whether generated or real-world videos obey physical laws. By combining diffusion-based generation with vision-language reasoning in a single checkpoint, Cosmos 3 enables researchers to synthesize robotic scenarios and immediately classify their physical feasibility without switching models.
Understanding the Cosmos 3 Architecture
The system is built on two synchronized towers that share the same checkpoint weights (available as nvidia/Cosmos3-Nano for fast prototyping or nvidia/Cosmos3-Super for production quality).
The Generator Tower
The generator is a diffusion-based visual synthesis model that produces video or images from textual prompts. According to the source code in evaluation/cosmos3/generator/rbench/run_with_cosmos_framework.ipynb, it uses a UNet-style diffusion backbone with a motion-model head, supporting multiple modes including text-to-image (t2i), text-to-video (t2v), and image-to-video (i2v).
The Reasoner Tower
The reasoner operates on the same transformer backbone as the generator but adds a projection head for text-only output. As implemented in cookbooks/cosmos3/reasoner/run_with_transformers.ipynb, this tower performs visual question answering, captioning, and binary classification (e.g., outputting "physically plausible" or "physically implausible").
Setting Up Physical Plausibility Analysis
Installation and Prerequisites
To run the complete pipeline, install the Cosmos Framework and select your inference backend. For high-throughput serving, use vLLM:
python -m venv .venv && source .venv/bin/activate
pip install "vllm[cosmos]>=0.23.0"
For local experimentation with Hugging Face libraries, install the base framework from the cosmos/framework directory.
Loading Model Checkpoints
The entry point for both towers is defined in cosmos/framework/__init__.py. You can load specific components depending on your task:
- Full pipeline:
Cosmos3OmniPipeline(Diffusers wrapper) - Reasoner only:
Cosmos3OmniForConditionalGeneration(Transformers wrapper)
Generation and Classification Workflow
Physical plausibility analysis follows a two-stage pipeline where media generation precedes classification.
Step 1: Generating Video with the Generator
Use the Cosmos3OmniPipeline to synthesize video from a text prompt describing a physical interaction. The generator supports configurable frame rates and durations:
from cosmos.framework import Cosmos3OmniPipeline
import torch
pipeline = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.float16,
device="cuda"
)
prompt = "A humanoid robot picks up a sachet from a table and places it into an open cardboard box."
video = pipeline(prompt, num_inference_steps=50, video_fps=24, video_length=5)
video.save("output.mp4")
Source: cookbooks/cosmos3/generator/action/run_id_with_cosmos_framework.ipynb
Step 2: Classifying Plausibility with the Reasoner
Feed the generated video or the original prompt to the reasoner to obtain a binary label. The reasoner predicts true for physically plausible scenarios and false for implausible ones, as defined in the benchmark annotations at evaluation/cosmos3/generator/rbench/assets/prompts/visual_reasoning_prompts.json:
from cosmos.framework import Cosmos3OmniForConditionalGeneration
from transformers import AutoTokenizer
import torch
model = Cosmos3OmniForConditionalGeneration.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.float16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("nvidia/Cosmos3-Nano")
input_text = "A humanoid robot lifts a sachet and drops it into a cardboard box."
inputs = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=20)
label = tokenizer.decode(outputs[0], skip_special_tokens=True)
print("Physical plausibility:", label)
Source: cookbooks/cosmos3/reasoner/run_with_transformers.ipynb
Implementation Methods
Depending on your throughput requirements and infrastructure, you can deploy Cosmos 3 through three distinct APIs.
Method 1: Diffusers Pipeline (Generation)
The Diffusers integration provides the most flexible sampling options for the generator tower. The Cosmos3OmniPipeline class handles checkpoint loading, tokenizer setup, and scheduler configuration automatically.
Method 2: Transformers (Reasoner Classification)
For classification tasks that only require the reasoner, load Cosmos3OmniForConditionalGeneration to avoid initializing the diffusion components. This reduces memory footprint by loading only the vision-language model head, as demonstrated in cookbooks/cosmos3/reasoner/run_with_transformers.ipynb.
Method 3: vLLM Serving (High-Throughput)
For benchmarking at scale, serve the model with tensor parallelism across multiple GPUs. The implementation in cookbooks/cosmos3/reasoner/run_with_vllm.ipynb supports 4-GPU tensor-parallel inference:
vllm serve nvidia/Cosmos3-Super \
--tensor-parallel-size 4 \
--port 8000 \
--max-model-len 4096
Query the service via HTTP:
curl -X POST http://localhost:8000/generate \
-H "Content-Type: application/json" \
-d '{"prompt":"A robot places a sachet into a box","max_new_tokens":20}'
Benchmarking with PhysicsIQ and RBench
The repository includes dedicated evaluation suites for physical plausibility. The PhysicsIQ benchmark (located in evaluation/cosmos3/generator/physics_iq/) and RBench provide JSON-annotated prompts with ground-truth labels for classification accuracy scoring.
To reproduce the PhysicsIQ scores referenced in inference_benchmarks.md, run the notebook at evaluation/cosmos3/generator/physics_iq/run_with_cosmos_framework.ipynb. This script automatically downloads the 90GB Cosmos3-Super checkpoint, generates videos for each benchmark prompt, runs the reasoner classification, and computes the final plausibility score.
Summary
- Cosmos 3 unifies generation and reasoning in a single omni-model with shared checkpoints (
nvidia/Cosmos3-Nanoornvidia/Cosmos3-Super). - The generator tower uses diffusion to synthesize video, while the reasoner tower classifies physical plausibility via the
Cosmos3OmniForConditionalGenerationclass. - Three entry points are available: PyTorch (reference), Diffusers (flexible sampling), and vLLM (scalable serving with tensor parallelism).
- Physical plausibility is evaluated against PhysicsIQ and RBench datasets, which provide labeled ground truth in
evaluation/cosmos3/generator/rbench/assets/prompts/visual_reasoning_prompts.json.
Frequently Asked Questions
What is the difference between Cosmos3-Nano and Cosmos3-Super?
Cosmos3-Nano is optimized for fast prototyping and lower GPU memory requirements, while Cosmos3-Super provides higher-quality video generation and more accurate reasoning at the cost of a 90GB checkpoint size and increased inference latency. Both share the same architecture but differ in model capacity and training scale.
Can I use the reasoner without generating video first?
Yes. The Cosmos3OmniForConditionalGeneration class loads only the reasoner tower from the checkpoint, allowing you to classify existing videos or textual descriptions of physical scenes without initializing the diffusion components. This is useful for batch-processing existing datasets.
How do I evaluate classification accuracy against the benchmarks?
Run the notebooks in evaluation/cosmos3/generator/physics_iq/ or evaluation/cosmos3/generator/rbench/. These scripts generate media for each benchmark prompt, query the reasoner for plausibility labels, and compare predictions against the ground-truth annotations stored in the respective JSON files to compute accuracy metrics.
What hardware is required for tensor-parallel inference?
The vLLM integration supports tensor parallelism across 4+ GPUs for the Cosmos3-Super checkpoint. According to inference_benchmarks.md, this configuration is necessary to achieve the throughput benchmarks listed for high-resolution video generation and classification tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →