# Using Cosmos 3 for Physical Plausibility Analysis and Classification: A Complete Guide

> Learn how to use Cosmos 3 for physical plausibility analysis and classification. This guide covers the omni-model's two-step workflow for video synthesis and LLM-vision classification.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Cosmos 3 is an omni-model that couples a generator tower for video synthesis with a reasoner tower for large-language-vision classification, enabling end-to-end physical plausibility analysis through a unified two-step workflow.**

The NVIDIA Cosmos repository provides a unified framework for evaluating whether generated or real-world videos obey physical laws. By combining diffusion-based generation with vision-language reasoning in a single checkpoint, Cosmos 3 enables researchers to synthesize robotic scenarios and immediately classify their physical feasibility without switching models.

## Understanding the Cosmos 3 Architecture

The system is built on two synchronized towers that share the same checkpoint weights (available as `nvidia/Cosmos3-Nano` for fast prototyping or `nvidia/Cosmos3-Super` for production quality).

### The Generator Tower

The **generator** is a diffusion-based visual synthesis model that produces video or images from textual prompts. According to the source code in `evaluation/cosmos3/generator/rbench/run_with_cosmos_framework.ipynb`, it uses a UNet-style diffusion backbone with a motion-model head, supporting multiple modes including text-to-image (t2i), text-to-video (t2v), and image-to-video (i2v).

### The Reasoner Tower

The **reasoner** operates on the same transformer backbone as the generator but adds a projection head for text-only output. As implemented in `cookbooks/cosmos3/reasoner/run_with_transformers.ipynb`, this tower performs visual question answering, captioning, and binary classification (e.g., outputting "physically plausible" or "physically implausible").

## Setting Up Physical Plausibility Analysis

### Installation and Prerequisites

To run the complete pipeline, install the Cosmos Framework and select your inference backend. For high-throughput serving, use vLLM:

```bash
python -m venv .venv && source .venv/bin/activate
pip install "vllm[cosmos]>=0.23.0"

```

For local experimentation with Hugging Face libraries, install the base framework from the `cosmos/framework` directory.

### Loading Model Checkpoints

The entry point for both towers is defined in [`cosmos/framework/__init__.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos/framework/__init__.py). You can load specific components depending on your task:

- **Full pipeline**: `Cosmos3OmniPipeline` (Diffusers wrapper)
- **Reasoner only**: `Cosmos3OmniForConditionalGeneration` (Transformers wrapper)

## Generation and Classification Workflow

Physical plausibility analysis follows a two-stage pipeline where media generation precedes classification.

### Step 1: Generating Video with the Generator

Use the `Cosmos3OmniPipeline` to synthesize video from a text prompt describing a physical interaction. The generator supports configurable frame rates and durations:

```python
from cosmos.framework import Cosmos3OmniPipeline
import torch

pipeline = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.float16,
    device="cuda"
)

prompt = "A humanoid robot picks up a sachet from a table and places it into an open cardboard box."
video = pipeline(prompt, num_inference_steps=50, video_fps=24, video_length=5)
video.save("output.mp4")

```

*Source*: `cookbooks/cosmos3/generator/action/run_id_with_cosmos_framework.ipynb`

### Step 2: Classifying Plausibility with the Reasoner

Feed the generated video or the original prompt to the reasoner to obtain a binary label. The reasoner predicts `true` for physically plausible scenarios and `false` for implausible ones, as defined in the benchmark annotations at [`evaluation/cosmos3/generator/rbench/assets/prompts/visual_reasoning_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/evaluation/cosmos3/generator/rbench/assets/prompts/visual_reasoning_prompts.json):

```python
from cosmos.framework import Cosmos3OmniForConditionalGeneration
from transformers import AutoTokenizer
import torch

model = Cosmos3OmniForConditionalGeneration.from_pretrained(
    "nvidia/Cosmos3-Nano", 
    torch_dtype=torch.float16, 
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("nvidia/Cosmos3-Nano")

input_text = "A humanoid robot lifts a sachet and drops it into a cardboard box."
inputs = tokenizer(input_text, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=20)
label = tokenizer.decode(outputs[0], skip_special_tokens=True)
print("Physical plausibility:", label)

```

*Source*: `cookbooks/cosmos3/reasoner/run_with_transformers.ipynb`

## Implementation Methods

Depending on your throughput requirements and infrastructure, you can deploy Cosmos 3 through three distinct APIs.

### Method 1: Diffusers Pipeline (Generation)

The **Diffusers** integration provides the most flexible sampling options for the generator tower. The `Cosmos3OmniPipeline` class handles checkpoint loading, tokenizer setup, and scheduler configuration automatically.

### Method 2: Transformers (Reasoner Classification)

For classification tasks that only require the reasoner, load `Cosmos3OmniForConditionalGeneration` to avoid initializing the diffusion components. This reduces memory footprint by loading only the vision-language model head, as demonstrated in `cookbooks/cosmos3/reasoner/run_with_transformers.ipynb`.

### Method 3: vLLM Serving (High-Throughput)

For benchmarking at scale, serve the model with tensor parallelism across multiple GPUs. The implementation in `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb` supports 4-GPU tensor-parallel inference:

```bash
vllm serve nvidia/Cosmos3-Super \
    --tensor-parallel-size 4 \
    --port 8000 \
    --max-model-len 4096

```

Query the service via HTTP:

```bash
curl -X POST http://localhost:8000/generate \
  -H "Content-Type: application/json" \
  -d '{"prompt":"A robot places a sachet into a box","max_new_tokens":20}'

```

## Benchmarking with PhysicsIQ and RBench

The repository includes dedicated evaluation suites for physical plausibility. The **PhysicsIQ** benchmark (located in `evaluation/cosmos3/generator/physics_iq/`) and **RBench** provide JSON-annotated prompts with ground-truth labels for classification accuracy scoring.

To reproduce the PhysicsIQ scores referenced in [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md), run the notebook at `evaluation/cosmos3/generator/physics_iq/run_with_cosmos_framework.ipynb`. This script automatically downloads the 90GB `Cosmos3-Super` checkpoint, generates videos for each benchmark prompt, runs the reasoner classification, and computes the final plausibility score.

## Summary

- **Cosmos 3** unifies generation and reasoning in a single omni-model with shared checkpoints (`nvidia/Cosmos3-Nano` or `nvidia/Cosmos3-Super`).
- The **generator** tower uses diffusion to synthesize video, while the **reasoner** tower classifies physical plausibility via the `Cosmos3OmniForConditionalGeneration` class.
- Three entry points are available: **PyTorch** (reference), **Diffusers** (flexible sampling), and **vLLM** (scalable serving with tensor parallelism).
- Physical plausibility is evaluated against **PhysicsIQ** and **RBench** datasets, which provide labeled ground truth in [`evaluation/cosmos3/generator/rbench/assets/prompts/visual_reasoning_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/evaluation/cosmos3/generator/rbench/assets/prompts/visual_reasoning_prompts.json).

## Frequently Asked Questions

### What is the difference between Cosmos3-Nano and Cosmos3-Super?

**Cosmos3-Nano** is optimized for fast prototyping and lower GPU memory requirements, while **Cosmos3-Super** provides higher-quality video generation and more accurate reasoning at the cost of a 90GB checkpoint size and increased inference latency. Both share the same architecture but differ in model capacity and training scale.

### Can I use the reasoner without generating video first?

Yes. The `Cosmos3OmniForConditionalGeneration` class loads only the reasoner tower from the checkpoint, allowing you to classify existing videos or textual descriptions of physical scenes without initializing the diffusion components. This is useful for batch-processing existing datasets.

### How do I evaluate classification accuracy against the benchmarks?

Run the notebooks in `evaluation/cosmos3/generator/physics_iq/` or `evaluation/cosmos3/generator/rbench/`. These scripts generate media for each benchmark prompt, query the reasoner for plausibility labels, and compare predictions against the ground-truth annotations stored in the respective JSON files to compute accuracy metrics.

### What hardware is required for tensor-parallel inference?

The vLLM integration supports tensor parallelism across 4+ GPUs for the `Cosmos3-Super` checkpoint. According to [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md), this configuration is necessary to achieve the throughput benchmarks listed for high-resolution video generation and classification tasks.