# How to Use Cosmos 3 for Embodied Reasoning and Next-Action Prediction

> Leverage Cosmos 3 an omnimodal world model for embodied reasoning and next-action prediction. Use its causal transformer Reasoner surface to plan tasks from video and text inputs.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Cosmos 3 is an omnimodal world model that uses a causal transformer-based Reasoner surface to predict next actions from video and text inputs, enabling embodied AI agents to perform physical reasoning and task planning.**

The NVIDIA Cosmos repository introduces Cosmos 3, an omnimodal world model designed specifically for embodied reasoning and next-action prediction. By unifying language, vision, audio, video, and action tokens within a single Mixture-of-Transformers (MoT) architecture, Cosmos 3 enables robots and autonomous agents to understand physical environments and predict future actions from multimodal inputs.

## Understanding Cosmos 3 Architecture for Embodied AI

### The Reasoner and Generator Surfaces

Cosmos 3 exposes two distinct runtime surfaces that share the same backbone but serve different purposes. The **Reasoner** surface accepts text and vision inputs to produce text outputs, making it ideal for embodied reasoning, world understanding, and next-action prediction. In contrast, the **Generator** surface handles text, vision, sound, and action inputs to synthesize vision, sound, and action outputs for world generation and simulation.

According to the surface table in the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), the Reasoner specifically targets "world understanding, grounding, physical reasoning, task planning, embodied reasoning, [and] next-action prediction."

### Unified Multimodal Processing with MoT

Both surfaces share core architectural components within the **Mixture-of-Transformers (MoT)** framework. The Reasoner employs a **causal transformer** running autoregressively with causal self-attention, while the Generator uses a **diffusion transformer** for denoising multimodal tokens.

Key shared components include:

- **Multimodal token embeddings** that project language tokens, image-patch tokens, video-frame tokens, audio-segment tokens, and action-state tokens into a common embedding space
- **3D multi-dimensional rotary position embedding (mRoPE)** that encodes spatial (height/width), temporal (frame index), and action-dimensional coordinates, maintaining consistent geometry across modalities

## Implementing Next-Action Prediction with the Reasoner

For embodied reasoning tasks, the Reasoner ingests video frames paired with textual queries (e.g., "What should the robot do next?") and outputs natural language descriptions of predicted actions. The model has been trained on diverse datasets including egocentric video, robot-arm actions, autonomous-vehicle trajectories, and DROID manipulation data.

### Local Inference with HuggingFace Transformers

To run embodied reasoning locally, use the `Cosmos3OmniForConditionalGeneration` class from the HuggingFace Transformers library. This implementation processes video inputs at specified frame rates and generates text responses.

```python
from pathlib import Path
import torch
from transformers import AutoProcessor, Cosmos3OmniForConditionalGeneration

model_id = "nvidia/Cosmos3-Nano"
video_path = Path("cookbooks/cosmos3/reasoner/assets/video_caption.mp4").resolve()

processor = AutoProcessor.from_pretrained(model_id)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "video", "path": str(video_path)},
            {"type": "text", "text": "What is the next action the robot should take?"},
        ],
    }
]

inputs = processor.apply_chat_template(
    messages,
    fps=2,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
).to("cuda", torch.bfloat16)

generated_ids = model.generate(**inputs, max_new_tokens=64, do_sample=False)
answer = processor.decode(generated_ids[0], skip_special_tokens=True)
print("Next‑action prediction →", answer)

```

This example, referenced in [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md), samples video at 2 frames per second and generates concise action predictions using **bfloat16** precision for efficient GPU inference.

### Deploying with vLLM-Omni

For production APIs, deploy the Reasoner using **vLLM-Omni** to create an OpenAI-compatible endpoint. This method supports higher throughput and standard HTTP interfaces.

```python
import requests
import json

url = "http://localhost:8000/v1/chat/completions"
headers = {"Content-Type": "application/json"}

payload = {
    "model": "nvidia/Cosmos3-Nano",
    "messages": [
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": "https://example.com/ego_clip.mp4"}},
                {"type": "text", "text": "Predict the next action for the robot."}
            ],
        },
    ],
    "max_tokens": 128,
    "temperature": 0.0,
    "extra_body": {"media_io_kwargs": {"video": {"fps": 4.0}}},
}
response = requests.post(url, headers=headers, data=json.dumps(payload))
print(response.json()["choices"][0]["message"]["content"])

```

The `extra_body` parameter configures video sampling at 4 FPS, as documented in the vLLM-Omni Reasoner section of the repository.

### Production Deployment with NVIDIA NIM

For containerized deployment, use the **NVIDIA NIM** microservice. The pre-built container exposes the same OpenAI-compatible API optimized for enterprise inference.

```python
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
response = client.chat.completions.create(
    model="nvidia/cosmos3-nano-reasoner",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": "https://example.com/ego_clip.mp4"}},
                {"type": "text", "text": "What should the robot do next?"},
            ],
        },
    ],
    max_tokens=128,
    temperature=0.0,
    extra_body={"media_io_kwargs": {"video": {"fps": 4.0}}},
)
print(response.choices[0].message.content)

```

## Key Source Files and Reference Implementation

The Cosmos 3 repository provides comprehensive resources for implementing embodied reasoning:

- [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (root): Overview of Reasoner capabilities and quick-start instructions
- [`cookbooks/cosmos3/reasoner/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/README.md): Detailed Reasoner setup and API formats
- `cookbooks/cosmos3/reasoner/run_with_vllm.ipynb`: Notebook demonstrating embodied reasoning against vLLM servers
- `cookbooks/cosmos3/reasoner/run_with_nim.ipynb`: NIM container deployment examples
- [`cosmos_framework/scripts/inference.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/inference.py): Framework entry point for both Reasoner and Generator inference
- `cookbooks/cosmos3/cosmos3-model-architecture.png`: Visual diagram of the unified MoT architecture

## Summary

- **Cosmos 3** unifies multimodal understanding through a **Mixture-of-Transformers (MoT)** architecture with specialized Reasoner and Generator surfaces.
- The **Reasoner** surface uses a **causal transformer** with **3D mRoPE** embeddings to predict next actions from video and text inputs.
- Three deployment options exist: **HuggingFace Transformers** for local development, **vLLM-Omni** for scalable APIs, and **NVIDIA NIM** for production containers.
- The model supports **variable frame rate sampling** (e.g., 2-4 FPS) for video inputs and outputs either natural language descriptions or structured action parameters.
- Implementation references are found in `cookbooks/cosmos3/reasoner/` and the root [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) in the NVIDIA/cosmos repository.

## Frequently Asked Questions

### What is the difference between the Reasoner and Generator surfaces in Cosmos 3?

The **Reasoner** employs a causal transformer with autoregressive generation to process text and vision inputs into text outputs, specializing in embodied reasoning and next-action prediction. The **Generator** uses a diffusion transformer to denoise and synthesize vision, sound, and action outputs from multimodal inputs, designed for world simulation and policy learning. Both share the same MoT backbone and mRoPE embedding system but operate with different attention mechanisms and generation strategies.

### How does Cosmos 3 handle spatial and temporal reasoning for robotics?

Cosmos 3 implements **3D multi-dimensional rotary position embedding (mRoPE)** that simultaneously encodes spatial coordinates (height/width), temporal frame indices, and action dimensions. This allows the model to maintain consistent geometric relationships across video frames and action sequences, critical for accurate embodied reasoning in robotic environments.

### Can Cosmos 3 output structured action parameters instead of text?

Yes, while the Reasoner naturally outputs natural language descriptions of next actions (e.g., "Move the arm to grasp the cup"), you can prompt the model to return **JSON payloads** with structured action parameters. The official README lists "Action modeling" capabilities including policy prediction, inverse dynamics, and forward dynamics, supporting structured outputs for robot control systems.

### What hardware requirements are needed to run Cosmos 3 for embodied reasoning?

The examples in the repository use **NVIDIA GPUs** with **bfloat16** precision support. The `Cosmos3-Nano` model can run on consumer GPUs, while larger variants require datacenter-class hardware. For production deployment, the vLLM-Omni and NIM implementations support multi-GPU configurations and optimized inference engines to handle real-time video processing for embodied AI applications.