How to Use Cosmos 3 for Embodied Reasoning and Next-Action Prediction

Cosmos 3 is an omnimodal world model that uses a causal transformer-based Reasoner surface to predict next actions from video and text inputs, enabling embodied AI agents to perform physical reasoning and task planning.

The NVIDIA Cosmos repository introduces Cosmos 3, an omnimodal world model designed specifically for embodied reasoning and next-action prediction. By unifying language, vision, audio, video, and action tokens within a single Mixture-of-Transformers (MoT) architecture, Cosmos 3 enables robots and autonomous agents to understand physical environments and predict future actions from multimodal inputs.

Understanding Cosmos 3 Architecture for Embodied AI

The Reasoner and Generator Surfaces

Cosmos 3 exposes two distinct runtime surfaces that share the same backbone but serve different purposes. The Reasoner surface accepts text and vision inputs to produce text outputs, making it ideal for embodied reasoning, world understanding, and next-action prediction. In contrast, the Generator surface handles text, vision, sound, and action inputs to synthesize vision, sound, and action outputs for world generation and simulation.

According to the surface table in the repository's README.md, the Reasoner specifically targets "world understanding, grounding, physical reasoning, task planning, embodied reasoning, [and] next-action prediction."

Unified Multimodal Processing with MoT

Both surfaces share core architectural components within the Mixture-of-Transformers (MoT) framework. The Reasoner employs a causal transformer running autoregressively with causal self-attention, while the Generator uses a diffusion transformer for denoising multimodal tokens.

Key shared components include:

  • Multimodal token embeddings that project language tokens, image-patch tokens, video-frame tokens, audio-segment tokens, and action-state tokens into a common embedding space
  • 3D multi-dimensional rotary position embedding (mRoPE) that encodes spatial (height/width), temporal (frame index), and action-dimensional coordinates, maintaining consistent geometry across modalities

Implementing Next-Action Prediction with the Reasoner

For embodied reasoning tasks, the Reasoner ingests video frames paired with textual queries (e.g., "What should the robot do next?") and outputs natural language descriptions of predicted actions. The model has been trained on diverse datasets including egocentric video, robot-arm actions, autonomous-vehicle trajectories, and DROID manipulation data.

Local Inference with HuggingFace Transformers

To run embodied reasoning locally, use the Cosmos3OmniForConditionalGeneration class from the HuggingFace Transformers library. This implementation processes video inputs at specified frame rates and generates text responses.

from pathlib import Path
import torch
from transformers import AutoProcessor, Cosmos3OmniForConditionalGeneration

model_id = "nvidia/Cosmos3-Nano"
video_path = Path("cookbooks/cosmos3/reasoner/assets/video_caption.mp4").resolve()

processor = AutoProcessor.from_pretrained(model_id)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "video", "path": str(video_path)},
            {"type": "text", "text": "What is the next action the robot should take?"},
        ],
    }
]

inputs = processor.apply_chat_template(
    messages,
    fps=2,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
).to("cuda", torch.bfloat16)

generated_ids = model.generate(**inputs, max_new_tokens=64, do_sample=False)
answer = processor.decode(generated_ids[0], skip_special_tokens=True)
print("Next‑action prediction →", answer)

This example, referenced in cookbooks/cosmos3/reasoner/README.md, samples video at 2 frames per second and generates concise action predictions using bfloat16 precision for efficient GPU inference.

Deploying with vLLM-Omni

For production APIs, deploy the Reasoner using vLLM-Omni to create an OpenAI-compatible endpoint. This method supports higher throughput and standard HTTP interfaces.

import requests
import json

url = "http://localhost:8000/v1/chat/completions"
headers = {"Content-Type": "application/json"}

payload = {
    "model": "nvidia/Cosmos3-Nano",
    "messages": [
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": "https://example.com/ego_clip.mp4"}},
                {"type": "text", "text": "Predict the next action for the robot."}
            ],
        },
    ],
    "max_tokens": 128,
    "temperature": 0.0,
    "extra_body": {"media_io_kwargs": {"video": {"fps": 4.0}}},
}
response = requests.post(url, headers=headers, data=json.dumps(payload))
print(response.json()["choices"][0]["message"]["content"])

The extra_body parameter configures video sampling at 4 FPS, as documented in the vLLM-Omni Reasoner section of the repository.

Production Deployment with NVIDIA NIM

For containerized deployment, use the NVIDIA NIM microservice. The pre-built container exposes the same OpenAI-compatible API optimized for enterprise inference.

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
response = client.chat.completions.create(
    model="nvidia/cosmos3-nano-reasoner",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {
            "role": "user",
            "content": [
                {"type": "video_url", "video_url": {"url": "https://example.com/ego_clip.mp4"}},
                {"type": "text", "text": "What should the robot do next?"},
            ],
        },
    ],
    max_tokens=128,
    temperature=0.0,
    extra_body={"media_io_kwargs": {"video": {"fps": 4.0}}},
)
print(response.choices[0].message.content)

Key Source Files and Reference Implementation

The Cosmos 3 repository provides comprehensive resources for implementing embodied reasoning:

  • README.md (root): Overview of Reasoner capabilities and quick-start instructions
  • cookbooks/cosmos3/reasoner/README.md: Detailed Reasoner setup and API formats
  • cookbooks/cosmos3/reasoner/run_with_vllm.ipynb: Notebook demonstrating embodied reasoning against vLLM servers
  • cookbooks/cosmos3/reasoner/run_with_nim.ipynb: NIM container deployment examples
  • cosmos_framework/scripts/inference.py: Framework entry point for both Reasoner and Generator inference
  • cookbooks/cosmos3/cosmos3-model-architecture.png: Visual diagram of the unified MoT architecture

Summary

  • Cosmos 3 unifies multimodal understanding through a Mixture-of-Transformers (MoT) architecture with specialized Reasoner and Generator surfaces.
  • The Reasoner surface uses a causal transformer with 3D mRoPE embeddings to predict next actions from video and text inputs.
  • Three deployment options exist: HuggingFace Transformers for local development, vLLM-Omni for scalable APIs, and NVIDIA NIM for production containers.
  • The model supports variable frame rate sampling (e.g., 2-4 FPS) for video inputs and outputs either natural language descriptions or structured action parameters.
  • Implementation references are found in cookbooks/cosmos3/reasoner/ and the root README.md in the NVIDIA/cosmos repository.

Frequently Asked Questions

What is the difference between the Reasoner and Generator surfaces in Cosmos 3?

The Reasoner employs a causal transformer with autoregressive generation to process text and vision inputs into text outputs, specializing in embodied reasoning and next-action prediction. The Generator uses a diffusion transformer to denoise and synthesize vision, sound, and action outputs from multimodal inputs, designed for world simulation and policy learning. Both share the same MoT backbone and mRoPE embedding system but operate with different attention mechanisms and generation strategies.

How does Cosmos 3 handle spatial and temporal reasoning for robotics?

Cosmos 3 implements 3D multi-dimensional rotary position embedding (mRoPE) that simultaneously encodes spatial coordinates (height/width), temporal frame indices, and action dimensions. This allows the model to maintain consistent geometric relationships across video frames and action sequences, critical for accurate embodied reasoning in robotic environments.

Can Cosmos 3 output structured action parameters instead of text?

Yes, while the Reasoner naturally outputs natural language descriptions of next actions (e.g., "Move the arm to grasp the cup"), you can prompt the model to return JSON payloads with structured action parameters. The official README lists "Action modeling" capabilities including policy prediction, inverse dynamics, and forward dynamics, supporting structured outputs for robot control systems.

What hardware requirements are needed to run Cosmos 3 for embodied reasoning?

The examples in the repository use NVIDIA GPUs with bfloat16 precision support. The Cosmos3-Nano model can run on consumer GPUs, while larger variants require datacenter-class hardware. For production deployment, the vLLM-Omni and NIM implementations support multi-GPU configurations and optimized inference engines to handle real-time video processing for embodied AI applications.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →