How to Use Forward Dynamics for Robot Trajectory Prediction with Action Conditioning in Cosmos 3

Cosmos 3 provides a Generator-mode "forward-dynamics" surface that predicts future visual observations of a robot from an initial image and a sequence of action commands, enabling physics-aware trajectory prediction without explicit simulators.

The NVIDIA/cosmos repository implements Cosmos 3, a world foundation model featuring a dedicated forward-dynamics pipeline for embodied AI. This mode allows you to perform forward dynamics for robot trajectory prediction by conditioning video generation on action sequences, effectively learning a world model that propagates robot actions through a diffusion-based generation process.

Architecture of Forward Dynamics

Cosmos 3's forward-dynamics implementation builds on the same Mixture-of-Transformers (MoT) architecture used throughout the model family. In this mode, the system processes a conditioning image alongside a domain-specific action tensor to autoregressively generate future visual states.

Multimodal Tokenization and MoT Backbone

The pipeline begins with a Multimodal Tokenizer that converts images, videos, and action vectors into a shared token space. The MoT Transformer then processes these tokens using causal self-attention for reasoning and full attention for diffusion generation. As implemented in the Cosmos framework, this transformer handles both text-only Reasoner requests and multimodal Generator requests, including the forward-dynamics mode.

3-D mRoPE Embedding

Spatial-temporal coherence is maintained through 3-D mRoPE (multimodal Rotary Position Embedding). This embedding provides position information that works across vision, audio, and action modalities, enabling the model to maintain coherent geometry when actions move the camera or robot end-effector through the scene.

Action Conditioning Mechanism

Action vectors are concatenated to the token stream as a dedicated "action block." The model learns to propagate the effect of each action step through the diffusion process, effectively learning a forward dynamics model of the world. When you set model_mode="forward_dynamics" in the request JSON, the diffusion transformer predicts only the visual stream without audio or additional actions, making it ideal for trajectory prediction and robot simulation.

Preparing Input Data

Forward dynamics requires two inputs: a conditioning image representing the initial observation, and an action tensor describing the trajectory to execute.

Conditioning Image

Prepare a single RGB frame that represents the robot's initial observation. This image should capture the workspace from the robot's camera perspective (ego-view) or a fixed external view, depending on your domain configuration.

Action Tensor Format

The action tensor must be a JSON array of shape (T, D), where T is the number of action steps and D is the domain-specific dimensionality. Cosmos 3 supports various robot embodiments:

  • 10-D for single-arm robots (e.g., DROID)
  • 20-D for dual-arm configurations
  • 57-D for egocentric motion scenarios

Action utilities for converting between pose representations reside in cosmos_framework/data/vfm/action/pose_utils.py. The framework includes helpers like pose_abs_to_rel and pose_rel_to_abs for transforming absolute camera poses to the relative 9-D format required by autonomous vehicle domains, with analogous helpers available for robot action formats.

Building the JSONL Specification

Cosmos 3 consumes forward-dynamics requests via a JSONL specification file. Each line describes one forward-dynamics run, specifying the image path, action path, domain name, and generation parameters.

The complete workflow is documented in cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb, which provides reference implementations for robotics, autonomous vehicle, and UMI domains.

import json
from pathlib import Path

vision_path = Path("cookbooks/cosmos3/generator/action/assets/images/robot_initial.jpg")

action_trajectory = [
    [0.05, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0],
    [0.05, 0.0, 0.0, 0.1, 0.0, 0.0, 0.0, 0.0, 0.0],
]

action_path = Path("tmp/robot_action.json")
action_path.write_text(json.dumps(action_trajectory, indent=2))

record = {
    "model_mode": "forward_dynamics",
    "domain_name": "droid_lerobot",
    "fps": 10,
    "image_size": 480,
    "view_point": "ego_view",
    "action_chunk_size": len(action_trajectory),
    "action_path": str(action_path),
    "vision_path": str(vision_path),
    "prompt": "A robot arm moving a cup on a table.",
    "seed": 0,
    "name": "robot_fd_example"
}

spec_path = Path("tmp/robot_fd_spec.jsonl")
spec_path.write_text(json.dumps(record) + "\n")

Running Forward Dynamics Inference

Execute the forward-dynamics generation using the CLI entry point cosmos_framework/scripts/inference.py, which parses the JSONL spec, loads the Cosmos 3 checkpoint, and streams the generated video to your output directory.

Command-Line Execution

export COSMOS3_REPO=$(pwd)/packages/cosmos3
export COSMOS3_CHECKPOINT_PATH=Cosmos3-Nano

python -m cosmos_framework.scripts.inference \
    -i tmp/robot_fd_spec.jsonl \
    -o outputs/robot_fd \
    --checkpoint-path $COSMOS3_CHECKPOINT_PATH \
    --image_size 480 \
    --seed 0 \
    --benchmark

The generated video appears under outputs/robot_fd/robot_fd_example/vision.mp4. Alternatively, you can stream results from a vLLM-Omni server via the /v1/videos/sync endpoint for production deployments.

Visualization

Preview the generated trajectory video directly in a Jupyter notebook:

import imageio_ffmpeg, subprocess
from IPython.display import Video, display
from pathlib import Path

FFMPEG = imageio_ffmpeg.get_ffmpeg_exe()
src = Path("outputs/robot_fd/robot_fd_example/vision.mp4")
preview = src.with_name(src.stem + "_preview.mp4")

subprocess.run([
    FFMPEG, "-y", "-i", str(src),
    "-c:v", "libx264", "-crf", "28",
    "-preset", "veryfast", "-an", "-pix_fmt", "yuv420p",
    str(preview)
], check=True)

display(Video(str(preview), embed=True))

Summary

  • Forward dynamics in Cosmos 3 generates future visual observations by conditioning on action sequences, implementing a learned world model for robot trajectory prediction.
  • The Mixture-of-Transformers architecture processes domain-specific action tensors (10-D, 20-D, or 57-D) alongside initial observations to produce physics-consistent video sequences.
  • Input preparation involves creating a JSONL specification that references your conditioning image and action trajectory file.
  • Use python -m cosmos_framework.scripts.inference to execute generation, or integrate with the vLLM-Omni API for scalable deployment.
  • Action utilities in cosmos_framework/data/vfm/action/pose_utils.py handle coordinate transformations between absolute and relative pose representations.

Frequently Asked Questions

What action space dimensions does Cosmos 3 support for robot trajectory prediction?

Cosmos 3 supports domain-specific action dimensions defined in the model configuration. Common schemas include 10-D for single-arm robots (such as DROID), 20-D for dual-arm configurations, and 57-D for egocentric motion. The domain_name field in your JSONL spec selects the appropriate tokenizer and action block format for your specific embodiment.

How do I convert between absolute and relative pose representations for action conditioning?

The framework provides utility functions in cosmos_framework/data/vfm/action/pose_utils.py for coordinate transformations. Use pose_abs_to_rel to convert absolute camera poses to the relative 9-D format required by autonomous vehicle domains, or apply the analogous robotics-specific helpers to normalize your action trajectories before serialization to JSON.

Can forward dynamics generate audio or only visual predictions?

In forward_dynamics mode, the diffusion transformer predicts only the visual stream. This mode explicitly excludes audio generation and does not output additional action tokens, optimizing the model specifically for visual trajectory prediction and robot simulation tasks.

What hardware requirements are needed for real-time forward dynamics inference?

Forward-dynamics inference requires NVIDIA GPUs with sufficient VRAM to load the Cosmos 3 checkpoints. The Cosmos 3-Nano model provides the fastest inference for development and testing, while larger variants require proportionally more GPU memory. Refer to inference_benchmarks.md in the repository root for latency metrics and recommended GPU configurations (such as H100 or A100 instances) for different throughput requirements.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →