# How to Use Forward Dynamics for Robot Trajectory Prediction with Action Conditioning in Cosmos 3

> Learn to predict robot trajectories using forward dynamics and action conditioning in Cosmos 3. Achieve physics-aware predictions without simulators and enhance your robotics projects.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-06

---

**Cosmos 3 provides a Generator-mode "forward-dynamics" surface that predicts future visual observations of a robot from an initial image and a sequence of action commands, enabling physics-aware trajectory prediction without explicit simulators.**

The NVIDIA/cosmos repository implements Cosmos 3, a world foundation model featuring a dedicated forward-dynamics pipeline for embodied AI. This mode allows you to perform forward dynamics for robot trajectory prediction by conditioning video generation on action sequences, effectively learning a world model that propagates robot actions through a diffusion-based generation process.

## Architecture of Forward Dynamics

Cosmos 3's forward-dynamics implementation builds on the same **Mixture-of-Transformers (MoT)** architecture used throughout the model family. In this mode, the system processes a conditioning image alongside a domain-specific action tensor to autoregressively generate future visual states.

### Multimodal Tokenization and MoT Backbone

The pipeline begins with a **Multimodal Tokenizer** that converts images, videos, and action vectors into a shared token space. The **MoT Transformer** then processes these tokens using causal self-attention for reasoning and full attention for diffusion generation. As implemented in the Cosmos framework, this transformer handles both text-only Reasoner requests and multimodal Generator requests, including the forward-dynamics mode.

### 3-D mRoPE Embedding

Spatial-temporal coherence is maintained through **3-D mRoPE (multimodal Rotary Position Embedding)**. This embedding provides position information that works across vision, audio, and action modalities, enabling the model to maintain coherent geometry when actions move the camera or robot end-effector through the scene.

### Action Conditioning Mechanism

Action vectors are concatenated to the token stream as a dedicated "action block." The model learns to propagate the effect of each action step through the diffusion process, effectively learning a **forward dynamics model** of the world. When you set `model_mode="forward_dynamics"` in the request JSON, the diffusion transformer predicts only the visual stream without audio or additional actions, making it ideal for trajectory prediction and robot simulation.

## Preparing Input Data

Forward dynamics requires two inputs: a conditioning image representing the initial observation, and an action tensor describing the trajectory to execute.

### Conditioning Image

Prepare a single RGB frame that represents the robot's initial observation. This image should capture the workspace from the robot's camera perspective (ego-view) or a fixed external view, depending on your domain configuration.

### Action Tensor Format

The action tensor must be a JSON array of shape **(T, D)**, where **T** is the number of action steps and **D** is the domain-specific dimensionality. Cosmos 3 supports various robot embodiments:

- **10-D** for single-arm robots (e.g., DROID)
- **20-D** for dual-arm configurations
- **57-D** for egocentric motion scenarios

Action utilities for converting between pose representations reside in [`cosmos_framework/data/vfm/action/pose_utils.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/data/vfm/action/pose_utils.py). The framework includes helpers like `pose_abs_to_rel` and `pose_rel_to_abs` for transforming absolute camera poses to the relative 9-D format required by autonomous vehicle domains, with analogous helpers available for robot action formats.

## Building the JSONL Specification

Cosmos 3 consumes forward-dynamics requests via a JSONL specification file. Each line describes one forward-dynamics run, specifying the image path, action path, domain name, and generation parameters.

The complete workflow is documented in `cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb`, which provides reference implementations for robotics, autonomous vehicle, and UMI domains.

```python
import json
from pathlib import Path

vision_path = Path("cookbooks/cosmos3/generator/action/assets/images/robot_initial.jpg")

action_trajectory = [
    [0.05, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0],
    [0.05, 0.0, 0.0, 0.1, 0.0, 0.0, 0.0, 0.0, 0.0],
]

action_path = Path("tmp/robot_action.json")
action_path.write_text(json.dumps(action_trajectory, indent=2))

record = {
    "model_mode": "forward_dynamics",
    "domain_name": "droid_lerobot",
    "fps": 10,
    "image_size": 480,
    "view_point": "ego_view",
    "action_chunk_size": len(action_trajectory),
    "action_path": str(action_path),
    "vision_path": str(vision_path),
    "prompt": "A robot arm moving a cup on a table.",
    "seed": 0,
    "name": "robot_fd_example"
}

spec_path = Path("tmp/robot_fd_spec.jsonl")
spec_path.write_text(json.dumps(record) + "\n")

```

## Running Forward Dynamics Inference

Execute the forward-dynamics generation using the CLI entry point [`cosmos_framework/scripts/inference.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/inference.py), which parses the JSONL spec, loads the Cosmos 3 checkpoint, and streams the generated video to your output directory.

### Command-Line Execution

```bash
export COSMOS3_REPO=$(pwd)/packages/cosmos3
export COSMOS3_CHECKPOINT_PATH=Cosmos3-Nano

python -m cosmos_framework.scripts.inference \
    -i tmp/robot_fd_spec.jsonl \
    -o outputs/robot_fd \
    --checkpoint-path $COSMOS3_CHECKPOINT_PATH \
    --image_size 480 \
    --seed 0 \
    --benchmark

```

The generated video appears under `outputs/robot_fd/robot_fd_example/vision.mp4`. Alternatively, you can stream results from a vLLM-Omni server via the `/v1/videos/sync` endpoint for production deployments.

### Visualization

Preview the generated trajectory video directly in a Jupyter notebook:

```python
import imageio_ffmpeg, subprocess
from IPython.display import Video, display
from pathlib import Path

FFMPEG = imageio_ffmpeg.get_ffmpeg_exe()
src = Path("outputs/robot_fd/robot_fd_example/vision.mp4")
preview = src.with_name(src.stem + "_preview.mp4")

subprocess.run([
    FFMPEG, "-y", "-i", str(src),
    "-c:v", "libx264", "-crf", "28",
    "-preset", "veryfast", "-an", "-pix_fmt", "yuv420p",
    str(preview)
], check=True)

display(Video(str(preview), embed=True))

```

## Summary

- **Forward dynamics** in Cosmos 3 generates future visual observations by conditioning on action sequences, implementing a learned world model for robot trajectory prediction.
- The **Mixture-of-Transformers** architecture processes domain-specific action tensors (10-D, 20-D, or 57-D) alongside initial observations to produce physics-consistent video sequences.
- Input preparation involves creating a **JSONL specification** that references your conditioning image and action trajectory file.
- Use `python -m cosmos_framework.scripts.inference` to execute generation, or integrate with the **vLLM-Omni** API for scalable deployment.
- Action utilities in [`cosmos_framework/data/vfm/action/pose_utils.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/data/vfm/action/pose_utils.py) handle coordinate transformations between absolute and relative pose representations.

## Frequently Asked Questions

### What action space dimensions does Cosmos 3 support for robot trajectory prediction?

Cosmos 3 supports domain-specific action dimensions defined in the model configuration. Common schemas include **10-D** for single-arm robots (such as DROID), **20-D** for dual-arm configurations, and **57-D** for egocentric motion. The `domain_name` field in your JSONL spec selects the appropriate tokenizer and action block format for your specific embodiment.

### How do I convert between absolute and relative pose representations for action conditioning?

The framework provides utility functions in [`cosmos_framework/data/vfm/action/pose_utils.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/data/vfm/action/pose_utils.py) for coordinate transformations. Use `pose_abs_to_rel` to convert absolute camera poses to the relative 9-D format required by autonomous vehicle domains, or apply the analogous robotics-specific helpers to normalize your action trajectories before serialization to JSON.

### Can forward dynamics generate audio or only visual predictions?

In `forward_dynamics` mode, the diffusion transformer predicts **only the visual stream**. This mode explicitly excludes audio generation and does not output additional action tokens, optimizing the model specifically for visual trajectory prediction and robot simulation tasks.

### What hardware requirements are needed for real-time forward dynamics inference?

Forward-dynamics inference requires NVIDIA GPUs with sufficient VRAM to load the Cosmos 3 checkpoints. The **Cosmos 3-Nano** model provides the fastest inference for development and testing, while larger variants require proportionally more GPU memory. Refer to [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) in the repository root for latency metrics and recommended GPU configurations (such as H100 or A100 instances) for different throughput requirements.