# Implementing Forward Dynamics with Cosmos 3 for Robotics Simulation: A Complete Technical Guide

> Implement forward dynamics with Cosmos 3 for robotics simulation. This guide shows how to use the forward_dynamics action mode to predict future visual observations, enabling physics-aware simulations.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-13

---

**Cosmos 3 provides a dedicated `forward_dynamics` action mode that accepts an initial image paired with robot-action tokens to predict future visual observations as a video, enabling physics-aware robotics simulation through either the Cosmos Framework or a vLLM-Omni server.**

Cosmos 3, NVIDIA's open-source world foundation model, natively supports forward dynamics for robotics by conditioning its diffusion transformer on action tokens rather than text prompts. According to the `NVIDIA/cosmos` repository, this capability allows the model to forecast how a robot's movements will alter the visual environment over time, outputting a complete MP4 video of the predicted future state. This implementation guide references the specific source files, API contracts, and code patterns found in the official cookbooks and README documentation.

## Understanding Forward Dynamics in Cosmos 3

Forward dynamics in Cosmos 3 operates as a **generator task** within the model's unified architecture. Unlike text-to-video generation, this mode consumes an initial visual frame (JPG/PNG) and a sequence of robot actions (JSON array of floats) to produce a video visualizing the physically plausible outcome of those actions.

As documented in the project README (lines 354–355), the model processes these inputs by concatenating image tokens with action tokens before running the diffusion pipeline. The output is a full video tensor post-processed into MP4 format via `export_to_video` (Diffusers) or streamed directly from the vLLM-Omni server.

## Architecture Overview

### Model Surface and API

The forward dynamics capability is exposed through the `action_mode: "forward_dynamics"` parameter in the vLLM-Omni API. The model requires:

- **Input**: An image token stream followed by an action chunk (JSON array of joint-space or 9-D ego-motion values)
- **Output**: MP4 video (optionally with sound) representing the predicted future state

The supported action dimensions vary by embodiment—for example, 10D for single-arm DROID manipulation or 57D for egocentric motion, as listed in the README table (lines 548–555).

### Cosmos Framework Pipeline

The `cosmos_framework.scripts.inference` module serves as the primary entry point for local execution. This script wraps the `Cosmos3OmniPipeline` (Diffusers) and handles:

- Loading model checkpoints (`Cosmos3-Nano` or `Cosmos3-Super`)
- Assembling the request payload with `model_mode="forward_dynamics"`
- Running the diffusion pipeline in generator mode

### vLLM-Omni Server

For production deployments, the vLLM-Omni server exposes OpenAI-compatible endpoints:
- `/v1/videos/sync` – synchronous generation
- `/v1/videos` – asynchronous generation

The server reads action files from mounted paths (`--allowed-local-media-path`) and processes `extra_params` containing the action mode and domain specifications.

## Implementation Workflows

### Method 1: Cosmos Framework (Python Script)

The `run_fd_with_cosmos_framework.ipynb` notebook in `cookbooks/cosmos3/generator/action/` demonstrates the complete setup. First, configure the environment variables:

```python
import os

os.environ["COSMOS3_REPO"] = "/path/to/cosmos-framework"
os.environ["COSMOS3_UV_GROUP"] = "cu130-train"
os.environ["COSMOS3_OUTPUT_ROOT"] = "/tmp/cosmos_fd_outputs"
os.environ["HF_HOME"] = "/tmp/hf_cache"

```

Then launch the inference entry point. This command reads the spec file `action_forward_dynamics_robotics_custom.jsonl` and writes the output to `<COSMOS3_OUTPUT_ROOT>/action_forward_dynamics_robotics_custom/<run>/vision.mp4`:

```bash
python -m cosmos_framework.scripts.inference

```

The script [`cosmos_framework/scripts/inference.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/inference.py) (in the separate cosmos-framework repository) parses the `model_mode="forward_dynamics"` flag and invokes the `Cosmos3OmniPipeline` with the appropriate conditioning tokens.

### Method 2: vLLM-Omni API (cURL)

To query a running vLLM-Omni server (started via `vllm serve nvidia/Cosmos3-Nano --omni --model-class-name Cosmos3OmniDiffusersPipeline`), send a POST request to the synchronous endpoint:

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=Robot arm moves a cube to a shelf." \
  --form-string "negative_prompt=blur, low-quality" \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=35" \
  --form-string "guidance_scale=6.0" \
  --form-string "seed=42" \
  --form-string 'extra_params={"action_mode":"forward_dynamics","domain_name":"bridge_orig_lerobot","raw_action_dim":10,"action_chunk_size":5}' \
  -o forward_dynamics_output.mp4

```

The `extra_params` field must specify the `action_mode`, `domain_name` (e.g., `bridge_orig_lerobot`, `av`, or `camera_pose`), `raw_action_dim`, and `action_chunk_size` to match the expected input tensor shape.

### Method 3: OpenAI-Compatible Python Client

For programmatic access, use the OpenAI Python client against the vLLM-Omni server:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")

response = client.videos.sync.create(
    prompt="A mobile robot pushes a box across a warehouse floor.",
    size="1280x720",
    num_frames=189,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    seed=123,
    extra_params={
        "action_mode": "forward_dynamics",
        "domain_name": "av",
        "raw_action_dim": 9,
        "action_chunk_size": 8,
    },
)

with open("fd_video.mp4", "wb") as f:
    f.write(response.content)

```

The `run_fd_with_vllm.ipynb` notebook provides the complete reference implementation for this approach, including visualization via `IPython.display.Video`.

## Preparing Action Specifications

Forward dynamics requires a JSONL spec file where each line contains:

```json
{ "action": [0.1, -0.2, ...], "prompt": "Robot description", "image_path": "frame.jpg" }

```

Action dimensions vary by domain:
- **10D**: Single-arm DROID manipulation
- **57D**: Egocentric motion (camera pose)
- **9D**: Autonomous vehicle (AV) ego-motion

The pipeline converts these float arrays into action tokens that are concatenated with the visual token stream before diffusion. Ensure the `action_chunk_size` matches the sequence length expected by the specific model checkpoint.

## Summary

- **Forward dynamics** in Cosmos 3 uses the `action_mode="forward_dynamics"` parameter to predict future visual states from initial images and robot actions.
- **Two primary interfaces** exist: the `cosmos_framework.scripts.inference` CLI for local Diffusers-based execution, and the vLLM-Omni server (`/v1/videos/sync`) for production API deployment.
- **Input requirements** include a JSONL file with action arrays (dimensions vary: 10D for DROID, 57D for egocentric, 9D for AV) and a reference image.
- **Key source files** include `cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb` for framework usage and `cookbooks/cosmos3/generator/action/run_fd_with_vllm.ipynb` for API usage.
- **Output** is always an MP4 video generated by the `Cosmos3OmniPipeline` via diffusion, with unified architecture shared across text-to-video and image-to-video modes.

## Frequently Asked Questions

### What action dimensions does Cosmos 3 support for forward dynamics?

According to the README (lines 354–355), Cosmos 3 supports varying action dimensions depending on the robot embodiment. Single-arm DROID manipulation uses 10D joint-space controls, egocentric motion (camera pose) uses 57D parameters, and autonomous vehicle (AV) domains typically use 9D ego-motion values. The `raw_action_dim` parameter in your API request must match these specifications for the chosen `domain_name`.

### How do I choose between the Cosmos Framework and vLLM-Omni for implementation?

Use the **Cosmos Framework** (`cosmos_framework.scripts.inference`) for local development, research experimentation, or when you need direct access to the Diffusers pipeline and checkpoints. Deploy **vLLM-Omni** when you require a production-grade, OpenAI-compatible API server with synchronous (`/v1/videos/sync`) and asynchronous (`/v1/videos`) endpoints for serving multiple clients or integrating with existing robotics stacks.

### Can I use custom robot embodiments not listed in the predefined domain names?

While the README lists standard domains like `bridge_orig_lerobot`, `av`, and `camera_pose`, the architecture supports custom embodiments provided you correctly specify the `raw_action_dim` and format your action JSONL files to match the expected token sequence. The model's flexibility stems from its unified diffusion transformer that processes arbitrary action token streams, though performance may vary on out-of-distribution robot morphologies not represented in the training data.

### What is the difference between forward dynamics and standard text-to-video generation in Cosmos 3?

Both modes share the same **diffusion transformer backbone** and generate video through iterative denoising. However, forward dynamics conditions the generation on **action tokens** derived from robot joint states or ego-motion vectors, while text-to-video uses **text tokens** from prompts. This architectural distinction allows forward dynamics to maintain physical consistency with the robot's actuator constraints and the initial world state, producing predictions that respect the laws of physics for the specific embodiment.