# How to Use Forward Dynamics for Action-Conditioned Video Prediction in Cosmos 3

> Learn to use forward dynamics for action-conditioned video prediction with NVIDIA Cosmos. Generate future video rollouts from robot commands or camera poses.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: tutorial
- Published: 2026-06-05

---

**Set `action_mode` to `"forward_dynamics"`, provide an action trajectory JSON via `action_path`, and invoke either the Cosmos Framework inference script or the vLLM-Omni `POST /v1/videos/sync` endpoint to generate a future video rollout from robot commands, steering angles, or camera poses.**

NVIDIA Cosmos 3 provides **forward dynamics for action-conditioned video prediction** as a native generator mode that predicts future frames from an action sequence rather than from an input video. The same Mixture-of-Transformers (MoT) backbone that powers text-to-video and image-to-video generation also drives this mode by injecting raw action tokens into the diffusion transformer. In this guide you will learn how to structure requests, prepare action trajectories, and run inference through both the Python Framework and the OpenAI-compatible vLLM-Omni server.

## Architecture Overview: How Action Tokens Drive Video Generation

Cosmos 3 treats forward dynamics as a diffusion generation task where the conditioning signal is a temporal sequence of actions instead of pixels.

### Mixture-of-Transformers (MoT) and the Diffusion Transformer

The **Mixture-of-Transformers (MoT) backbone** processes multimodal tokens—text, vision, audio, and action—inside a unified transformer. According to [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), the identical MoT is shared between the Reasoner (causal) path and the Generator (diffusion) path, ensuring consistent token representations across tasks.

Inside the generator, the **diffusion transformer (DM)** iteratively denoises noisy multimodal tokens. Action tokens are injected as conditioning tokens (`action_chunk`) alongside any text prompt or visual frame conditioning.

### 3-D Multi-Dimensional Rotary Position Embedding (mRoPE)

To handle diverse embodiment spaces, Cosmos 3 employs a **3-D multi-dimensional rotary position embedding (mRoPE)** that jointly encodes spatial, temporal, and action dimensions. Whether the action space is 9-D for autonomous vehicles or 57-D for humanoid robots, this embedding preserves coherent world simulation across modalities.

## Request Specification for Forward Dynamics

A forward-dynamics request is defined by fields inside `extra_params`. The server interprets `action_mode` as the scheduling switch that routes tokens through the action-conditioning path.

Set the following parameters:

- `action_mode`: `"forward_dynamics"` — tells the diffusion model to ignore any input video and consume the supplied action trajectory.
- `action_path`: filesystem path to a JSON array of raw action vectors; the server must be launched with `--allowed-local-media-path` covering this directory.
- `domain_name`: embodiment tag such as `"av"` or `"droid_orig_lerobot"`.
- `raw_action_dim`: per-step action dimensionality (e.g., `9`, `10`, `57`).
- `action_chunk_size`: temporal length of the action sequence consumed by the model (commonly `16`).

Here is the complete request payload used by the Generator API:

```json
{
  "prompt": "A robot arm picks up a red block and places it on a shelf.",
  "size": "1280x720",
  "num_frames": 189,
  "fps": 24,
  "num_inference_steps": 35,
  "guidance_scale": 6.0,
  "extra_params": {
    "action_mode": "forward_dynamics",
    "domain_name": "droid_orig_lerobot",
    "raw_action_dim": 10,
    "action_chunk_size": 16,
    "action_path": "assets/actions/av_traj_forward.json"
  }
}

```

All fields are documented in the Generator API table in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) under the Forward Dynamics section.

## Running Forward Dynamics with the Cosmos Framework

The Python entry point is `cosmos_framework/scripts/inference`. When you provide `--extra_params` with `action_mode="forward_dynamics"`, the script loads the pipeline via `Cosmos3OmniPipeline.from_pretrained(...)`, executes the full diffusion schedule, and writes an MP4 file to disk.

Install the dependencies and launch inference as follows:

```bash

# 1. Install the framework (Diffusers and required deps)

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
uv pip install \
  "diffusers @ git+https://github.com/huggingface/diffusers.git" \
  accelerate av torch torchvision transformers \
  cosmos_guardrail

# 2. Launch the inference script

cosmos_framework/scripts/inference \
  --prompt "A small warehouse robot follows a path." \
  --size 1280x720 \
  --num_frames 189 \
  --fps 24 \
  --extra_params='{
      "action_mode":"forward_dynamics",
      "domain_name":"av",
      "raw_action_dim":9,
      "action_chunk_size":16,
      "action_path":"cookbooks/cosmos3/generator/action/assets/actions/av_traj_forward.json"
  }'

```

The script returns `output_video.mp4`. The end-to-end notebook is available at `cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb`.

## Running Forward Dynamics via the vLLM-Omni API

If you prefer an OpenAI-compatible HTTP interface, use the **vLLM-Omni** server. Submit a multipart `POST` request to `/v1/videos/sync` with the same `extra_params` JSON string. The server streams the generated MP4 bytes directly.

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=A robot arm lifts a cup." \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=35" \
  --form-string "guidance_scale=6.0" \
  --form-string 'extra_params={
      "action_mode":"forward_dynamics",
      "domain_name":"droid_orig_lerobot",
      "raw_action_dim":10,
      "action_chunk_size":16,
      "action_path":"cookbooks/cosmos3/generator/action/assets/actions/av_traj_forward.json"
  }' \
  -o forward_dynamics_output.mp4

```

The corresponding reference notebook is `cookbooks/cosmos3/generator/action/run_fd_with_vllm.ipynb`.

## Preparing the Action Trajectory JSON

The action file must contain a flat array of raw per-step action values. The format matches the examples under `cookbooks/cosmos3/generator/action/assets/actions/`.

```json
{
  "action": [
    [0.0, 0.01, -0.02, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0],
    [0.0, 0.02, -0.01, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0]
  ]
}

```

Each inner array length must equal `raw_action_dim` for the chosen `domain_name`. Consult the **Input & Output** table in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) to map embodiments to their required dimensions.

## Visualizing the Output

After generation, you can inspect the rollout with standard video tools. If you run the Jupyter notebook, the final cell typically calls:

```python
export_to_video(result.video, "fd_demo.mp4", fps=24)

```

Alternatively, open the saved MP4 with `ffplay` or any compatible player.

## Summary

- **Forward dynamics** in Cosmos 3 is a generator mode that predicts video rollouts from action trajectories instead of input video.
- The same **Mixture-of-Transformers (MoT)** and **diffusion transformer (DM)** architecture handles action tokens via the `action_chunk` conditioning path.
- A valid request requires `action_mode: "forward_dynamics"` and an `action_path` pointing to a JSON array of raw actions.
- You can execute inference through the **Cosmos Framework** Python script (`cosmos_framework/scripts/inference`) or the **vLLM-Omni** REST API (`POST /v1/videos/sync`).
- Action dimensions and embodiment identifiers are standardized in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) under the Input & Output reference table.

## Frequently Asked Questions

### What action dimension should I use for my robot or vehicle?

The required `raw_action_dim` depends on the embodiment. For example, the AV domain uses 9-D actions, the DROID-original LeRobot domain uses 10-D, and other humanoid embodiments may use 57-D. The **Input & Output** table in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) lists the exact mapping of `domain_name` to dimensionality, so you should reference that table before building your JSON trajectory.

### Do I need an input video to run forward dynamics?

No. Setting `action_mode` to `"forward_dynamics"` explicitly instructs the diffusion model to ignore any input video and use the supplied action trajectory to drive the rollout. You may still provide an initial image or text prompt for visual or semantic conditioning, but the temporal dynamics are governed entirely by the action tokens.

### Can I combine text prompts with action conditioning?

Yes. The forward-dynamics pipeline supports multimodal conditioning. In both the Cosmos Framework script and the vLLM-Omni API, you can supply a `prompt` string alongside the `action_path`. The MoT backbone fuses text tokens, optional visual tokens, and action tokens (`action_chunk`) before the DM transformer performs diffusion sampling.

### What is the performance difference between the Cosmos Framework and vLLM-Omni?

Both deployment paths expose the same forward-dynamics surface and call the same underlying model; the only difference is request routing. The Cosmos Framework runner is a local Python entry point ideal for interactive development, while vLLM-Omni offers an OpenAI-compatible HTTP server (`POST /v1/videos/sync`) better suited for production serving. Latency numbers for various hardware configurations are listed in [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md).