How to Use Forward Dynamics for Action-Conditioned Video Prediction in Cosmos 3

Set action_mode to "forward_dynamics", provide an action trajectory JSON via action_path, and invoke either the Cosmos Framework inference script or the vLLM-Omni POST /v1/videos/sync endpoint to generate a future video rollout from robot commands, steering angles, or camera poses.

NVIDIA Cosmos 3 provides forward dynamics for action-conditioned video prediction as a native generator mode that predicts future frames from an action sequence rather than from an input video. The same Mixture-of-Transformers (MoT) backbone that powers text-to-video and image-to-video generation also drives this mode by injecting raw action tokens into the diffusion transformer. In this guide you will learn how to structure requests, prepare action trajectories, and run inference through both the Python Framework and the OpenAI-compatible vLLM-Omni server.

Architecture Overview: How Action Tokens Drive Video Generation

Cosmos 3 treats forward dynamics as a diffusion generation task where the conditioning signal is a temporal sequence of actions instead of pixels.

Mixture-of-Transformers (MoT) and the Diffusion Transformer

The Mixture-of-Transformers (MoT) backbone processes multimodal tokens—text, vision, audio, and action—inside a unified transformer. According to README.md, the identical MoT is shared between the Reasoner (causal) path and the Generator (diffusion) path, ensuring consistent token representations across tasks.

Inside the generator, the diffusion transformer (DM) iteratively denoises noisy multimodal tokens. Action tokens are injected as conditioning tokens (action_chunk) alongside any text prompt or visual frame conditioning.

3-D Multi-Dimensional Rotary Position Embedding (mRoPE)

To handle diverse embodiment spaces, Cosmos 3 employs a 3-D multi-dimensional rotary position embedding (mRoPE) that jointly encodes spatial, temporal, and action dimensions. Whether the action space is 9-D for autonomous vehicles or 57-D for humanoid robots, this embedding preserves coherent world simulation across modalities.

Request Specification for Forward Dynamics

A forward-dynamics request is defined by fields inside extra_params. The server interprets action_mode as the scheduling switch that routes tokens through the action-conditioning path.

Set the following parameters:

  • action_mode: "forward_dynamics" — tells the diffusion model to ignore any input video and consume the supplied action trajectory.
  • action_path: filesystem path to a JSON array of raw action vectors; the server must be launched with --allowed-local-media-path covering this directory.
  • domain_name: embodiment tag such as "av" or "droid_orig_lerobot".
  • raw_action_dim: per-step action dimensionality (e.g., 9, 10, 57).
  • action_chunk_size: temporal length of the action sequence consumed by the model (commonly 16).

Here is the complete request payload used by the Generator API:

{
  "prompt": "A robot arm picks up a red block and places it on a shelf.",
  "size": "1280x720",
  "num_frames": 189,
  "fps": 24,
  "num_inference_steps": 35,
  "guidance_scale": 6.0,
  "extra_params": {
    "action_mode": "forward_dynamics",
    "domain_name": "droid_orig_lerobot",
    "raw_action_dim": 10,
    "action_chunk_size": 16,
    "action_path": "assets/actions/av_traj_forward.json"
  }
}

All fields are documented in the Generator API table in README.md under the Forward Dynamics section.

Running Forward Dynamics with the Cosmos Framework

The Python entry point is cosmos_framework/scripts/inference. When you provide --extra_params with action_mode="forward_dynamics", the script loads the pipeline via Cosmos3OmniPipeline.from_pretrained(...), executes the full diffusion schedule, and writes an MP4 file to disk.

Install the dependencies and launch inference as follows:


# 1. Install the framework (Diffusers and required deps)

uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
uv pip install \
  "diffusers @ git+https://github.com/huggingface/diffusers.git" \
  accelerate av torch torchvision transformers \
  cosmos_guardrail

# 2. Launch the inference script

cosmos_framework/scripts/inference \
  --prompt "A small warehouse robot follows a path." \
  --size 1280x720 \
  --num_frames 189 \
  --fps 24 \
  --extra_params='{
      "action_mode":"forward_dynamics",
      "domain_name":"av",
      "raw_action_dim":9,
      "action_chunk_size":16,
      "action_path":"cookbooks/cosmos3/generator/action/assets/actions/av_traj_forward.json"
  }'

The script returns output_video.mp4. The end-to-end notebook is available at cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb.

Running Forward Dynamics via the vLLM-Omni API

If you prefer an OpenAI-compatible HTTP interface, use the vLLM-Omni server. Submit a multipart POST request to /v1/videos/sync with the same extra_params JSON string. The server streams the generated MP4 bytes directly.

curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=A robot arm lifts a cup." \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=35" \
  --form-string "guidance_scale=6.0" \
  --form-string 'extra_params={
      "action_mode":"forward_dynamics",
      "domain_name":"droid_orig_lerobot",
      "raw_action_dim":10,
      "action_chunk_size":16,
      "action_path":"cookbooks/cosmos3/generator/action/assets/actions/av_traj_forward.json"
  }' \
  -o forward_dynamics_output.mp4

The corresponding reference notebook is cookbooks/cosmos3/generator/action/run_fd_with_vllm.ipynb.

Preparing the Action Trajectory JSON

The action file must contain a flat array of raw per-step action values. The format matches the examples under cookbooks/cosmos3/generator/action/assets/actions/.

{
  "action": [
    [0.0, 0.01, -0.02, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0],
    [0.0, 0.02, -0.01, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0]
  ]
}

Each inner array length must equal raw_action_dim for the chosen domain_name. Consult the Input & Output table in README.md to map embodiments to their required dimensions.

Visualizing the Output

After generation, you can inspect the rollout with standard video tools. If you run the Jupyter notebook, the final cell typically calls:

export_to_video(result.video, "fd_demo.mp4", fps=24)

Alternatively, open the saved MP4 with ffplay or any compatible player.

Summary

  • Forward dynamics in Cosmos 3 is a generator mode that predicts video rollouts from action trajectories instead of input video.
  • The same Mixture-of-Transformers (MoT) and diffusion transformer (DM) architecture handles action tokens via the action_chunk conditioning path.
  • A valid request requires action_mode: "forward_dynamics" and an action_path pointing to a JSON array of raw actions.
  • You can execute inference through the Cosmos Framework Python script (cosmos_framework/scripts/inference) or the vLLM-Omni REST API (POST /v1/videos/sync).
  • Action dimensions and embodiment identifiers are standardized in README.md under the Input & Output reference table.

Frequently Asked Questions

What action dimension should I use for my robot or vehicle?

The required raw_action_dim depends on the embodiment. For example, the AV domain uses 9-D actions, the DROID-original LeRobot domain uses 10-D, and other humanoid embodiments may use 57-D. The Input & Output table in README.md lists the exact mapping of domain_name to dimensionality, so you should reference that table before building your JSON trajectory.

Do I need an input video to run forward dynamics?

No. Setting action_mode to "forward_dynamics" explicitly instructs the diffusion model to ignore any input video and use the supplied action trajectory to drive the rollout. You may still provide an initial image or text prompt for visual or semantic conditioning, but the temporal dynamics are governed entirely by the action tokens.

Can I combine text prompts with action conditioning?

Yes. The forward-dynamics pipeline supports multimodal conditioning. In both the Cosmos Framework script and the vLLM-Omni API, you can supply a prompt string alongside the action_path. The MoT backbone fuses text tokens, optional visual tokens, and action tokens (action_chunk) before the DM transformer performs diffusion sampling.

What is the performance difference between the Cosmos Framework and vLLM-Omni?

Both deployment paths expose the same forward-dynamics surface and call the same underlying model; the only difference is request routing. The Cosmos Framework runner is a local Python entry point ideal for interactive development, while vLLM-Omni offers an OpenAI-compatible HTTP server (POST /v1/videos/sync) better suited for production serving. Latency numbers for various hardware configurations are listed in inference_benchmarks.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →