Implementing Forward Dynamics with Cosmos 3 for Robotics Simulation: A Complete Technical Guide
Cosmos 3 provides a dedicated forward_dynamics action mode that accepts an initial image paired with robot-action tokens to predict future visual observations as a video, enabling physics-aware robotics simulation through either the Cosmos Framework or a vLLM-Omni server.
Cosmos 3, NVIDIA's open-source world foundation model, natively supports forward dynamics for robotics by conditioning its diffusion transformer on action tokens rather than text prompts. According to the NVIDIA/cosmos repository, this capability allows the model to forecast how a robot's movements will alter the visual environment over time, outputting a complete MP4 video of the predicted future state. This implementation guide references the specific source files, API contracts, and code patterns found in the official cookbooks and README documentation.
Understanding Forward Dynamics in Cosmos 3
Forward dynamics in Cosmos 3 operates as a generator task within the model's unified architecture. Unlike text-to-video generation, this mode consumes an initial visual frame (JPG/PNG) and a sequence of robot actions (JSON array of floats) to produce a video visualizing the physically plausible outcome of those actions.
As documented in the project README (lines 354–355), the model processes these inputs by concatenating image tokens with action tokens before running the diffusion pipeline. The output is a full video tensor post-processed into MP4 format via export_to_video (Diffusers) or streamed directly from the vLLM-Omni server.
Architecture Overview
Model Surface and API
The forward dynamics capability is exposed through the action_mode: "forward_dynamics" parameter in the vLLM-Omni API. The model requires:
- Input: An image token stream followed by an action chunk (JSON array of joint-space or 9-D ego-motion values)
- Output: MP4 video (optionally with sound) representing the predicted future state
The supported action dimensions vary by embodiment—for example, 10D for single-arm DROID manipulation or 57D for egocentric motion, as listed in the README table (lines 548–555).
Cosmos Framework Pipeline
The cosmos_framework.scripts.inference module serves as the primary entry point for local execution. This script wraps the Cosmos3OmniPipeline (Diffusers) and handles:
- Loading model checkpoints (
Cosmos3-NanoorCosmos3-Super) - Assembling the request payload with
model_mode="forward_dynamics" - Running the diffusion pipeline in generator mode
vLLM-Omni Server
For production deployments, the vLLM-Omni server exposes OpenAI-compatible endpoints:
/v1/videos/sync– synchronous generation/v1/videos– asynchronous generation
The server reads action files from mounted paths (--allowed-local-media-path) and processes extra_params containing the action mode and domain specifications.
Implementation Workflows
Method 1: Cosmos Framework (Python Script)
The run_fd_with_cosmos_framework.ipynb notebook in cookbooks/cosmos3/generator/action/ demonstrates the complete setup. First, configure the environment variables:
import os
os.environ["COSMOS3_REPO"] = "/path/to/cosmos-framework"
os.environ["COSMOS3_UV_GROUP"] = "cu130-train"
os.environ["COSMOS3_OUTPUT_ROOT"] = "/tmp/cosmos_fd_outputs"
os.environ["HF_HOME"] = "/tmp/hf_cache"
Then launch the inference entry point. This command reads the spec file action_forward_dynamics_robotics_custom.jsonl and writes the output to <COSMOS3_OUTPUT_ROOT>/action_forward_dynamics_robotics_custom/<run>/vision.mp4:
python -m cosmos_framework.scripts.inference
The script cosmos_framework/scripts/inference.py (in the separate cosmos-framework repository) parses the model_mode="forward_dynamics" flag and invokes the Cosmos3OmniPipeline with the appropriate conditioning tokens.
Method 2: vLLM-Omni API (cURL)
To query a running vLLM-Omni server (started via vllm serve nvidia/Cosmos3-Nano --omni --model-class-name Cosmos3OmniDiffusersPipeline), send a POST request to the synchronous endpoint:
curl -sS -X POST http://localhost:8000/v1/videos/sync \
--form-string "prompt=Robot arm moves a cube to a shelf." \
--form-string "negative_prompt=blur, low-quality" \
--form-string "size=1280x720" \
--form-string "num_frames=189" \
--form-string "fps=24" \
--form-string "num_inference_steps=35" \
--form-string "guidance_scale=6.0" \
--form-string "seed=42" \
--form-string 'extra_params={"action_mode":"forward_dynamics","domain_name":"bridge_orig_lerobot","raw_action_dim":10,"action_chunk_size":5}' \
-o forward_dynamics_output.mp4
The extra_params field must specify the action_mode, domain_name (e.g., bridge_orig_lerobot, av, or camera_pose), raw_action_dim, and action_chunk_size to match the expected input tensor shape.
Method 3: OpenAI-Compatible Python Client
For programmatic access, use the OpenAI Python client against the vLLM-Omni server:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
response = client.videos.sync.create(
prompt="A mobile robot pushes a box across a warehouse floor.",
size="1280x720",
num_frames=189,
fps=24,
num_inference_steps=35,
guidance_scale=6.0,
seed=123,
extra_params={
"action_mode": "forward_dynamics",
"domain_name": "av",
"raw_action_dim": 9,
"action_chunk_size": 8,
},
)
with open("fd_video.mp4", "wb") as f:
f.write(response.content)
The run_fd_with_vllm.ipynb notebook provides the complete reference implementation for this approach, including visualization via IPython.display.Video.
Preparing Action Specifications
Forward dynamics requires a JSONL spec file where each line contains:
{ "action": [0.1, -0.2, ...], "prompt": "Robot description", "image_path": "frame.jpg" }
Action dimensions vary by domain:
- 10D: Single-arm DROID manipulation
- 57D: Egocentric motion (camera pose)
- 9D: Autonomous vehicle (AV) ego-motion
The pipeline converts these float arrays into action tokens that are concatenated with the visual token stream before diffusion. Ensure the action_chunk_size matches the sequence length expected by the specific model checkpoint.
Summary
- Forward dynamics in Cosmos 3 uses the
action_mode="forward_dynamics"parameter to predict future visual states from initial images and robot actions. - Two primary interfaces exist: the
cosmos_framework.scripts.inferenceCLI for local Diffusers-based execution, and the vLLM-Omni server (/v1/videos/sync) for production API deployment. - Input requirements include a JSONL file with action arrays (dimensions vary: 10D for DROID, 57D for egocentric, 9D for AV) and a reference image.
- Key source files include
cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynbfor framework usage andcookbooks/cosmos3/generator/action/run_fd_with_vllm.ipynbfor API usage. - Output is always an MP4 video generated by the
Cosmos3OmniPipelinevia diffusion, with unified architecture shared across text-to-video and image-to-video modes.
Frequently Asked Questions
What action dimensions does Cosmos 3 support for forward dynamics?
According to the README (lines 354–355), Cosmos 3 supports varying action dimensions depending on the robot embodiment. Single-arm DROID manipulation uses 10D joint-space controls, egocentric motion (camera pose) uses 57D parameters, and autonomous vehicle (AV) domains typically use 9D ego-motion values. The raw_action_dim parameter in your API request must match these specifications for the chosen domain_name.
How do I choose between the Cosmos Framework and vLLM-Omni for implementation?
Use the Cosmos Framework (cosmos_framework.scripts.inference) for local development, research experimentation, or when you need direct access to the Diffusers pipeline and checkpoints. Deploy vLLM-Omni when you require a production-grade, OpenAI-compatible API server with synchronous (/v1/videos/sync) and asynchronous (/v1/videos) endpoints for serving multiple clients or integrating with existing robotics stacks.
Can I use custom robot embodiments not listed in the predefined domain names?
While the README lists standard domains like bridge_orig_lerobot, av, and camera_pose, the architecture supports custom embodiments provided you correctly specify the raw_action_dim and format your action JSONL files to match the expected token sequence. The model's flexibility stems from its unified diffusion transformer that processes arbitrary action token streams, though performance may vary on out-of-distribution robot morphologies not represented in the training data.
What is the difference between forward dynamics and standard text-to-video generation in Cosmos 3?
Both modes share the same diffusion transformer backbone and generate video through iterative denoising. However, forward dynamics conditions the generation on action tokens derived from robot joint states or ego-motion vectors, while text-to-video uses text tokens from prompts. This architectural distinction allows forward dynamics to maintain physical consistency with the robot's actuator constraints and the initial world state, producing predictions that respect the laws of physics for the specific embodiment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →