How to Use Forward Dynamics for Robot Trajectory Prediction with Action Conditioning in Cosmos 3
Cosmos 3 provides a Generator-mode "forward-dynamics" surface that predicts future visual observations of a robot from an initial image and a sequence of action commands, enabling physics-aware trajectory prediction without explicit simulators.
The NVIDIA/cosmos repository implements Cosmos 3, a world foundation model featuring a dedicated forward-dynamics pipeline for embodied AI. This mode allows you to perform forward dynamics for robot trajectory prediction by conditioning video generation on action sequences, effectively learning a world model that propagates robot actions through a diffusion-based generation process.
Architecture of Forward Dynamics
Cosmos 3's forward-dynamics implementation builds on the same Mixture-of-Transformers (MoT) architecture used throughout the model family. In this mode, the system processes a conditioning image alongside a domain-specific action tensor to autoregressively generate future visual states.
Multimodal Tokenization and MoT Backbone
The pipeline begins with a Multimodal Tokenizer that converts images, videos, and action vectors into a shared token space. The MoT Transformer then processes these tokens using causal self-attention for reasoning and full attention for diffusion generation. As implemented in the Cosmos framework, this transformer handles both text-only Reasoner requests and multimodal Generator requests, including the forward-dynamics mode.
3-D mRoPE Embedding
Spatial-temporal coherence is maintained through 3-D mRoPE (multimodal Rotary Position Embedding). This embedding provides position information that works across vision, audio, and action modalities, enabling the model to maintain coherent geometry when actions move the camera or robot end-effector through the scene.
Action Conditioning Mechanism
Action vectors are concatenated to the token stream as a dedicated "action block." The model learns to propagate the effect of each action step through the diffusion process, effectively learning a forward dynamics model of the world. When you set model_mode="forward_dynamics" in the request JSON, the diffusion transformer predicts only the visual stream without audio or additional actions, making it ideal for trajectory prediction and robot simulation.
Preparing Input Data
Forward dynamics requires two inputs: a conditioning image representing the initial observation, and an action tensor describing the trajectory to execute.
Conditioning Image
Prepare a single RGB frame that represents the robot's initial observation. This image should capture the workspace from the robot's camera perspective (ego-view) or a fixed external view, depending on your domain configuration.
Action Tensor Format
The action tensor must be a JSON array of shape (T, D), where T is the number of action steps and D is the domain-specific dimensionality. Cosmos 3 supports various robot embodiments:
- 10-D for single-arm robots (e.g., DROID)
- 20-D for dual-arm configurations
- 57-D for egocentric motion scenarios
Action utilities for converting between pose representations reside in cosmos_framework/data/vfm/action/pose_utils.py. The framework includes helpers like pose_abs_to_rel and pose_rel_to_abs for transforming absolute camera poses to the relative 9-D format required by autonomous vehicle domains, with analogous helpers available for robot action formats.
Building the JSONL Specification
Cosmos 3 consumes forward-dynamics requests via a JSONL specification file. Each line describes one forward-dynamics run, specifying the image path, action path, domain name, and generation parameters.
The complete workflow is documented in cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb, which provides reference implementations for robotics, autonomous vehicle, and UMI domains.
import json
from pathlib import Path
vision_path = Path("cookbooks/cosmos3/generator/action/assets/images/robot_initial.jpg")
action_trajectory = [
[0.05, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0],
[0.05, 0.0, 0.0, 0.1, 0.0, 0.0, 0.0, 0.0, 0.0],
]
action_path = Path("tmp/robot_action.json")
action_path.write_text(json.dumps(action_trajectory, indent=2))
record = {
"model_mode": "forward_dynamics",
"domain_name": "droid_lerobot",
"fps": 10,
"image_size": 480,
"view_point": "ego_view",
"action_chunk_size": len(action_trajectory),
"action_path": str(action_path),
"vision_path": str(vision_path),
"prompt": "A robot arm moving a cup on a table.",
"seed": 0,
"name": "robot_fd_example"
}
spec_path = Path("tmp/robot_fd_spec.jsonl")
spec_path.write_text(json.dumps(record) + "\n")
Running Forward Dynamics Inference
Execute the forward-dynamics generation using the CLI entry point cosmos_framework/scripts/inference.py, which parses the JSONL spec, loads the Cosmos 3 checkpoint, and streams the generated video to your output directory.
Command-Line Execution
export COSMOS3_REPO=$(pwd)/packages/cosmos3
export COSMOS3_CHECKPOINT_PATH=Cosmos3-Nano
python -m cosmos_framework.scripts.inference \
-i tmp/robot_fd_spec.jsonl \
-o outputs/robot_fd \
--checkpoint-path $COSMOS3_CHECKPOINT_PATH \
--image_size 480 \
--seed 0 \
--benchmark
The generated video appears under outputs/robot_fd/robot_fd_example/vision.mp4. Alternatively, you can stream results from a vLLM-Omni server via the /v1/videos/sync endpoint for production deployments.
Visualization
Preview the generated trajectory video directly in a Jupyter notebook:
import imageio_ffmpeg, subprocess
from IPython.display import Video, display
from pathlib import Path
FFMPEG = imageio_ffmpeg.get_ffmpeg_exe()
src = Path("outputs/robot_fd/robot_fd_example/vision.mp4")
preview = src.with_name(src.stem + "_preview.mp4")
subprocess.run([
FFMPEG, "-y", "-i", str(src),
"-c:v", "libx264", "-crf", "28",
"-preset", "veryfast", "-an", "-pix_fmt", "yuv420p",
str(preview)
], check=True)
display(Video(str(preview), embed=True))
Summary
- Forward dynamics in Cosmos 3 generates future visual observations by conditioning on action sequences, implementing a learned world model for robot trajectory prediction.
- The Mixture-of-Transformers architecture processes domain-specific action tensors (10-D, 20-D, or 57-D) alongside initial observations to produce physics-consistent video sequences.
- Input preparation involves creating a JSONL specification that references your conditioning image and action trajectory file.
- Use
python -m cosmos_framework.scripts.inferenceto execute generation, or integrate with the vLLM-Omni API for scalable deployment. - Action utilities in
cosmos_framework/data/vfm/action/pose_utils.pyhandle coordinate transformations between absolute and relative pose representations.
Frequently Asked Questions
What action space dimensions does Cosmos 3 support for robot trajectory prediction?
Cosmos 3 supports domain-specific action dimensions defined in the model configuration. Common schemas include 10-D for single-arm robots (such as DROID), 20-D for dual-arm configurations, and 57-D for egocentric motion. The domain_name field in your JSONL spec selects the appropriate tokenizer and action block format for your specific embodiment.
How do I convert between absolute and relative pose representations for action conditioning?
The framework provides utility functions in cosmos_framework/data/vfm/action/pose_utils.py for coordinate transformations. Use pose_abs_to_rel to convert absolute camera poses to the relative 9-D format required by autonomous vehicle domains, or apply the analogous robotics-specific helpers to normalize your action trajectories before serialization to JSON.
Can forward dynamics generate audio or only visual predictions?
In forward_dynamics mode, the diffusion transformer predicts only the visual stream. This mode explicitly excludes audio generation and does not output additional action tokens, optimizing the model specifically for visual trajectory prediction and robot simulation tasks.
What hardware requirements are needed for real-time forward dynamics inference?
Forward-dynamics inference requires NVIDIA GPUs with sufficient VRAM to load the Cosmos 3 checkpoints. The Cosmos 3-Nano model provides the fastest inference for development and testing, while larger variants require proportionally more GPU memory. Refer to inference_benchmarks.md in the repository root for latency metrics and recommended GPU configurations (such as H100 or A100 instances) for different throughput requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →