How to Configure Action Conditioning for Robot Embodiments in NVIDIA Cosmos
Action conditioning in NVIDIA Cosmos requires a JSON array with dimensionality specific to your embodiment: 10-D for single-arm robots like DROID and UMI, 9-D for autonomous vehicles, and 57-D for egocentric motion, processed in 16-frame chunks for forward-dynamics generation.
NVIDIA Cosmos enables action-conditioned video generation and forward-dynamics simulation across diverse robotic platforms. To generate accurate motion sequences, you must configure the action array to match the specific dimensionality requirements of your robot type, whether it's a single-arm manipulator, humanoid, or autonomous vehicle. The repository provides sample configurations and validation logic in the forward-dynamics cookbooks to ensure your action data aligns with the model's expectations.
Action Representation by Embodiment
Cosmos 3 accepts action-conditioned requests where the action field is a JSON array. The dimensionality must match the specific embodiment you are modeling, as defined in the repository's README.
| Embodiment | Action Dimensions | Typical Units | Normalization |
|---|---|---|---|
| Camera motion | 9-D (rotation + translation) | metres / radians | -1 → 1 |
| Autonomous vehicle | 9-D (ego pose per frame) | metres / radians | -1 → 1 |
| Egocentric motion | 57-D (full body pose) | metres / radians | -1 → 1 |
| Single-arm robot (DROID, UMI) | 10-D (end-effector pose + gripper) | metres / radians | -1 → 1 |
| Dual-arm robot | 20-D (two 10-D arms) | metres / radians | -1 → 1 |
| Humanoid robot (AgiBot) | 29-D | metres / radians | -1 → 1 |
Request Format and Chunking Mechanism
The forward-dynamics pipeline processes action data in specific chunks and requires a structured JSON payload.
Chunking Requirements
For forward-dynamics generation, Cosmos processes chunks of 16 consecutive frames. Your action file must contain 16 × N rows, where N represents the number of chunks you want to generate. Each row must have the exact dimensionality of the selected embodiment.
Conditioning Image Logic
Chunk 0 uses a static conditioning image shipped with the repository (such as those in cookbooks/cosmos3/generator/action/assets/images/). Subsequent chunks automatically use the last generated frame of the previous chunk as their conditioning image. This logic is implemented in the forward-dynamics notebooks run_fd_with_cosmos_framework.ipynb and run_fd_with_vllm.ipynb.
JSON Request Structure
Every generator call requires the following JSON structure:
{
"prompt": "<optional-text-prompt>",
"image": "<path-or-URL-to-conditioning-image>",
"action": [ <list-of-action-vectors> ],
"size": "<output-resolution-tier>",
"frame_rate": <int>,
"num_frames": <int>
}
The action field is the only component that varies between embodiments.
Embodiment-Specific Configuration
DROID Configuration
For the DROID single-arm robot, use 10-D action vectors representing end-effector pose and gripper state.
Sample file: cookbooks/cosmos3/generator/action/assets/actions/droid.json
Array format: [[x, y, z, rx, ry, rz, gripper_open], ...] where each inner list contains position, orientation, and gripper state in meters and radians.
The forward-dynamics notebooks automatically split this into 16-row chunks and validate dimensionality with assertions like assert all(len(row) == 10 for row in action).
UMI Configuration
The UMI (Universally Manipulation Interface) embodiment uses the same 10-D format as DROID but operates in the UMI coordinate system (right-handed, meters).
Sample file: cookbooks/cosmos3/generator/action/assets/actions/umi.json
Validation: The notebook run_fd_with_vllm.ipynb validates UMI actions using:
assert all(len(row) == umi_raw_action_dim for row in umi_action)
Autonomous Vehicle Configuration
Autonomous vehicles require 9-D ego pose trajectories.
Sample file: cookbooks/cosmos3/generator/action/assets/actions/av_traj_forward.json
Array format: [x, y, z, yaw, pitch, roll, speed, steering, throttle] per frame.
Implementation Examples
Python: DROID Forward Dynamics
This example from run_fd_with_cosmos_framework.ipynb demonstrates loading and validating DROID actions:
from cosmos_framework.scripts.inference import run_forward_dynamics
import json
# Load the DROID action JSON
with open("cookbooks/cosmos3/generator/action/assets/actions/droid.json") as f:
droid_action = json.load(f)
# Validate 10-D structure
assert all(len(row) == 10 for row in droid_action), "DROID action must be 10-D"
# Construct request
request = {
"prompt": "You are a robot manipulator solving a block-stacking task.",
"image": "cookbooks/cosmos3/generator/action/assets/images/droid.png",
"action": droid_action,
"size": "480p",
"frame_rate": 15,
"num_frames": len(droid_action),
}
# Run inference (automatically handles 16-frame chunking)
run_forward_dynamics(request, output_path="outputs/droid_forward.mp4")
Python: UMI with Environment Variables
From run_fd_with_vllm.ipynb, this snippet shows UMI-specific setup:
import json
import os
umi_path = "cookbooks/cosmos3/generator/action/assets/actions/umi.json"
with open(umi_path) as f:
umi_action = json.load(f)
# Ensure 10-D rows
assert all(len(row) == 10 for row in umi_action), "UMI action must be 10-D"
# Set environment vars for the notebook pipeline
os.environ["COSMOS3_UMI_FD_INPUT"] = str(umi_path)
os.environ["COSMOS3_UMI_FD_OUTPUT"] = "outputs/umi_forward"
Bash: CLI Invocation for Autonomous Vehicles
Use the Cosmos Framework CLI entry point to run action-conditioned generation:
cosmos_framework.scripts.inference \
--model cosmos3-nano \
--task forward_dynamics \
--embodiment av \
--action-file cookbooks/cosmos3/generator/action/assets/actions/av_traj_forward.json \
--image assets/av/first_frame.png \
--output av_forward.mp4
Summary
- Match dimensions exactly: Use 10-D arrays for DROID and UMI, 9-D for autonomous vehicles, and 57-D for egocentric motion.
- Chunk into 16-frame sequences: Ensure your action file contains multiples of 16 rows; the model processes forward-dynamics in 16-frame chunks.
- Validate before inference: Use assertions like
assert all(len(row) == 10 for row in action)to catch dimensionality mismatches early. - Use provided samples: Reference
cookbooks/cosmos3/generator/action/assets/actions/for correctly formatted JSON templates. - Leverage automatic chunking: The
cosmos_framework.scripts.inferenceCLI and notebooks handle the 16-frame folding automatically when you provide the full action array.
Frequently Asked Questions
What is the difference between DROID and UMI action conditioning?
Both DROID and UMI use 10-D action representations (end-effector position, orientation, and gripper state), but they differ in coordinate system conventions. DROID uses its native coordinate frame, while UMI uses a right-handed coordinate system specific to the Universally Manipulation Interface. Both require normalization to the [-1, 1] range, and both validate dimensionality using assert all(len(row) == 10 for row in action) in their respective notebooks.
How does the 16-frame chunking mechanism work in forward dynamics?
Cosmos processes forward-dynamics generation in fixed chunks of 16 consecutive frames. Your action file must contain 16 × N rows, where N is the number of chunks. The first chunk uses your provided conditioning image, while subsequent chunks automatically use the last generated frame from the previous chunk as their conditioning image. The run_fd_with_cosmos_framework.ipynb notebook implements this logic transparently.
Can I use custom action data instead of the provided samples?
Yes, you can use custom action data as long as you maintain the correct dimensionality for your embodiment and ensure the row count is a multiple of 16. Convert your poses into flat lists of normalized floats (-1 to 1), validate the length matches the expected dimension (e.g., 10 for single-arm robots), and pass the JSON array to the inference script via --action-file or directly in the Python API.
What normalization range should action values use?
According to the NVIDIA Cosmos source code, all action values must be normalized to the -1 to 1 range regardless of embodiment type. This applies to positions, orientations, gripper states, and vehicle controls. The sample files in cookbooks/cosmos3/generator/action/assets/actions/ demonstrate this normalization for each supported robot type.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →