Supported Action Dimensions for Robot Embodiments in NVIDIA Cosmos
NVIDIA Cosmos 3 supports robot embodiments with raw action dimensions ranging from 6-DOF for camera motion to 10-DOF for UMI robots, configured via the extra_params.domain_name parameter in action-related API requests.
NVIDIA Cosmos 3 treats every robot embodiment as a domain that defines the expected shape of the action vector. When calling action endpoints such as policy, inverse_dynamics, or forward_dynamics, clients must specify the embodiment through extra_params.domain_name and provide action payloads matching the corresponding raw_action_dim defined for that domain.
Supported Robot Embodiments and Action Dimensions
The repository defines specific dimensionality for each embodiment type in the request schema. Each domain expects a precise raw_action_dim that the server validates against incoming requests.
Camera Motion (Egocentric)
For egocentric camera motion, set domain_name to camera_pose. This embodiment expects a 6-dimensional action vector representing 3-D translation and 3-D rotation (6 DOF). According to the README documentation, this domain is explicitly designed for first-person perspective camera control [README.md†L356-L357].
Standard Robotic Arm
The bridge_orig_lerobot domain represents standard robotic arms with 9-dimensional action vectors. As shown in run_id_with_vllm.ipynb, this configuration handles 3-D end-effector position, 3-D orientation, and gripper state [run_id_with_vllm.ipynb†L250]. The raw action dimension is set to 9 to accommodate these combined degrees of freedom.
Autonomous Vehicle
For autonomous vehicle control, use the av domain with 9-dimensional actions. The same notebook that configures robotic arms sets raw_action_dim: 9 for the av domain, representing 3-D vehicle pose, 3-D velocity, and steering/brake/throttle signals [run_id_with_vllm.ipynb†L132-L133][run_id_with_vllm.ipynb†L250].
Mobile Droid
The droid_lerobot domain supports mobile droid robots with 9-dimensional action vectors. Following the pattern of other lerobot variants documented in the README examples, this domain handles 3-D base pose, 3-D arm pose, and gripper state [README.md†L356-L357].
UMI Dexterous Manipulation
The umi domain requires 10-dimensional raw action vectors, the highest dimensionality currently supported. Notebooks such as run_fd_with_vllm.ipynb explicitly define umi_raw_action_dim = 10 and enforce that each action row maintains this 10-dimensional structure [run_fd_with_vllm.ipynb†L945-L953]. This 10-D raw vector is subsequently converted to a 16-D "action chunk" required by the Cosmos Framework [run_fd_with_cosmos_framework.ipynb†L1089-L1097].
Configuring Action Dimensions in API Requests
When invoking action endpoints, the extra_params object must specify both the domain and the expected dimensionality. The server validates that incoming action payloads match the raw_action_dim for the specified domain_name before processing.
For forward-dynamics requests, clients must also provide an action_path pointing to a file containing action rows with the appropriate dimensionality for the selected embodiment.
Implementation Examples
The following examples demonstrate how to configure requests for different embodiments according to the source code:
Standard Robotic Arm (9-DOF):
import requests, json
extra = {
"action_mode": "policy",
"domain_name": "bridge_orig_lerobot",
"raw_action_dim": 9,
"action_chunk_size": 60,
"guardrails": True,
}
files = {
"prompt": (None, "A small warehouse robot lifts a box."),
"extra_params": (None, json.dumps(extra)),
}
resp = requests.post(
"http://localhost:8000/v1/videos",
files=files,
)
print(resp.json())
Egocentric Camera Motion (6-DOF):
extra = {
"action_mode": "policy",
"domain_name": "camera_pose",
"raw_action_dim": 6,
"action_chunk_size": 30,
}
# Equivalent curl command:
# curl -X POST http://localhost:8000/v1/videos/sync \
# --form-string "prompt=First-person view of a person walking forward." \
# --form-string 'extra_params={"action_mode":"policy","domain_name":"camera_pose","raw_action_dim":6,"action_chunk_size":30}'
Summary
- Camera Motion: Use
camera_posedomain with 6 dimensions (3-D translation + 3-D rotation). - Standard Robotic Arms: Use
bridge_orig_lerobotdomain with 9 dimensions (position + orientation + gripper). - Autonomous Vehicles: Use
avdomain with 9 dimensions (pose + velocity + controls). - Mobile Droids: Use
droid_lerobotdomain with 9 dimensions (base pose + arm pose + gripper). - UMI Robots: Use
umidomain with 10 dimensions (raw 10-D converted to 16-D action chunks). - Configuration requires setting both
extra_params.domain_nameandextra_params.raw_action_dimto match the embodiment specification.
Frequently Asked Questions
What happens if I provide the wrong action dimension for a robot embodiment?
The server expects action payloads that match the raw_action_dim specified for the selected domain_name. Providing incorrect dimensions will cause validation errors or malformed processing, as the Cosmos Framework strictly enforces the vector shape defined in extra_params.
Can I use action dimensions other than those documented for each domain?
No, the dimensions are fixed per domain as implemented in the source code. For example, camera_pose strictly requires 6 dimensions according to the README [README.md†L356-L357], while umi specifically requires 10 dimensions [run_fd_with_vllm.ipynb†L945-L953]. Deviating from these values will result in compatibility issues with the pretrained models.
How does the UMI robot's 10-dimensional action differ from other embodiments?
The UMI robot uses a 10-dimensional raw action vector that is internally converted to a 16-dimensional "action chunk" by the Cosmos Framework [run_fd_with_cosmos_framework.ipynb†L1089-L1097]. This differs from other domains where the raw_action_dim typically matches the final action space size, making UMI unique in requiring this additional post-processing step.
Where are the action dimension constants defined in the codebase?
The action dimensions are primarily configured in example notebooks: run_id_with_vllm.ipynb sets dimensions for av and bridge_orig_lerobot [run_id_with_vllm.ipynb†L132-L133][run_id_with_vllm.ipynb†L250], while run_fd_with_vllm.ipynb and run_fd_with_cosmos_framework.ipynb define the UMI-specific 10-dimensional requirement [run_fd_with_vllm.ipynb†L945-L953][run_fd_with_cosmos_framework.ipynb†L1089-L1097]. The README documents the schema at lines 356-357 [README.md†L356-L357].
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →