Supported Action Dimensions for Different Robot Embodiments in Cosmos 3 Action Models
Cosmos 3 supports six distinct robot embodiments with raw action dimensions ranging from 9D to 57D, allowing a single Mixture-of-Transformers model to handle everything from camera motion and autonomous vehicles to single-arm, dual-arm, and humanoid robots through configurable action vectors.
Cosmos 3 is a multimodal generative model that conditions video generation on robot actions through a unified token interface. When using the action models—whether for policy prediction, inverse dynamics, or forward dynamics—you must specify the embodiment type via domain_name and provide action vectors matching the raw action dimension defined for that specific robot platform. These dimensions are defined in the Action Conditioning table of the repository's README.md.
Supported Embodiments and Action Dimensions
Cosmos 3 treats action as a generic multimodal token sequence, but each embodiment requires a specific vector length corresponding to its controllable degrees of freedom (DoFs). The supported raw action dimensions are:
- Camera motion: 9D — Global camera pose (3-D translation + 3-D rotation + 3-D intrinsics)
- Autonomous vehicle: 9D — Vehicle pose and steering (3-D translation + 3-D rotation + 3-D control signals)
- Egocentric motion: 57D — Full-body human-centric pose (joint angles, root translation, etc.)
- Single-arm robot: 10D — End-effector pose + gripper state (used by DROID, UR, Fractal, Bridge, UMI datasets)
- Dual-arm robot: 20D — Two single-arm vectors concatenated (dual DROID arms)
- Humanoid robot: 29D — Whole-body pose (torso, arms, legs, head) exemplified by the AgiBot robot
These values are defined in [README.md](https://github.com/NVIDIA/cosmos/blob/main/README.md) at lines 107-108 and must be supplied via the extra_params field alongside the domain_name that identifies the specific embodiment (e.g., bridge_orig_lerobot, av, camera_pose).
How Action Dimensions Work in the Architecture
The model does not hard-code any specific robot API. Instead, it learns a mapping from a fixed-size action vector to the latent space using a unified Mixture-of-Transformers architecture. This design can ingest any action vector length as long as it matches the configured raw_action_dim for the request.
This flexibility enables the same model checkpoint to serve multiple robot platforms without retraining. The server validates that supplied action data matches the declared dimension, then processes it through the same transformer layers that handle vision, text, and audio tokens.
Practical Implementation: Three Action Modes
When calling Cosmos 3 action endpoints, you must set raw_action_dim to the corresponding dimension from the table above and provide an action chunk (a sequence of vectors of that size). Below are implementations for each of the three action modes using the vLLM-Omni server format.
Policy Mode
Policy mode predicts a robot action chunk from a text prompt and optional visual input. The response contains a generated video and a JSON array of predicted actions.
POST http://localhost:8000/v1/videos
Content-Type: multipart/form-data
{
"prompt": "Pick up the red block and place it on the blue platform.",
"negative_prompt": "",
"size": "1280x720",
"num_frames": 189,
"fps": 24,
"extra_params": {
"action_mode": "policy",
"domain_name": "bridge_orig_lerobot",
"raw_action_dim": 10,
"action_chunk_size": 16
}
}
The response returns a video and a JSON array of shape [16, 10] containing the predicted robot action vectors for the single-arm embodiment.
Inverse Dynamics Mode
Inverse dynamics mode predicts the action sequence that produced a given video observation.
POST http://localhost:8000/v1/videos
Content-Type: multipart/form-data
{
"prompt": "Explain the robot motion.",
"input_reference": "file:///data/robot_demo.mp4",
"extra_params": {
"action_mode": "inverse_dynamics",
"domain_name": "av",
"raw_action_dim": 9,
"action_chunk_size": 30
}
}
The server returns the original video unchanged plus a JSON array [30, 9] describing the vehicle's control trajectory.
Forward Dynamics Mode
Forward dynamics mode generates future video conditioned on an initial observation and a predefined action sequence.
POST http://localhost:8000/v1/videos
Content-Type: multipart/form-data
{
"prompt": "",
"input_reference": "file:///data/start_frame.png",
"extra_params": {
"action_mode": "forward_dynamics",
"domain_name": "humanoid",
"raw_action_dim": 29,
"action_path": "/data/agi_action.json"
}
}
Only video is returned; the server consumes the supplied action file (containing vectors of shape [N, 29]) to roll out future observations for the humanoid embodiment.
Key Source Files
The following files in the NVIDIA Cosmos repository provide definitive specifications and examples for action dimensions:
- [
README.md](https://github.com/NVIDIA/cosmos/blob/main/README.md) — Contains the Action Conditioning table listing supported dimensions for each embodiment (lines 107-108) - [
cookbooks/cosmos3/generator/action/README.md](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/README.md) — Provides practical tutorials for the three action modes cookbooks/cosmos3/generator/action/run_fd_with_vllm.ipynb— Example notebook demonstrating forward-dynamics inferencecookbooks/cosmos3/generator/action/run_id_with_vllm.ipynb— Example notebook for inverse-dynamics inferencecookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.ipynb— Shows policy endpoint invocation via the Cosmos Framework CLI
Summary
- Cosmos 3 supports six embodiment types with raw action dimensions ranging from 9D to 57D, covering camera motion, vehicles, egocentric human motion, and robotic arms.
- The Mixture-of-Transformers architecture processes any action vector length as long as it matches the declared
raw_action_dimfor the specific embodiment. - Action conditioning requires setting both
domain_name(embodiment identifier) andraw_action_dim(vector size) in the request'sextra_params. - Three operational modes—policy, inverse dynamics, and forward dynamics—all use the same dimension specifications but differ in input/output behavior.
Frequently Asked Questions
What happens if I provide the wrong action dimension for an embodiment?
The server validates that the supplied action data matches the declared raw_action_dim. If the dimensions mismatch, the request will fail validation before processing through the transformer layers, as the model expects the specific vector length defined for that embodiment's kinematics.
Can I use Cosmos 3 with custom robot embodiments not listed in the table?
No, the model checkpoint is trained on specific embodiments. While the architecture supports variable-length action vectors through the Mixture-of-Transformers design, you must use one of the six predefined domain_name values (such as bridge_orig_lerobot, av, or humanoid) with their corresponding dimensions (10D, 9D, or 29D respectively).
Do dual-arm robots use a different dimension than single-arm robots?
Yes, dual-arm robots use 20D action vectors, which represents two concatenated single-arm vectors (10D + 10D). This allows the model to control both arms simultaneously within a single action chunk, whereas single-arm robots only require 10D vectors for end-effector pose and gripper state.
How do I select the correct action chunk size?
The action_chunk_size parameter determines the number of timesteps to predict or process, independent of the raw_action_dim. For policy mode, typical values are 16-32 timesteps, while inverse dynamics might use 30 or more depending on the video length. The action chunk size does not affect the dimensionality of individual action vectors, only the sequence length.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →