# Supported Action Dimensions for Different Robot Embodiments in Cosmos 3 Action Models

> Explore supported action dimensions for Cosmos 3 robot embodiments. Discover how raw action dimensions from 9D to 57D enable single MoT models for diverse robotic applications.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: api-reference
- Published: 2026-06-14

---

**Cosmos 3 supports six distinct robot embodiments with raw action dimensions ranging from 9D to 57D, allowing a single Mixture-of-Transformers model to handle everything from camera motion and autonomous vehicles to single-arm, dual-arm, and humanoid robots through configurable action vectors.**

Cosmos 3 is a multimodal generative model that conditions video generation on robot actions through a unified token interface. When using the action models—whether for policy prediction, inverse dynamics, or forward dynamics—you must specify the embodiment type via `domain_name` and provide action vectors matching the **raw action dimension** defined for that specific robot platform. These dimensions are defined in the Action Conditioning table of the repository's [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md).

## Supported Embodiments and Action Dimensions

Cosmos 3 treats action as a generic multimodal token sequence, but each embodiment requires a specific vector length corresponding to its controllable degrees of freedom (DoFs). The supported **raw action dimensions** are:

- **Camera motion: 9D** — Global camera pose (3-D translation + 3-D rotation + 3-D intrinsics)
- **Autonomous vehicle: 9D** — Vehicle pose and steering (3-D translation + 3-D rotation + 3-D control signals)
- **Egocentric motion: 57D** — Full-body human-centric pose (joint angles, root translation, etc.)
- **Single-arm robot: 10D** — End-effector pose + gripper state (used by DROID, UR, Fractal, Bridge, UMI datasets)
- **Dual-arm robot: 20D** — Two single-arm vectors concatenated (dual DROID arms)
- **Humanoid robot: 29D** — Whole-body pose (torso, arms, legs, head) exemplified by the AgiBot robot

These values are defined in [[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)](https://github.com/NVIDIA/cosmos/blob/main/README.md) at lines 107-108 and must be supplied via the `extra_params` field alongside the `domain_name` that identifies the specific embodiment (e.g., `bridge_orig_lerobot`, `av`, `camera_pose`).

## How Action Dimensions Work in the Architecture

The model does not hard-code any specific robot API. Instead, it learns a mapping from a fixed-size action vector to the latent space using a unified **Mixture-of-Transformers** architecture. This design can ingest any action vector length as long as it matches the configured `raw_action_dim` for the request.

This flexibility enables the same model checkpoint to serve multiple robot platforms without retraining. The server validates that supplied action data matches the declared dimension, then processes it through the same transformer layers that handle vision, text, and audio tokens.

## Practical Implementation: Three Action Modes

When calling Cosmos 3 action endpoints, you must set `raw_action_dim` to the corresponding dimension from the table above and provide an action chunk (a sequence of vectors of that size). Below are implementations for each of the three action modes using the vLLM-Omni server format.

### Policy Mode

Policy mode predicts a robot action chunk from a text prompt and optional visual input. The response contains a generated video and a JSON array of predicted actions.

```json
POST http://localhost:8000/v1/videos
Content-Type: multipart/form-data

{
  "prompt": "Pick up the red block and place it on the blue platform.",
  "negative_prompt": "",
  "size": "1280x720",
  "num_frames": 189,
  "fps": 24,
  "extra_params": {
    "action_mode": "policy",
    "domain_name": "bridge_orig_lerobot",
    "raw_action_dim": 10,
    "action_chunk_size": 16
  }
}

```

The response returns a video and a JSON array of shape `[16, 10]` containing the predicted robot action vectors for the single-arm embodiment.

### Inverse Dynamics Mode

Inverse dynamics mode predicts the action sequence that produced a given video observation.

```json
POST http://localhost:8000/v1/videos
Content-Type: multipart/form-data

{
  "prompt": "Explain the robot motion.",
  "input_reference": "file:///data/robot_demo.mp4",
  "extra_params": {
    "action_mode": "inverse_dynamics",
    "domain_name": "av",
    "raw_action_dim": 9,
    "action_chunk_size": 30
  }
}

```

The server returns the original video unchanged plus a JSON array `[30, 9]` describing the vehicle's control trajectory.

### Forward Dynamics Mode

Forward dynamics mode generates future video conditioned on an initial observation and a predefined action sequence.

```json
POST http://localhost:8000/v1/videos
Content-Type: multipart/form-data

{
  "prompt": "",
  "input_reference": "file:///data/start_frame.png",
  "extra_params": {
    "action_mode": "forward_dynamics",
    "domain_name": "humanoid",
    "raw_action_dim": 29,
    "action_path": "/data/agi_action.json"
  }
}

```

Only video is returned; the server consumes the supplied action file (containing vectors of shape `[N, 29]`) to roll out future observations for the humanoid embodiment.

## Key Source Files

The following files in the NVIDIA Cosmos repository provide definitive specifications and examples for action dimensions:

- [[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)](https://github.com/NVIDIA/cosmos/blob/main/README.md) — Contains the Action Conditioning table listing supported dimensions for each embodiment (lines 107-108)
- [[`cookbooks/cosmos3/generator/action/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/README.md)](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/README.md) — Provides practical tutorials for the three action modes
- [`cookbooks/cosmos3/generator/action/run_fd_with_vllm.ipynb`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_fd_with_vllm.ipynb) — Example notebook demonstrating forward-dynamics inference
- [`cookbooks/cosmos3/generator/action/run_id_with_vllm.ipynb`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_id_with_vllm.ipynb) — Example notebook for inverse-dynamics inference
- [`cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.ipynb`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.ipynb) — Shows policy endpoint invocation via the Cosmos Framework CLI

## Summary

- Cosmos 3 supports **six embodiment types** with raw action dimensions ranging from **9D to 57D**, covering camera motion, vehicles, egocentric human motion, and robotic arms.
- The **Mixture-of-Transformers** architecture processes any action vector length as long as it matches the declared `raw_action_dim` for the specific embodiment.
- Action conditioning requires setting both `domain_name` (embodiment identifier) and `raw_action_dim` (vector size) in the request's `extra_params`.
- Three operational modes—**policy**, **inverse dynamics**, and **forward dynamics**—all use the same dimension specifications but differ in input/output behavior.

## Frequently Asked Questions

### What happens if I provide the wrong action dimension for an embodiment?

The server validates that the supplied action data matches the declared `raw_action_dim`. If the dimensions mismatch, the request will fail validation before processing through the transformer layers, as the model expects the specific vector length defined for that embodiment's kinematics.

### Can I use Cosmos 3 with custom robot embodiments not listed in the table?

No, the model checkpoint is trained on specific embodiments. While the architecture supports variable-length action vectors through the Mixture-of-Transformers design, you must use one of the six predefined `domain_name` values (such as `bridge_orig_lerobot`, `av`, or `humanoid`) with their corresponding dimensions (10D, 9D, or 29D respectively).

### Do dual-arm robots use a different dimension than single-arm robots?

Yes, dual-arm robots use **20D** action vectors, which represents two concatenated single-arm vectors (10D + 10D). This allows the model to control both arms simultaneously within a single action chunk, whereas single-arm robots only require 10D vectors for end-effector pose and gripper state.

### How do I select the correct action chunk size?

The `action_chunk_size` parameter determines the number of timesteps to predict or process, independent of the `raw_action_dim`. For policy mode, typical values are 16-32 timesteps, while inverse dynamics might use 30 or more depending on the video length. The action chunk size does not affect the dimensionality of individual action vectors, only the sequence length.