# Supported Action Dimensions for Robot Embodiments in NVIDIA Cosmos

> Explore supported action dimensions for robot embodiments in NVIDIA Cosmos. Discover 6-DOF camera motion to 10-DOF UMI robots for your projects.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: api-reference
- Published: 2026-06-06

---

**NVIDIA Cosmos 3 supports robot embodiments with raw action dimensions ranging from 6-DOF for camera motion to 10-DOF for UMI robots, configured via the `extra_params.domain_name` parameter in action-related API requests.**

NVIDIA Cosmos 3 treats every robot embodiment as a **domain** that defines the expected shape of the action vector. When calling action endpoints such as `policy`, `inverse_dynamics`, or `forward_dynamics`, clients must specify the embodiment through `extra_params.domain_name` and provide action payloads matching the corresponding `raw_action_dim` defined for that domain.

## Supported Robot Embodiments and Action Dimensions

The repository defines specific dimensionality for each embodiment type in the request schema. Each domain expects a precise `raw_action_dim` that the server validates against incoming requests.

### Camera Motion (Egocentric)

For **egocentric camera motion**, set `domain_name` to `camera_pose`. This embodiment expects a **6-dimensional** action vector representing 3-D translation and 3-D rotation (6 DOF). According to the README documentation, this domain is explicitly designed for first-person perspective camera control [README.md†L356-L357].

### Standard Robotic Arm

The **bridge_orig_lerobot** domain represents standard robotic arms with **9-dimensional** action vectors. As shown in `run_id_with_vllm.ipynb`, this configuration handles 3-D end-effector position, 3-D orientation, and gripper state [run_id_with_vllm.ipynb†L250]. The raw action dimension is set to 9 to accommodate these combined degrees of freedom.

### Autonomous Vehicle

For autonomous vehicle control, use the **av** domain with **9-dimensional** actions. The same notebook that configures robotic arms sets `raw_action_dim: 9` for the `av` domain, representing 3-D vehicle pose, 3-D velocity, and steering/brake/throttle signals [run_id_with_vllm.ipynb†L132-L133][run_id_with_vllm.ipynb†L250].

### Mobile Droid

The **droid_lerobot** domain supports mobile droid robots with **9-dimensional** action vectors. Following the pattern of other lerobot variants documented in the README examples, this domain handles 3-D base pose, 3-D arm pose, and gripper state [README.md†L356-L357].

### UMI Dexterous Manipulation

The **umi** domain requires **10-dimensional** raw action vectors, the highest dimensionality currently supported. Notebooks such as `run_fd_with_vllm.ipynb` explicitly define `umi_raw_action_dim = 10` and enforce that each action row maintains this 10-dimensional structure [run_fd_with_vllm.ipynb†L945-L953]. This 10-D raw vector is subsequently converted to a 16-D "action chunk" required by the Cosmos Framework [run_fd_with_cosmos_framework.ipynb†L1089-L1097].

## Configuring Action Dimensions in API Requests

When invoking action endpoints, the `extra_params` object must specify both the domain and the expected dimensionality. The server validates that incoming action payloads match the `raw_action_dim` for the specified `domain_name` before processing.

For **forward-dynamics** requests, clients must also provide an `action_path` pointing to a file containing action rows with the appropriate dimensionality for the selected embodiment.

## Implementation Examples

The following examples demonstrate how to configure requests for different embodiments according to the source code:

**Standard Robotic Arm (9-DOF):**

```python
import requests, json

extra = {
    "action_mode": "policy",
    "domain_name": "bridge_orig_lerobot",
    "raw_action_dim": 9,
    "action_chunk_size": 60,
    "guardrails": True,
}
files = {
    "prompt": (None, "A small warehouse robot lifts a box."),
    "extra_params": (None, json.dumps(extra)),
}
resp = requests.post(
    "http://localhost:8000/v1/videos",
    files=files,
)
print(resp.json())

```

**Egocentric Camera Motion (6-DOF):**

```python
extra = {
    "action_mode": "policy",
    "domain_name": "camera_pose",
    "raw_action_dim": 6,
    "action_chunk_size": 30,
}

# Equivalent curl command:

# curl -X POST http://localhost:8000/v1/videos/sync \

#   --form-string "prompt=First-person view of a person walking forward." \

#   --form-string 'extra_params={"action_mode":"policy","domain_name":"camera_pose","raw_action_dim":6,"action_chunk_size":30}'

```

## Summary

- **Camera Motion**: Use `camera_pose` domain with **6** dimensions (3-D translation + 3-D rotation).
- **Standard Robotic Arms**: Use `bridge_orig_lerobot` domain with **9** dimensions (position + orientation + gripper).
- **Autonomous Vehicles**: Use `av` domain with **9** dimensions (pose + velocity + controls).
- **Mobile Droids**: Use `droid_lerobot` domain with **9** dimensions (base pose + arm pose + gripper).
- **UMI Robots**: Use `umi` domain with **10** dimensions (raw 10-D converted to 16-D action chunks).
- Configuration requires setting both `extra_params.domain_name` and `extra_params.raw_action_dim` to match the embodiment specification.

## Frequently Asked Questions

### What happens if I provide the wrong action dimension for a robot embodiment?

The server expects action payloads that match the `raw_action_dim` specified for the selected `domain_name`. Providing incorrect dimensions will cause validation errors or malformed processing, as the Cosmos Framework strictly enforces the vector shape defined in `extra_params`.

### Can I use action dimensions other than those documented for each domain?

No, the dimensions are fixed per domain as implemented in the source code. For example, `camera_pose` strictly requires 6 dimensions according to the README [README.md†L356-L357], while `umi` specifically requires 10 dimensions [run_fd_with_vllm.ipynb†L945-L953]. Deviating from these values will result in compatibility issues with the pretrained models.

### How does the UMI robot's 10-dimensional action differ from other embodiments?

The UMI robot uses a 10-dimensional raw action vector that is internally converted to a 16-dimensional "action chunk" by the Cosmos Framework [run_fd_with_cosmos_framework.ipynb†L1089-L1097]. This differs from other domains where the `raw_action_dim` typically matches the final action space size, making UMI unique in requiring this additional post-processing step.

### Where are the action dimension constants defined in the codebase?

The action dimensions are primarily configured in example notebooks: `run_id_with_vllm.ipynb` sets dimensions for `av` and `bridge_orig_lerobot` [run_id_with_vllm.ipynb†L132-L133][run_id_with_vllm.ipynb†L250], while `run_fd_with_vllm.ipynb` and `run_fd_with_cosmos_framework.ipynb` define the UMI-specific 10-dimensional requirement [run_fd_with_vllm.ipynb†L945-L953][run_fd_with_cosmos_framework.ipynb†L1089-L1097]. The README documents the schema at lines 356-357 [README.md†L356-L357].