# How to Configure Action Conditioning for Robot Embodiments in NVIDIA Cosmos

> Configure action conditioning for robot embodiments like DROID or UMI in NVIDIA Cosmos. Learn JSON array dimensionality for 10-D robots, autonomous vehicles, and egocentric motion.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-06

---

**Action conditioning in NVIDIA Cosmos requires a JSON array with dimensionality specific to your embodiment: 10-D for single-arm robots like DROID and UMI, 9-D for autonomous vehicles, and 57-D for egocentric motion, processed in 16-frame chunks for forward-dynamics generation.**

NVIDIA Cosmos enables action-conditioned video generation and forward-dynamics simulation across diverse robotic platforms. To generate accurate motion sequences, you must configure the action array to match the specific dimensionality requirements of your robot type, whether it's a single-arm manipulator, humanoid, or autonomous vehicle. The repository provides sample configurations and validation logic in the forward-dynamics cookbooks to ensure your action data aligns with the model's expectations.

## Action Representation by Embodiment

Cosmos 3 accepts action-conditioned requests where the *action* field is a JSON array. The dimensionality must match the specific embodiment you are modeling, as defined in the repository's README.

| Embodiment | Action Dimensions | Typical Units | Normalization |
|------------|------------------|---------------|---------------|
| Camera motion | 9-D (rotation + translation) | metres / radians | -1 → 1 |
| Autonomous vehicle | 9-D (ego pose per frame) | metres / radians | -1 → 1 |
| Egocentric motion | 57-D (full body pose) | metres / radians | -1 → 1 |
| Single-arm robot (DROID, UMI) | 10-D (end-effector pose + gripper) | metres / radians | -1 → 1 |
| Dual-arm robot | 20-D (two 10-D arms) | metres / radians | -1 → 1 |
| Humanoid robot (AgiBot) | 29-D | metres / radians | -1 → 1 |

## Request Format and Chunking Mechanism

The forward-dynamics pipeline processes action data in specific chunks and requires a structured JSON payload.

### Chunking Requirements

For forward-dynamics generation, Cosmos processes chunks of **16 consecutive frames**. Your action file must contain `16 × N` rows, where `N` represents the number of chunks you want to generate. Each row must have the exact dimensionality of the selected embodiment.

### Conditioning Image Logic

Chunk 0 uses a static conditioning image shipped with the repository (such as those in `cookbooks/cosmos3/generator/action/assets/images/`). Subsequent chunks automatically use the **last generated frame** of the previous chunk as their conditioning image. This logic is implemented in the forward-dynamics notebooks `run_fd_with_cosmos_framework.ipynb` and `run_fd_with_vllm.ipynb`.

### JSON Request Structure

Every generator call requires the following JSON structure:

```json
{
  "prompt": "<optional-text-prompt>",
  "image": "<path-or-URL-to-conditioning-image>",
  "action": [ <list-of-action-vectors> ],
  "size": "<output-resolution-tier>",
  "frame_rate": <int>,
  "num_frames": <int>
}

```

The **`action`** field is the only component that varies between embodiments.

## Embodiment-Specific Configuration

### DROID Configuration

For the DROID single-arm robot, use 10-D action vectors representing end-effector pose and gripper state.

**Sample file:** [`cookbooks/cosmos3/generator/action/assets/actions/droid.json`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/assets/actions/droid.json)

**Array format:** `[[x, y, z, rx, ry, rz, gripper_open], ...]` where each inner list contains position, orientation, and gripper state in meters and radians.

The forward-dynamics notebooks automatically split this into 16-row chunks and validate dimensionality with assertions like `assert all(len(row) == 10 for row in action)`.

### UMI Configuration

The UMI (Universally Manipulation Interface) embodiment uses the same 10-D format as DROID but operates in the UMI coordinate system (right-handed, meters).

**Sample file:** [`cookbooks/cosmos3/generator/action/assets/actions/umi.json`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/assets/actions/umi.json)

**Validation:** The notebook `run_fd_with_vllm.ipynb` validates UMI actions using:

```python
assert all(len(row) == umi_raw_action_dim for row in umi_action)

```

### Autonomous Vehicle Configuration

Autonomous vehicles require 9-D ego pose trajectories.

**Sample file:** [`cookbooks/cosmos3/generator/action/assets/actions/av_traj_forward.json`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/assets/actions/av_traj_forward.json)

**Array format:** `[x, y, z, yaw, pitch, roll, speed, steering, throttle]` per frame.

## Implementation Examples

### Python: DROID Forward Dynamics

This example from `run_fd_with_cosmos_framework.ipynb` demonstrates loading and validating DROID actions:

```python
from cosmos_framework.scripts.inference import run_forward_dynamics
import json

# Load the DROID action JSON

with open("cookbooks/cosmos3/generator/action/assets/actions/droid.json") as f:
    droid_action = json.load(f)

# Validate 10-D structure

assert all(len(row) == 10 for row in droid_action), "DROID action must be 10-D"

# Construct request

request = {
    "prompt": "You are a robot manipulator solving a block-stacking task.",
    "image": "cookbooks/cosmos3/generator/action/assets/images/droid.png",
    "action": droid_action,
    "size": "480p",
    "frame_rate": 15,
    "num_frames": len(droid_action),
}

# Run inference (automatically handles 16-frame chunking)

run_forward_dynamics(request, output_path="outputs/droid_forward.mp4")

```

### Python: UMI with Environment Variables

From `run_fd_with_vllm.ipynb`, this snippet shows UMI-specific setup:

```python
import json
import os

umi_path = "cookbooks/cosmos3/generator/action/assets/actions/umi.json"
with open(umi_path) as f:
    umi_action = json.load(f)

# Ensure 10-D rows

assert all(len(row) == 10 for row in umi_action), "UMI action must be 10-D"

# Set environment vars for the notebook pipeline

os.environ["COSMOS3_UMI_FD_INPUT"] = str(umi_path)
os.environ["COSMOS3_UMI_FD_OUTPUT"] = "outputs/umi_forward"

```

### Bash: CLI Invocation for Autonomous Vehicles

Use the Cosmos Framework CLI entry point to run action-conditioned generation:

```bash
cosmos_framework.scripts.inference \
    --model cosmos3-nano \
    --task forward_dynamics \
    --embodiment av \
    --action-file cookbooks/cosmos3/generator/action/assets/actions/av_traj_forward.json \
    --image assets/av/first_frame.png \
    --output av_forward.mp4

```

## Summary

- **Match dimensions exactly:** Use 10-D arrays for DROID and UMI, 9-D for autonomous vehicles, and 57-D for egocentric motion.
- **Chunk into 16-frame sequences:** Ensure your action file contains multiples of 16 rows; the model processes forward-dynamics in 16-frame chunks.
- **Validate before inference:** Use assertions like `assert all(len(row) == 10 for row in action)` to catch dimensionality mismatches early.
- **Use provided samples:** Reference `cookbooks/cosmos3/generator/action/assets/actions/` for correctly formatted JSON templates.
- **Leverage automatic chunking:** The `cosmos_framework.scripts.inference` CLI and notebooks handle the 16-frame folding automatically when you provide the full action array.

## Frequently Asked Questions

### What is the difference between DROID and UMI action conditioning?

Both DROID and UMI use **10-D action representations** (end-effector position, orientation, and gripper state), but they differ in coordinate system conventions. DROID uses its native coordinate frame, while UMI uses a right-handed coordinate system specific to the Universally Manipulation Interface. Both require normalization to the [-1, 1] range, and both validate dimensionality using `assert all(len(row) == 10 for row in action)` in their respective notebooks.

### How does the 16-frame chunking mechanism work in forward dynamics?

Cosmos processes forward-dynamics generation in fixed chunks of **16 consecutive frames**. Your action file must contain `16 × N` rows, where `N` is the number of chunks. The first chunk uses your provided conditioning image, while subsequent chunks automatically use the last generated frame from the previous chunk as their conditioning image. The `run_fd_with_cosmos_framework.ipynb` notebook implements this logic transparently.

### Can I use custom action data instead of the provided samples?

Yes, you can use custom action data as long as you maintain the correct dimensionality for your embodiment and ensure the row count is a multiple of 16. Convert your poses into flat lists of normalized floats (-1 to 1), validate the length matches the expected dimension (e.g., 10 for single-arm robots), and pass the JSON array to the inference script via `--action-file` or directly in the Python API.

### What normalization range should action values use?

According to the NVIDIA Cosmos source code, all action values must be normalized to the **-1 to 1 range** regardless of embodiment type. This applies to positions, orientations, gripper states, and vehicle controls. The sample files in `cookbooks/cosmos3/generator/action/assets/actions/` demonstrate this normalization for each supported robot type.