# Configuring Action Policy Predictions for DROID Robot Embodiment in Cosmos 3

> Learn to configure DROID robot action policy predictions in Cosmos 3. Set up the environment, register the experiment, and run the training script for seamless embodiment.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Cosmos 3 configures DROID robot action policies using a 10-dimensional action space (9D end-effector pose plus gripper state), requiring the `DROID_ROOT` environment variable, registration of the `action_policy_droid_nano` experiment, and execution of the [`launch_sft_action_policy_droid.sh`](https://github.com/NVIDIA/cosmos/blob/main/launch_sft_action_policy_droid.sh) training script.**

The NVIDIA Cosmos repository provides a unified framework for generating and evaluating robot action policies across diverse embodiments. Configuring action policy predictions for the DROID robot embodiment in Cosmos 3 involves preparing the LeRobot dataset, defining hyperparameters in TOML configuration files, and deploying through specialized training scripts. The following sections detail the complete pipeline from data staging to closed-loop inference.

## Understanding the DROID Embodiment and Action Space

The DROID embodiment represents a single-arm robot configuration within the Cosmos 3 ecosystem. It processes multiview 480p video streams to predict future observations and generate corresponding robotic actions.

### Action Space Dimensions

According to the Cosmos 3 source code, the DROID embodiment supports two action representations:

- **End-effector pose**: A 10-dimensional vector comprising 9D end-effector pose plus 1D gripper state
- **Joint position**: An 8-dimensional `joint_pos` vector combined with proprioceptive state feedback

The finetuning configuration typically uses the joint position representation with a chunk length of 32 frames for temporal prediction.

## Preparing the DROID Dataset

The DROID LeRobot dataset must be properly staged before training begins. The framework accesses data through the `cosmos_framework.data.generator.action.datasets.DROIDLeRobotDataset` class.

### Environment Variables and Paths

Set the `DROID_ROOT` or `DATASET_PATH` environment variable to point to the dataset location:

```bash
export DROID_ROOT=/path/to/Cosmos3-DROID/success

# Alternative variable used by launch scripts

export DATASET_PATH=$DROID_ROOT

```

The finetuning cookbook validates the presence of [`data/Cosmos3-DROID/success/meta/info.json`](https://github.com/NVIDIA/cosmos/blob/main/data/Cosmos3-DROID/success/meta/info.json) to ensure dataset integrity.

### Downloading the Dataset

If the dataset is not locally staged, use the Hugging Face CLI to download it automatically:

```bash
uvx hf@latest download --repo-type dataset nvidia/Cosmos3-DROID --local-dir data/Cosmos3-DROID

```

This command fetches the DROID dataset and stores it in the expected directory structure.

## Configuring the Finetuning Experiment

Cosmos 3 uses TOML configuration files to define experiment parameters and model architectures.

### TOML Configuration File

The DROID-specific configuration resides in [`cookbooks/cosmos3/generator/action/finetune/toml/sft_config/action_policy_droid_repro.toml`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/finetune/toml/sft_config/action_policy_droid_repro.toml). This file registers the `action_policy_droid_nano` experiment, which activates DROID-specific action space handling and data loading logic.

Key configuration parameters include:

- **Experiment registration**: `experiment = "action_policy_droid_nano"`
- **Action type**: Joint position actions (`joint_pos`) with proprioceptive state
- **Video processing**: 480p multiview video chunks
- **Temporal window**: Chunk length of 32 frames
- **Data loading**: Episode-shuffle streaming with optional filtering via [`keep_ranges_1_0_1.json`](https://github.com/NVIDIA/cosmos/blob/main/keep_ranges_1_0_1.json)

### Dataset Integration

The configuration references the `DROIDLeRobotDataset` class to load training episodes. The dataset root path is resolved from the `DROID_ROOT` environment variable during runtime initialization.

## Launching the Training Pipeline

The [`launch_sft_action_policy_droid.sh`](https://github.com/NVIDIA/cosmos/blob/main/launch_sft_action_policy_droid.sh) script orchestrates the complete finetuning workflow.

### Environment Setup

The launch script performs several critical setup steps:

1. Validates the `DATASET_PATH` or falls back to `$PWD/data/Cosmos3-DROID/success`
2. Exports required environment variables:
   - `DROID_ROOT`: Path to dataset root
   - `BASE_CHECKPOINT_PATH`: Pretrained model checkpoint
   - `WAN_VAE_PATH`: VAE model weights

3. Downloads the dataset if not present using `uvx hf@latest download`

### Execution

Run the training script from the repository root:

```bash
bash cookbooks/cosmos3/generator/action/finetune/launch_sft_action_policy_droid.sh

```

This executes the `cosmos_framework.scripts.train` entrypoint with the TOML configuration, producing a finetuned model saved to the specified `RUN_DIR`.

## Serving the Trained Policy

After training, deploy the model using the Cosmos 3 inference server to enable real-time action prediction.

### Starting the Inference Server

Serve the trained policy using the **Cosmos 3-Nano-Policy-DROID** server configuration:

```bash
COSMOS_MODEL=$RUN_DIR/model  # Path to finetuned checkpoint

python -m cosmos_framework.scripts.inference \
    --model $COSMOS_MODEL \
    --device cuda \
    --port 8080

```

The server loads the trained weights and listens for incoming prediction requests.

### Client Request Format

During inference, clients send POST requests containing the current robot state and visual observations. The request body must include:

- **`joint_pos`**: Current 8-dimensional joint positions
- **`concat_view`**: Latest multiview video frames (480p)
- **`proprio`**: Optional proprioceptive sensor data

Example client implementation:

```python
import requests
import os

payload = {
    "joint_pos": current_joint_positions.tolist(),  # 8-dim vector

    "proprio": current_proprio_state,              # Optional sensor data

    "concat_view": latest_video_frames,            # 480p multiview stack

}

response = requests.post("http://localhost:8080/predict", json=payload)
future_actions = response.json()["actions"]

```

The server returns a sequence of future actions and predicted observations, enabling closed-loop control for the DROID robot in simulation or physical deployment.

## Forward Dynamics Generation

For offline evaluation and debugging, use the provided Jupyter notebook `cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb`. This notebook demonstrates forward-dynamics generation for DROID using the Cosmos 3 framework without requiring a live robot connection.

## Summary

- **DROID embodiment**: Uses a 10-dimensional action space (9D end-effector pose + gripper) or 8D joint positions with proprioception, paired with 480p multiview video
- **Dataset preparation**: Requires `DROID_ROOT` environment variable and uses `DROIDLeRobotDataset` class located at `cosmos_framework.data.generator.action.datasets`
- **Configuration**: Defined in [`action_policy_droid_repro.toml`](https://github.com/NVIDIA/cosmos/blob/main/action_policy_droid_repro.toml) with experiment name `action_policy_droid_nano`, chunk length 32, and episode-shuffle streaming
- **Training**: Executed via [`launch_sft_action_policy_droid.sh`](https://github.com/NVIDIA/cosmos/blob/main/launch_sft_action_policy_droid.sh) which sets environment variables and calls `cosmos_framework.scripts.train`
- **Inference**: Deployed using `cosmos_framework.scripts.inference` with client requests containing `joint_pos`, `concat_view`, and optional `proprio` data

## Frequently Asked Questions

### What is the action space dimension for DROID in Cosmos 3?

The DROID embodiment supports a 10-dimensional action space comprising 9D end-effector pose plus 1D gripper state, or alternatively 8D joint positions (`joint_pos`) combined with proprioceptive state. The configuration file specifies which representation to use during training and inference.

### How do I set up the dataset environment for DROID training?

Set the `DROID_ROOT` environment variable to point to your `Cosmos3-DROID/success` directory, or use `DATASET_PATH` as an alternative. The training scripts automatically validate the presence of [`data/Cosmos3-DROID/success/meta/info.json`](https://github.com/NVIDIA/cosmos/blob/main/data/Cosmos3-DROID/success/meta/info.json) and can download the dataset via `uvx hf@latest download` if not present locally.

### What is the role of the `action_policy_droid_nano` experiment?

The `action_policy_droid_nano` experiment identifier registers DROID-specific data loading and model configuration within the Cosmos 3 framework. It is defined in the TOML config file [`action_policy_droid_repro.toml`](https://github.com/NVIDIA/cosmos/blob/main/action_policy_droid_repro.toml) and activates the appropriate action space handling, chunk length settings (32 frames), and episode shuffling for the DROID LeRobot dataset.

### How do I deploy the trained policy for real-time inference?

After training, deploy the model using `python -m cosmos_framework.scripts.inference` with the `--model` path set to your output directory. The server accepts JSON POST requests containing `joint_pos` (8D vector), `concat_view` (multiview video), and optional `proprio` data, returning predicted future actions for robot control.