Configuring Action Policy Predictions for DROID Robot Embodiment in Cosmos 3

Cosmos 3 configures DROID robot action policies using a 10-dimensional action space (9D end-effector pose plus gripper state), requiring the DROID_ROOT environment variable, registration of the action_policy_droid_nano experiment, and execution of the launch_sft_action_policy_droid.sh training script.

The NVIDIA Cosmos repository provides a unified framework for generating and evaluating robot action policies across diverse embodiments. Configuring action policy predictions for the DROID robot embodiment in Cosmos 3 involves preparing the LeRobot dataset, defining hyperparameters in TOML configuration files, and deploying through specialized training scripts. The following sections detail the complete pipeline from data staging to closed-loop inference.

Understanding the DROID Embodiment and Action Space

The DROID embodiment represents a single-arm robot configuration within the Cosmos 3 ecosystem. It processes multiview 480p video streams to predict future observations and generate corresponding robotic actions.

Action Space Dimensions

According to the Cosmos 3 source code, the DROID embodiment supports two action representations:

  • End-effector pose: A 10-dimensional vector comprising 9D end-effector pose plus 1D gripper state
  • Joint position: An 8-dimensional joint_pos vector combined with proprioceptive state feedback

The finetuning configuration typically uses the joint position representation with a chunk length of 32 frames for temporal prediction.

Preparing the DROID Dataset

The DROID LeRobot dataset must be properly staged before training begins. The framework accesses data through the cosmos_framework.data.generator.action.datasets.DROIDLeRobotDataset class.

Environment Variables and Paths

Set the DROID_ROOT or DATASET_PATH environment variable to point to the dataset location:

export DROID_ROOT=/path/to/Cosmos3-DROID/success

# Alternative variable used by launch scripts

export DATASET_PATH=$DROID_ROOT

The finetuning cookbook validates the presence of data/Cosmos3-DROID/success/meta/info.json to ensure dataset integrity.

Downloading the Dataset

If the dataset is not locally staged, use the Hugging Face CLI to download it automatically:

uvx hf@latest download --repo-type dataset nvidia/Cosmos3-DROID --local-dir data/Cosmos3-DROID

This command fetches the DROID dataset and stores it in the expected directory structure.

Configuring the Finetuning Experiment

Cosmos 3 uses TOML configuration files to define experiment parameters and model architectures.

TOML Configuration File

The DROID-specific configuration resides in cookbooks/cosmos3/generator/action/finetune/toml/sft_config/action_policy_droid_repro.toml. This file registers the action_policy_droid_nano experiment, which activates DROID-specific action space handling and data loading logic.

Key configuration parameters include:

  • Experiment registration: experiment = "action_policy_droid_nano"
  • Action type: Joint position actions (joint_pos) with proprioceptive state
  • Video processing: 480p multiview video chunks
  • Temporal window: Chunk length of 32 frames
  • Data loading: Episode-shuffle streaming with optional filtering via keep_ranges_1_0_1.json

Dataset Integration

The configuration references the DROIDLeRobotDataset class to load training episodes. The dataset root path is resolved from the DROID_ROOT environment variable during runtime initialization.

Launching the Training Pipeline

The launch_sft_action_policy_droid.sh script orchestrates the complete finetuning workflow.

Environment Setup

The launch script performs several critical setup steps:

  1. Validates the DATASET_PATH or falls back to $PWD/data/Cosmos3-DROID/success

  2. Exports required environment variables:

    • DROID_ROOT: Path to dataset root
    • BASE_CHECKPOINT_PATH: Pretrained model checkpoint
    • WAN_VAE_PATH: VAE model weights
  3. Downloads the dataset if not present using uvx hf@latest download

Execution

Run the training script from the repository root:

bash cookbooks/cosmos3/generator/action/finetune/launch_sft_action_policy_droid.sh

This executes the cosmos_framework.scripts.train entrypoint with the TOML configuration, producing a finetuned model saved to the specified RUN_DIR.

Serving the Trained Policy

After training, deploy the model using the Cosmos 3 inference server to enable real-time action prediction.

Starting the Inference Server

Serve the trained policy using the Cosmos 3-Nano-Policy-DROID server configuration:

COSMOS_MODEL=$RUN_DIR/model  # Path to finetuned checkpoint

python -m cosmos_framework.scripts.inference \
    --model $COSMOS_MODEL \
    --device cuda \
    --port 8080

The server loads the trained weights and listens for incoming prediction requests.

Client Request Format

During inference, clients send POST requests containing the current robot state and visual observations. The request body must include:

  • joint_pos: Current 8-dimensional joint positions
  • concat_view: Latest multiview video frames (480p)
  • proprio: Optional proprioceptive sensor data

Example client implementation:

import requests
import os

payload = {
    "joint_pos": current_joint_positions.tolist(),  # 8-dim vector

    "proprio": current_proprio_state,              # Optional sensor data

    "concat_view": latest_video_frames,            # 480p multiview stack

}

response = requests.post("http://localhost:8080/predict", json=payload)
future_actions = response.json()["actions"]

The server returns a sequence of future actions and predicted observations, enabling closed-loop control for the DROID robot in simulation or physical deployment.

Forward Dynamics Generation

For offline evaluation and debugging, use the provided Jupyter notebook cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb. This notebook demonstrates forward-dynamics generation for DROID using the Cosmos 3 framework without requiring a live robot connection.

Summary

  • DROID embodiment: Uses a 10-dimensional action space (9D end-effector pose + gripper) or 8D joint positions with proprioception, paired with 480p multiview video
  • Dataset preparation: Requires DROID_ROOT environment variable and uses DROIDLeRobotDataset class located at cosmos_framework.data.generator.action.datasets
  • Configuration: Defined in action_policy_droid_repro.toml with experiment name action_policy_droid_nano, chunk length 32, and episode-shuffle streaming
  • Training: Executed via launch_sft_action_policy_droid.sh which sets environment variables and calls cosmos_framework.scripts.train
  • Inference: Deployed using cosmos_framework.scripts.inference with client requests containing joint_pos, concat_view, and optional proprio data

Frequently Asked Questions

What is the action space dimension for DROID in Cosmos 3?

The DROID embodiment supports a 10-dimensional action space comprising 9D end-effector pose plus 1D gripper state, or alternatively 8D joint positions (joint_pos) combined with proprioceptive state. The configuration file specifies which representation to use during training and inference.

How do I set up the dataset environment for DROID training?

Set the DROID_ROOT environment variable to point to your Cosmos3-DROID/success directory, or use DATASET_PATH as an alternative. The training scripts automatically validate the presence of data/Cosmos3-DROID/success/meta/info.json and can download the dataset via uvx hf@latest download if not present locally.

What is the role of the action_policy_droid_nano experiment?

The action_policy_droid_nano experiment identifier registers DROID-specific data loading and model configuration within the Cosmos 3 framework. It is defined in the TOML config file action_policy_droid_repro.toml and activates the appropriate action space handling, chunk length settings (32 frames), and episode shuffling for the DROID LeRobot dataset.

How do I deploy the trained policy for real-time inference?

After training, deploy the model using python -m cosmos_framework.scripts.inference with the --model path set to your output directory. The server accepts JSON POST requests containing joint_pos (8D vector), concat_view (multiview video), and optional proprio data, returning predicted future actions for robot control.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →