Configuring Action Policy Predictions for DROID Robot Embodiment in Cosmos 3
Cosmos 3 configures DROID robot action policies using a 10-dimensional action space (9D end-effector pose plus gripper state), requiring the DROID_ROOT environment variable, registration of the action_policy_droid_nano experiment, and execution of the launch_sft_action_policy_droid.sh training script.
The NVIDIA Cosmos repository provides a unified framework for generating and evaluating robot action policies across diverse embodiments. Configuring action policy predictions for the DROID robot embodiment in Cosmos 3 involves preparing the LeRobot dataset, defining hyperparameters in TOML configuration files, and deploying through specialized training scripts. The following sections detail the complete pipeline from data staging to closed-loop inference.
Understanding the DROID Embodiment and Action Space
The DROID embodiment represents a single-arm robot configuration within the Cosmos 3 ecosystem. It processes multiview 480p video streams to predict future observations and generate corresponding robotic actions.
Action Space Dimensions
According to the Cosmos 3 source code, the DROID embodiment supports two action representations:
- End-effector pose: A 10-dimensional vector comprising 9D end-effector pose plus 1D gripper state
- Joint position: An 8-dimensional
joint_posvector combined with proprioceptive state feedback
The finetuning configuration typically uses the joint position representation with a chunk length of 32 frames for temporal prediction.
Preparing the DROID Dataset
The DROID LeRobot dataset must be properly staged before training begins. The framework accesses data through the cosmos_framework.data.generator.action.datasets.DROIDLeRobotDataset class.
Environment Variables and Paths
Set the DROID_ROOT or DATASET_PATH environment variable to point to the dataset location:
export DROID_ROOT=/path/to/Cosmos3-DROID/success
# Alternative variable used by launch scripts
export DATASET_PATH=$DROID_ROOT
The finetuning cookbook validates the presence of data/Cosmos3-DROID/success/meta/info.json to ensure dataset integrity.
Downloading the Dataset
If the dataset is not locally staged, use the Hugging Face CLI to download it automatically:
uvx hf@latest download --repo-type dataset nvidia/Cosmos3-DROID --local-dir data/Cosmos3-DROID
This command fetches the DROID dataset and stores it in the expected directory structure.
Configuring the Finetuning Experiment
Cosmos 3 uses TOML configuration files to define experiment parameters and model architectures.
TOML Configuration File
The DROID-specific configuration resides in cookbooks/cosmos3/generator/action/finetune/toml/sft_config/action_policy_droid_repro.toml. This file registers the action_policy_droid_nano experiment, which activates DROID-specific action space handling and data loading logic.
Key configuration parameters include:
- Experiment registration:
experiment = "action_policy_droid_nano" - Action type: Joint position actions (
joint_pos) with proprioceptive state - Video processing: 480p multiview video chunks
- Temporal window: Chunk length of 32 frames
- Data loading: Episode-shuffle streaming with optional filtering via
keep_ranges_1_0_1.json
Dataset Integration
The configuration references the DROIDLeRobotDataset class to load training episodes. The dataset root path is resolved from the DROID_ROOT environment variable during runtime initialization.
Launching the Training Pipeline
The launch_sft_action_policy_droid.sh script orchestrates the complete finetuning workflow.
Environment Setup
The launch script performs several critical setup steps:
-
Validates the
DATASET_PATHor falls back to$PWD/data/Cosmos3-DROID/success -
Exports required environment variables:
DROID_ROOT: Path to dataset rootBASE_CHECKPOINT_PATH: Pretrained model checkpointWAN_VAE_PATH: VAE model weights
-
Downloads the dataset if not present using
uvx hf@latest download
Execution
Run the training script from the repository root:
bash cookbooks/cosmos3/generator/action/finetune/launch_sft_action_policy_droid.sh
This executes the cosmos_framework.scripts.train entrypoint with the TOML configuration, producing a finetuned model saved to the specified RUN_DIR.
Serving the Trained Policy
After training, deploy the model using the Cosmos 3 inference server to enable real-time action prediction.
Starting the Inference Server
Serve the trained policy using the Cosmos 3-Nano-Policy-DROID server configuration:
COSMOS_MODEL=$RUN_DIR/model # Path to finetuned checkpoint
python -m cosmos_framework.scripts.inference \
--model $COSMOS_MODEL \
--device cuda \
--port 8080
The server loads the trained weights and listens for incoming prediction requests.
Client Request Format
During inference, clients send POST requests containing the current robot state and visual observations. The request body must include:
joint_pos: Current 8-dimensional joint positionsconcat_view: Latest multiview video frames (480p)proprio: Optional proprioceptive sensor data
Example client implementation:
import requests
import os
payload = {
"joint_pos": current_joint_positions.tolist(), # 8-dim vector
"proprio": current_proprio_state, # Optional sensor data
"concat_view": latest_video_frames, # 480p multiview stack
}
response = requests.post("http://localhost:8080/predict", json=payload)
future_actions = response.json()["actions"]
The server returns a sequence of future actions and predicted observations, enabling closed-loop control for the DROID robot in simulation or physical deployment.
Forward Dynamics Generation
For offline evaluation and debugging, use the provided Jupyter notebook cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb. This notebook demonstrates forward-dynamics generation for DROID using the Cosmos 3 framework without requiring a live robot connection.
Summary
- DROID embodiment: Uses a 10-dimensional action space (9D end-effector pose + gripper) or 8D joint positions with proprioception, paired with 480p multiview video
- Dataset preparation: Requires
DROID_ROOTenvironment variable and usesDROIDLeRobotDatasetclass located atcosmos_framework.data.generator.action.datasets - Configuration: Defined in
action_policy_droid_repro.tomlwith experiment nameaction_policy_droid_nano, chunk length 32, and episode-shuffle streaming - Training: Executed via
launch_sft_action_policy_droid.shwhich sets environment variables and callscosmos_framework.scripts.train - Inference: Deployed using
cosmos_framework.scripts.inferencewith client requests containingjoint_pos,concat_view, and optionalpropriodata
Frequently Asked Questions
What is the action space dimension for DROID in Cosmos 3?
The DROID embodiment supports a 10-dimensional action space comprising 9D end-effector pose plus 1D gripper state, or alternatively 8D joint positions (joint_pos) combined with proprioceptive state. The configuration file specifies which representation to use during training and inference.
How do I set up the dataset environment for DROID training?
Set the DROID_ROOT environment variable to point to your Cosmos3-DROID/success directory, or use DATASET_PATH as an alternative. The training scripts automatically validate the presence of data/Cosmos3-DROID/success/meta/info.json and can download the dataset via uvx hf@latest download if not present locally.
What is the role of the action_policy_droid_nano experiment?
The action_policy_droid_nano experiment identifier registers DROID-specific data loading and model configuration within the Cosmos 3 framework. It is defined in the TOML config file action_policy_droid_repro.toml and activates the appropriate action space handling, chunk length settings (32 frames), and episode shuffling for the DROID LeRobot dataset.
How do I deploy the trained policy for real-time inference?
After training, deploy the model using python -m cosmos_framework.scripts.inference with the --model path set to your output directory. The server accepts JSON POST requests containing joint_pos (8D vector), concat_view (multiview video), and optional proprio data, returning predicted future actions for robot control.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →