How to Integrate Cosmos 3 with Custom Robot Embodiments: DROID, UMI, and Bridge Guide
Cosmos 3 treats robot control as an action modality that requires specifying the embodiment's domain_name and raw_action_dim in the extra_params payload to correctly tokenize 10-dimensional action sequences.
NVIDIA Cosmos represents robot control as a tokenized action modality rather than raw vectors, enabling seamless integration with diverse robot embodiments. When working with datasets like DROID, UMI, or Bridge, you must configure the model's action space parameters to match the specific robot's degrees of freedom. This guide explains how to integrate these custom embodiments using the vLLM-Omni endpoint or the Cosmos Framework inference scripts.
Understanding Action Modalities and Tokenization
Cosmos 3 processes robot movements as sequences of action tokens representing pose and gripper state. According to the action cookbook at cookbooks/cosmos3/generator/action/README.md (lines 37-42), each action sequence encodes a 9-D pose concatenated with a 1-D grasp value, resulting in a 10-dimensional action space for single-arm robots like DROID, UMI, and Bridge.
The model requires two critical fields to parse these sequences correctly:
domain_name: Selects the specific action space tokenizer (e.g.,droid_orig_lerobot,umi_orig_lerobot,bridge_orig_lerobot)raw_action_dim: Specifies the dimensionality of the action vector (10 for these single-arm embodiments)
Supported Embodiments and Configuration
The supported embodiments are enumerated in the main README.md at line 107. For robotics research, the following configurations apply:
| Embodiment | domain_name | raw_action_dim | Description |
|---|---|---|---|
| DROID | droid_orig_lerobot |
10 | 9D pose + 1D gripper |
| UMI | umi_orig_lerobot |
10 | 9D pose + 1D gripper |
| Bridge | bridge_orig_lerobot |
10 | 9D pose + 1D gripper |
Integration via vLLM-Omni Endpoint
When using the vLLM-Omni API, pass embodiment parameters through the extra_params JSON payload as documented in README.md (line 356). This approach supports forward-dynamics, inverse-dynamics, and policy generation tasks.
Example for forward-dynamics with a Bridge robot:
import json
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
# Path to a JSON file containing the 10-D action chunk for the Bridge robot
action_path = "/workspace/assets/bridge_action.json"
response = client.videos.create(
model="nvidia/Cosmos3-Nano",
prompt="A robot picks up a red cube and places it on a shelf.",
negative_prompt="blur, distortion",
size="1280x720",
num_frames=189,
fps=24,
num_inference_steps=30,
guidance_scale=1.0,
flow_shift=10.0,
seed=42,
input_reference=open("assets/start_image.jpg", "rb"),
extra_params={
"domain_name": "bridge_orig_lerobot",
"raw_action_dim": 10,
"action_path": action_path
}
)
Integration via Cosmos Framework Scripts
Alternatively, use the cosmos_framework.scripts.inference module for batched or programmatic access. This method accepts the same extra_params dictionary containing domain_name and raw_action_dim, allowing you to specify the action modality directly in your Python script without running the vLLM server.
Step-by-Step Integration Workflow
-
Select your inference surface: Choose between the Generator surface (for forward-dynamics, inverse-dynamics, or policy tasks) or the Predictor surface for evaluation.
-
Identify your embodiment: Map your robot to the correct
domain_namefrom the supported embodiments table. -
Prepare action data: Format your action sequences as JSON arrays with shape
(sequence_length, 10)where columns represent the 9D pose and 1D grasp state. -
Configure the request: Include
domain_nameandraw_action_diminextra_params, and provide theaction_pathor embedded action array along with the start image (input_reference).
Summary
- Cosmos 3 uses action modalities to represent robot control as tokenized sequences.
- Specify
domain_name(e.g.,droid_orig_lerobot) andraw_action_dim(10 for single-arm robots) inextra_params. - Configure requests via the vLLM-Omni endpoint or Cosmos Framework scripts (
cosmos_framework.scripts.inference). - Action files must contain 10-dimensional vectors encoding 9D pose plus 1D gripper state.
Frequently Asked Questions
What is the action dimensionality for DROID, UMI, and Bridge robots?
All three embodiments use a 10-dimensional action space as defined in cookbooks/cosmos3/generator/action/README.md (lines 37-42). This consists of a 9-dimensional pose representation combined with a 1-dimensional grasp value, regardless of whether you are using DROID, UMI, or Bridge datasets.
How do I specify the robot embodiment when calling the Cosmos 3 API?
You must include the domain_name and raw_action_dim parameters inside the extra_params JSON payload. For example, use "domain_name": "umi_orig_lerobot" and "raw_action_dim": 10 when working with UMI data, as documented in the main README.md at line 356.
Can I use Cosmos 3 for inverse dynamics and policy generation?
Yes. The Generator surface supports three robot-centric tasks: forward-dynamics (predicting future states from actions), inverse-dynamics (predicting actions from state changes), and policy generation (behavioral cloning). Configure the task type and provide the appropriate action or state references in your request payload.
What file format should I use for the action chunk?
Action chunks should be stored as JSON files containing arrays of 10-dimensional float values. When using forward-dynamics, specify the path to this file via the action_path parameter in extra_params, or embed the array directly depending on your chosen inference method.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →