How to Integrate Cosmos 3 with Custom Robotics Embodiments: A Complete Guide
You can integrate Cosmos 3 with custom robotics embodiments by defining a unique domain_name, setting raw_action_dim to match your robot's state vector dimensionality, and passing action trajectories via the extra_params JSON field—no modifications to the core model code are required.
Cosmos 3 is NVIDIA's open-source omnimodal world model built on a unified Mixture-of-Transformers (MoT) architecture that processes text, vision, audio, and action tokens in a single transformer stack. While the repository includes pre-configured support for DROID, UMI, and autonomous vehicle embodiments, the action modality interface is designed to accept arbitrary fixed-size vectors. This guide explains how to extend Cosmos 3 to custom robotic hardware by leveraging the generic action token representation and extra_params configuration system.
Understanding Cosmos 3 Action Modality and Embodiment Parameters
Cosmos 3 treats action tokens as a core modality that represents sequences of robot or vehicle poses alongside visual and audio inputs. According to the repository's architecture documentation, the model uses a generic action interface that decodes fixed-dimensional vectors without hardcoding specific robot morphologies.
The system defines pre-canned embodiments with fixed token dimensions in the root README:
| Embodiment | Representation | Dimensionality |
|---|---|---|
| Autonomous vehicle | Ego pose (9 D) | 9 |
| DROID / UMI | End-effector pose (9 D) + gripper state (1 D) | 10 |
When invoking the Generator API, you specify your embodiment configuration through the extra_params JSON object:
action_mode– Operating mode:policy,inverse_dynamics, orforward_dynamicsdomain_name– Unique string identifier (e.g.,av,bridge_orig_lerobot, or your custom name)raw_action_dim– Dimensionality of a single action token (must match your robot's state vector)action_chunk_size– Number of action tokens in the conditioning trajectoryaction_path– Path to JSON/NumPy file containing the action sequence (required for forward dynamics)
These parameters are parsed in vllm/omni for API requests and in cosmos_framework/scripts/inference.py for native PyTorch execution. The server does not enforce a fixed list of domain names, enabling arbitrary robot definitions.
Step-by-Step Integration for Custom Robotics
To add a custom robotics embodiment beyond the built-in DROID, UMI, and vehicle configurations, follow this workflow:
-
Define the action representation
Determine the floating-point vector that describes your robot's state. For example, a 6-DoF arm with gripper might use 12 dimensions (6 for pose, 6 for velocity, or 6 for pose plus gripper width). -
Create the action trajectory file
Format your action sequence as a NumPy array or JSON file containing vectors of lengthraw_action_dim. Store this incookbooks/cosmos3/generator/action/assets/or your preferred path. -
Select a unique domain identifier
Choose a descriptive string such asmy6d_armorbimanual_14d. This identifier is passed directly in the request without requiring source code changes. -
Configure inference parameters
Setraw_action_dimto match your vector size andaction_chunk_sizeto the temporal length of your trajectory. -
Encode non-numeric metadata (optional)
If your action space includes discrete modes (e.g., gripper open/closed), encode these as additional continuous channels so the total dimension matchesraw_action_dim.
The tokenizers, diffusion transformer, and rotary position embeddings remain unchanged because the model processes action tokens as a generic stream of continuous values.
Inference Code Examples
Native PyTorch with Cosmos Framework
For direct model execution using the Cosmos Framework, create a JSON specification file and invoke the inference script:
# Create action specification for a custom 12-DoF robot
cat > my_robot_action.json <<'EOF'
{
"action_mode": "forward_dynamics",
"domain_name": "my12d_robot",
"raw_action_dim": 12,
"action_chunk_size": 20,
"action_path": "assets/my_robot_action.npy"
}
EOF
# Launch inference via torchrun
torchrun -m cosmos_framework.scripts.inference \
--parallelism-preset=latency \
-i my_robot_action.json \
-o /tmp/cosmos3_my_robot \
--checkpoint-path Cosmos3-Nano \
--seed 42
This entry point, located in cosmos_framework/scripts/inference.py, builds the input specification and invokes the diffusion model using the provided action conditioning. The JSON schema matches the examples in cookbooks/cosmos3/generator/action/assets/.
vLLM-Omni API via cURL
For production deployments using the OpenAI-compatible vLLM-Omni server, send a POST request with the custom embodiment parameters in extra_params:
curl -sS -X POST http://localhost:8000/v1/videos/sync \
--form-string "prompt=Industrial robot arm welding components" \
--form-string "negative_prompt=blur,artifacts" \
--form-string "size=1280x720" \
--form-string "num_frames=120" \
--form-string "fps=24" \
--form-string "num_inference_steps=30" \
--form-string "guidance_scale=6.0" \
--form-string "seed=1234" \
--form-string 'extra_params={
"action_mode":"forward_dynamics",
"domain_name":"my12d_robot",
"raw_action_dim":12,
"action_chunk_size":20,
"action_path":"assets/my_robot_action.npy"
}' \
-o custom_robot_output.mp4
The vLLM-Omni handler parses the extra_params field and maps it to the model's action interface, allowing the same action file to work across both inference methods.
Python Client with OpenAI SDK
For programmatic access, use the OpenAI Python SDK to send requests to your local vLLM-Omni endpoint:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
# Define custom embodiment parameters
extra = {
"action_mode": "forward_dynamics",
"domain_name": "my12d_robot",
"raw_action_dim": 12,
"action_chunk_size": 20,
"action_path": "assets/my_robot_action.npy",
}
response = client.videos.sync.create(
model="nvidia/Cosmos3-Nano",
prompt="Industrial robot arm welding components",
size="1280x720",
num_frames=120,
fps=24,
num_inference_steps=30,
guidance_scale=6.0,
seed=1234,
extra_params=extra,
)
with open("custom_robot_output.mp4", "wb") as f:
f.write(response.video)
This approach leverages the same extra_params dictionary used by the Cosmos Framework, ensuring consistency across native and API-based inference.
Key Source Files and Architecture References
Understanding the following files helps when debugging custom embodiment integration:
cookbooks/cosmos3/generator/action/README.md– Documents the action modality specification, dimension tables, and asset format requirements.cosmos_framework/scripts/inference.py– Entry point for native PyTorch inference that constructs the model input from JSON specifications.cookbooks/cosmos3/generator/action/assets/– Contains example action trajectory files (*.npy,*.json) serving as templates for custom robot data.vllm/omni(external repository) – Implements the request parsing logic forextra_paramsand maps them to the model's action interface.
Summary
- Cosmos 3 uses a generic action modality that accepts arbitrary fixed-size vectors through the
extra_paramsconfiguration field. - No source code changes are required to add custom embodiments—simply define a unique
domain_nameand setraw_action_dimto match your robot's state vector. - Action trajectories are passed as JSON or NumPy arrays via the
action_pathparameter, supporting both the Cosmos Framework and vLLM-Omni APIs. - Pre-canned configurations exist for 9D autonomous vehicles and 10D DROID/UMI robots, but any dimensionality is supported by the underlying MoT architecture.
- All three inference methods (native PyTorch, cURL, Python SDK) use identical
extra_paramsschemas, enabling seamless portability across deployment targets.
Frequently Asked Questions
What is the maximum action dimension supported by Cosmos 3?
The repository documentation does not specify a hard limit on raw_action_dim. The Mixture-of-Transformers architecture treats action tokens as a continuous stream, so you can theoretically use any fixed dimension that fits within your GPU memory constraints. In practice, keep dimensions below 100 to maintain inference efficiency, as each action token is processed through the same transformer layers as vision and audio tokens.
Do I need to retrain the model to support a new robot embodiment?
No, you do not need to retrain Cosmos 3 to integrate custom robotics embodiments. The model is trained on diverse action-conditioned video data and generalizes to new domain_name identifiers through the generic action token interface. Simply provide correctly formatted action trajectories with the appropriate raw_action_dim and the model will condition its video generation on your robot's state sequence.
Can I use Cosmos 3 with robots that have discrete action spaces?
Yes, but you must encode discrete actions as continuous values. For example, if your robot has a binary gripper state (open/closed), represent this as a 0.0 or 1.0 float value within the action vector. The model expects continuous action tokens, so discrete modes should be one-hot encoded or mapped to continuous ranges that match the raw_action_dim specified in your request.
Where should I store custom action trajectory files?
Store custom action files in the cookbooks/cosmos3/generator/action/assets/ directory to match the repository's example structure, or reference any accessible path via the action_path parameter. The NumPy or JSON format must contain an array of vectors where each vector has length equal to raw_action_dim, ordered temporally for the action_chunk_size duration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →