# How to Integrate Cosmos 3 with Custom Robotics Embodiments: A Complete Guide

> Integrate Cosmos 3 with custom robotics embodiments by defining domain name, raw action dim, and passing action trajectories via extra_params. No core model modifications needed. Get the complete guide from NVIDIA/cosmos.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-14

---

**You can integrate Cosmos 3 with custom robotics embodiments by defining a unique `domain_name`, setting `raw_action_dim` to match your robot's state vector dimensionality, and passing action trajectories via the `extra_params` JSON field—no modifications to the core model code are required.**

Cosmos 3 is NVIDIA's open-source **omnimodal world model** built on a unified Mixture-of-Transformers (MoT) architecture that processes text, vision, audio, and action tokens in a single transformer stack. While the repository includes pre-configured support for DROID, UMI, and autonomous vehicle embodiments, the action modality interface is designed to accept arbitrary fixed-size vectors. This guide explains how to extend Cosmos 3 to custom robotic hardware by leveraging the generic action token representation and `extra_params` configuration system.

## Understanding Cosmos 3 Action Modality and Embodiment Parameters

Cosmos 3 treats action tokens as a core modality that represents sequences of robot or vehicle poses alongside visual and audio inputs. According to the repository's architecture documentation, the model uses a generic action interface that decodes fixed-dimensional vectors without hardcoding specific robot morphologies.

The system defines pre-canned embodiments with fixed token dimensions in the root README:

| Embodiment | Representation | Dimensionality |
|-----------|----------------|----------------|
| Autonomous vehicle | Ego pose (9 D) | 9 |
| DROID / UMI | End-effector pose (9 D) + gripper state (1 D) | 10 |

When invoking the Generator API, you specify your embodiment configuration through the `extra_params` JSON object:

- **`action_mode`** – Operating mode: `policy`, `inverse_dynamics`, or `forward_dynamics`
- **`domain_name`** – Unique string identifier (e.g., `av`, `bridge_orig_lerobot`, or your custom name)
- **`raw_action_dim`** – Dimensionality of a single action token (must match your robot's state vector)
- **`action_chunk_size`** – Number of action tokens in the conditioning trajectory
- **`action_path`** – Path to JSON/NumPy file containing the action sequence (required for forward dynamics)

These parameters are parsed in `vllm/omni` for API requests and in [`cosmos_framework/scripts/inference.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/inference.py) for native PyTorch execution. The server does not enforce a fixed list of domain names, enabling arbitrary robot definitions.

## Step-by-Step Integration for Custom Robotics

To add a custom robotics embodiment beyond the built-in DROID, UMI, and vehicle configurations, follow this workflow:

1. **Define the action representation**  
   Determine the floating-point vector that describes your robot's state. For example, a 6-DoF arm with gripper might use 12 dimensions (6 for pose, 6 for velocity, or 6 for pose plus gripper width).

2. **Create the action trajectory file**  
   Format your action sequence as a NumPy array or JSON file containing vectors of length `raw_action_dim`. Store this in `cookbooks/cosmos3/generator/action/assets/` or your preferred path.

3. **Select a unique domain identifier**  
   Choose a descriptive string such as `my6d_arm` or `bimanual_14d`. This identifier is passed directly in the request without requiring source code changes.

4. **Configure inference parameters**  
   Set `raw_action_dim` to match your vector size and `action_chunk_size` to the temporal length of your trajectory.

5. **Encode non-numeric metadata (optional)**  
   If your action space includes discrete modes (e.g., gripper open/closed), encode these as additional continuous channels so the total dimension matches `raw_action_dim`.

The tokenizers, diffusion transformer, and rotary position embeddings remain unchanged because the model processes action tokens as a generic stream of continuous values.

## Inference Code Examples

### Native PyTorch with Cosmos Framework

For direct model execution using the Cosmos Framework, create a JSON specification file and invoke the inference script:

```bash

# Create action specification for a custom 12-DoF robot

cat > my_robot_action.json <<'EOF'
{
  "action_mode": "forward_dynamics",
  "domain_name": "my12d_robot",
  "raw_action_dim": 12,
  "action_chunk_size": 20,
  "action_path": "assets/my_robot_action.npy"
}
EOF

```

```bash

# Launch inference via torchrun

torchrun -m cosmos_framework.scripts.inference \
  --parallelism-preset=latency \
  -i my_robot_action.json \
  -o /tmp/cosmos3_my_robot \
  --checkpoint-path Cosmos3-Nano \
  --seed 42

```

This entry point, located in [`cosmos_framework/scripts/inference.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/inference.py), builds the input specification and invokes the diffusion model using the provided action conditioning. The JSON schema matches the examples in `cookbooks/cosmos3/generator/action/assets/`.

### vLLM-Omni API via cURL

For production deployments using the OpenAI-compatible vLLM-Omni server, send a POST request with the custom embodiment parameters in `extra_params`:

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=Industrial robot arm welding components" \
  --form-string "negative_prompt=blur,artifacts" \
  --form-string "size=1280x720" \
  --form-string "num_frames=120" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=30" \
  --form-string "guidance_scale=6.0" \
  --form-string "seed=1234" \
  --form-string 'extra_params={
    "action_mode":"forward_dynamics",
    "domain_name":"my12d_robot",
    "raw_action_dim":12,
    "action_chunk_size":20,
    "action_path":"assets/my_robot_action.npy"
  }' \
  -o custom_robot_output.mp4

```

The vLLM-Omni handler parses the `extra_params` field and maps it to the model's action interface, allowing the same action file to work across both inference methods.

### Python Client with OpenAI SDK

For programmatic access, use the OpenAI Python SDK to send requests to your local vLLM-Omni endpoint:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")

# Define custom embodiment parameters

extra = {
    "action_mode": "forward_dynamics",
    "domain_name": "my12d_robot",
    "raw_action_dim": 12,
    "action_chunk_size": 20,
    "action_path": "assets/my_robot_action.npy",
}

response = client.videos.sync.create(
    model="nvidia/Cosmos3-Nano",
    prompt="Industrial robot arm welding components",
    size="1280x720",
    num_frames=120,
    fps=24,
    num_inference_steps=30,
    guidance_scale=6.0,
    seed=1234,
    extra_params=extra,
)

with open("custom_robot_output.mp4", "wb") as f:
    f.write(response.video)

```

This approach leverages the same `extra_params` dictionary used by the Cosmos Framework, ensuring consistency across native and API-based inference.

## Key Source Files and Architecture References

Understanding the following files helps when debugging custom embodiment integration:

- **[`cookbooks/cosmos3/generator/action/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/README.md)** – Documents the action modality specification, dimension tables, and asset format requirements.
- **[`cosmos_framework/scripts/inference.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/inference.py)** – Entry point for native PyTorch inference that constructs the model input from JSON specifications.
- **`cookbooks/cosmos3/generator/action/assets/`** – Contains example action trajectory files (`*.npy`, `*.json`) serving as templates for custom robot data.
- **`vllm/omni` (external repository)** – Implements the request parsing logic for `extra_params` and maps them to the model's action interface.

## Summary

- **Cosmos 3 uses a generic action modality** that accepts arbitrary fixed-size vectors through the `extra_params` configuration field.
- **No source code changes are required** to add custom embodiments—simply define a unique `domain_name` and set `raw_action_dim` to match your robot's state vector.
- **Action trajectories** are passed as JSON or NumPy arrays via the `action_path` parameter, supporting both the Cosmos Framework and vLLM-Omni APIs.
- **Pre-canned configurations** exist for 9D autonomous vehicles and 10D DROID/UMI robots, but any dimensionality is supported by the underlying MoT architecture.
- **All three inference methods** (native PyTorch, cURL, Python SDK) use identical `extra_params` schemas, enabling seamless portability across deployment targets.

## Frequently Asked Questions

### What is the maximum action dimension supported by Cosmos 3?

The repository documentation does not specify a hard limit on `raw_action_dim`. The Mixture-of-Transformers architecture treats action tokens as a continuous stream, so you can theoretically use any fixed dimension that fits within your GPU memory constraints. In practice, keep dimensions below 100 to maintain inference efficiency, as each action token is processed through the same transformer layers as vision and audio tokens.

### Do I need to retrain the model to support a new robot embodiment?

No, you do not need to retrain Cosmos 3 to integrate custom robotics embodiments. The model is trained on diverse action-conditioned video data and generalizes to new `domain_name` identifiers through the generic action token interface. Simply provide correctly formatted action trajectories with the appropriate `raw_action_dim` and the model will condition its video generation on your robot's state sequence.

### Can I use Cosmos 3 with robots that have discrete action spaces?

Yes, but you must encode discrete actions as continuous values. For example, if your robot has a binary gripper state (open/closed), represent this as a 0.0 or 1.0 float value within the action vector. The model expects continuous action tokens, so discrete modes should be one-hot encoded or mapped to continuous ranges that match the `raw_action_dim` specified in your request.

### Where should I store custom action trajectory files?

Store custom action files in the `cookbooks/cosmos3/generator/action/assets/` directory to match the repository's example structure, or reference any accessible path via the `action_path` parameter. The NumPy or JSON format must contain an array of vectors where each vector has length equal to `raw_action_dim`, ordered temporally for the `action_chunk_size` duration.