# Cosmos 3 Action Policy Models for Robot Manipulation: Architecture and Deployment

> Explore Cosmos 3 action policy models for robot manipulation. Convert vision-language prompts into executable robot trajectories with the NVIDIA/cosmos repository. Deploy DROID arm joint commands from scene images and text inst...

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: architecture
- Published: 2026-06-13

---

**Cosmos 3 action policy models convert vision-language prompts into executable robot trajectories, enabling the Cosmos3-Nano-Policy-DROID model to generate 10-dimensional DROID arm joint commands from scene images and text instructions.**

Cosmos 3 introduces a specialized family of action-policy models that transform the world-model framework into a robot controller. According to the NVIDIA Cosmos repository, these models leverage a unified Mixture-of-Transformers architecture to bridge high-level language instructions with low-level motor execution. The flagship **Cosmos3-Nano-Policy-DROID** model represents a 16-billion-parameter vision-language system specifically trained for manipulation tasks.

## Mixture-of-Transformers Architecture for Robot Control

The policy models inherit the same **Mixture-of-Transformers (MoT)** architecture used throughout Cosmos 3, as listed in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) lines 80-86. This design enables the model to process multimodal inputs and generate coherent action trajectories rather than purely visual content.

### Core Components in Policy Mode

When operating in policy mode, the model functions as a generator conditioned on visual and textual inputs. The architecture consists of four key components:

- **Autoregressive transformer**: Processes concatenated text-vision token streams to understand task instructions and scene context
- **Diffusion transformer (DM)**: Denoises latent action sequences while producing visualization frames, yielding coherent 10-dimensional DROID arm joint commands
- **Multimodal attention layers**: Share weights across vision, text, and action tokens, enabling joint reasoning about visual state and motor commands  
- **3-D multi-dimensional RoPE**: Encodes spatial-temporal relationships through 3-D rotary position embeddings, ensuring generated actions respect the geometry of the scene

### Safety Guardrails

Built-in guardrails filter harmful prompts and blur faces in generated videos. These safety checks operate by default during inference but can be disabled per-request if required for specific research applications.

## Cosmos3-Nano-Policy-DROID Specifications

The flagship **Cosmos3-Nano-Policy-DROID** model contains approximately 16 billion parameters and is highlighted in the repository's model table as a "Vision-language robot policy for DROID manipulation and control." Unlike standard video generation models, this system predicts action trajectories—specifically sequences of 10-dimensional DROID arm joint commands—while simultaneously producing short rollout videos that visualize the predicted motion.

The model operates as a conditional generator: it takes an initial scene image and a textual instruction (e.g., "Pick up the red block and place it on the blue platform"), then denoises a latent representation of the action trajectory through the diffusion transformer pipeline. The **action_chunk_size** parameter (typically set to 20) controls the temporal horizon of predicted motion.

## Deployment and Integration

Developers can deploy Cosmos 3 action policy models through two primary interfaces: the Cosmos Framework policy server or the vLLM-Omni API.

### Running the Policy Server

The policy server runs as a lightweight Python script within the Cosmos Framework Docker container. According to [`cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md) lines 57-64, launch the server with:

```bash
python -m cosmos_framework.scripts.action_policy_server_robolab \
  --port 8000

```

Once the server is active on port 8000, RoboLab clients can stream actions to simulated or real robots.

### RoboLab Client Integration

Connect a RoboLab task to the policy server using the client script described in [`run_policy_with_cosmos_framework.md`](https://github.com/NVIDIA/cosmos/blob/main/run_policy_with_cosmos_framework.md) lines 64-90:

```bash
python policies/cosmos3/run.py --task BananaInBowlTask

```

This setup enables real-time action streaming from the Cosmos 3 policy model to robot hardware.

### vLLM-Omni API Access

For programmatic access, the policy surface exposes an OpenAI-compatible endpoint through vLLM-Omni. As documented in [`cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md) lines 84-102, send requests using curl:

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=Pick up the red block and place it on the blue platform." \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=35" \
  --form-string "guidance_scale=6.0" \
  --form-string 'extra_params={"action_mode":"policy","domain_name":"droid","raw_action_dim":10,"action_chunk_size":20}' \
  -o policy_output.mp4

```

The `extra_params` JSON payload configures the policy mode, selects the DROID embodiment, and specifies the raw action dimensionality and chunk size.

### Python Client Example

Alternatively, use the OpenAI SDK to interact with the policy endpoint:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")

resp = client.videos.create(
    model="nvidia/Cosmos3-Nano",
    prompt="Pick up the red block and place it on the blue platform.",
    size="1280x720",
    num_frames=189,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    extra_params={"action_mode": "policy", "domain_name": "droid",
                  "raw_action_dim": 10, "action_chunk_size": 20},
)

with open("policy_output.mp4", "wb") as f:
    f.write(resp.video)

```

The response contains an MP4 video visualizing the predicted motion. When using the asynchronous `/v1/videos` endpoint, a companion JSON file provides the raw action sequence.

## Performance Characteristics

Inference benchmarks documented in [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) lines 352-356 provide latency and throughput expectations for production deployments. The action-policy mode maintains competitive performance when serving through vLLM-Omni, supporting scalable robot control pipelines.

## Summary

- **Cosmos3-Nano-Policy-DROID** is a 16-billion-parameter vision-language model that generates robot actions from text and image inputs
- The **Mixture-of-Transformers** architecture processes multimodal tokens and denoises latent action trajectories using a diffusion transformer
- Action outputs follow the **DROID** embodiment format with 10-dimensional joint commands and configurable action chunk sizes
- Deployment options include the **Cosmos Framework policy server** for RoboLab integration and **vLLM-Omni** for REST API access
- Built-in **guardrails** provide safety filtering and face blurring capabilities

## Frequently Asked Questions

### What is the parameter count of Cosmos 3 policy models?

The flagship policy model, **Cosmos3-Nano-Policy-DROID**, contains approximately 16 billion parameters. This makes it the largest model in the Cosmos 3 action-policy family while maintaining efficient inference through the Mixture-of-Transformers architecture.

### How does the policy model generate robot actions?

The model uses a **diffusion transformer** to denoise a latent representation of the action trajectory. Conditioned on the initial scene image and text instruction, it predicts sequences of 10-dimensional DROID arm joint commands while simultaneously generating a visualization video that shows the predicted robot motion.

### Can I run the policy model without the Cosmos Framework?

Yes. While the Cosmos Framework provides the recommended policy server (`cosmos_framework.scripts.action_policy_server_robolab`), you can also access the model through the **vLLM-Omni** OpenAI-compatible API. This allows direct HTTP requests using standard tools like curl or the OpenAI Python SDK, making it possible to integrate with custom robotics stacks outside the official framework.

### What safety mechanisms protect against harmful outputs?

Cosmos 3 policy models include **built-in guardrails** that filter harmful prompts and automatically blur faces in generated videos. These safety checks operate by default during inference but can be disabled per-request when necessary for specific research applications.