# How to Use Cosmos 3 for Robot Action Prediction: Complete Policy Server Setup Guide

> Learn to set up the Cosmos 3 policy server for robot action prediction. Stream predicted actions to clients for simulation or real world execution.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-05

---

**Cosmos 3 provides a policy server that runs the Cosmos3-Nano-Policy-DROID model, streaming predicted robot actions to a client that executes them in simulation or real environments.**

The NVIDIA Cosmos repository enables end-to-end robot manipulation through a 16-billion parameter Vision-Language policy model. This guide explains how to deploy the **policy server** from the `cosmos-framework` package and connect it to a simulation client for closed-loop robot control.

## Architecture Overview

The Cosmos 3 robot action prediction system consists of three coordinated components working across Docker containers.

### Policy Server

The server loads the **Cosmos3-Nano-Policy-DROID** checkpoint from HuggingFace and exposes a TCP endpoint on port 8000. It receives observation-instruction requests and returns continuous action tokens including trajectory predictions and gripper states. The main implementation resides in `cosmos_framework.scripts.action_policy_server_robolab` within the `cosmos-framework` Docker image.

### Simulation Client

The RoboLab client runs the robot environment and forwards observations to the policy server. It applies returned action chunks to the simulator. The entry point is [`policies/cosmos3/run.py`](https://github.com/NVIDIA/cosmos/blob/main/policies/cosmos3/run.py) in the RoboLab repository.

### Docker Orchestration

GPU-enabled containers isolate dependencies for both server and client. The configuration handles CUDA drivers, HuggingFace cache sharing, and networking between components.

## Prerequisites and Environment Setup

Before launching the policy server, ensure you have the following:

- **HuggingFace access token** (`HF_TOKEN`) to download the Cosmos3-Nano-Policy-DROID model
- **CUDA 13 drivers** for the `cu130-train` dependency group (substitute with `cu128-train` for older systems)
- Docker with NVIDIA runtime support (`--runtime nvidia`)

Export your token before building containers:

```bash
export HF_TOKEN=<your_hf_token>

```

## Step-by-Step Implementation

### Building the Policy Server

Clone the framework repository and build the GPU-enabled image:

```bash
git clone https://github.com/NVIDIA/cosmos-framework.git
cd cosmos-framework
docker build -t cosmos-framework:latest .

```

### Launching the Server Container

Run the container with GPU access and mount the HuggingFace cache to avoid redundant downloads:

```bash
docker run -it \
  -e HF_HOME=/workspace/.cache/huggingface \
  -e HF_TOKEN=$HF_TOKEN \
  --net host \
  --rm \
  --runtime nvidia \
  -v .:/workspace \
  -v $HOME/.cache/huggingface:/root/.cache/huggingface \
  cosmos-framework:latest \
  bash -c '
    uv sync --all-extras --group=cu130-train --group=policy-server && \
    python -m cosmos_framework.scripts.action_policy_server_robolab --port 8000
  '

```

The server initializes on port 8000 and loads the model into GPU memory.

### Setting Up the RoboLab Client

From a separate terminal, clone and build the RoboLab environment:

```bash
git clone https://github.com/NVlabs/RoboLab.git
cd RoboLab
./docker/build_docker.sh latest
./docker/run_docker.sh latest

```

### Running Interactive Policy Evaluation

Inside the RoboLab container, execute a policy-driven task with visualization:

```bash
python policies/cosmos3/run.py --task BananaInBowlTask

```

The client connects to `localhost:8000`, streams camera images and robot state to the server, and applies the returned action chunks in real-time.

### Headless Batch Evaluation

For evaluation across multiple environments without a GUI:

```bash
python policies/cosmos3/run.py \
  --task BananaInBowlTask \
  --num-envs 10 \
  --headless

```

This runs 10 parallel simulations, prints per-environment metrics, and saves rollout videos to `policies/cosmos3/results/`.

## Key Implementation Details

### Data Flow

The per-step loop follows this pattern:

1. Client captures current RGB camera image, robot joint positions, and gripper state
2. Client transmits observations plus high-level instruction (e.g., "pick the banana") to the policy server
3. Server processes inputs through the 16B Vision-Language Transformer
4. Server returns a **predicted action chunk** (typically 10-30 frames of future trajectory)
5. Client executes actions in the simulator or forwards them to physical hardware

The model jointly processes image and text inputs to emit **future-observation + action** sequences, supporting both open-loop video generation and closed-loop robotic manipulation.

### Reference Documentation

For detailed configuration options, consult the cookbooks in the NVIDIA Cosmos repository:

- [`cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md) provides the complete server launch tutorial
- [`cookbooks/cosmos3/generator/action/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/README.md) covers action-generation recipes
- [`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md) contains throughput and latency metrics for the policy model

## Summary

- **Cosmos 3 robot action prediction** relies on a client-server architecture separating the policy model from the execution environment
- The policy server runs via `cosmos_framework.scripts.action_policy_server_robolab` and requires `HF_TOKEN` for model access
- Simulation clients use [`policies/cosmos3/run.py`](https://github.com/NVIDIA/cosmos/blob/main/policies/cosmos3/run.py) to connect to the server and execute actions in environments like `BananaInBowlTask`
- Docker containers with `--runtime nvidia` provide isolated GPU access for both components
- The system supports interactive visualization and headless batch evaluation via `--num-envs` and `--headless` flags

## Frequently Asked Questions

### What hardware requirements does the Cosmos 3 policy server need?

The server requires an NVIDIA GPU with CUDA 13 support for optimal performance with the `cu130-train` dependency group. If your system uses older drivers, substitute with the `cu128-train` group during the `uv sync` step. The 16-billion parameter model requires significant GPU memory; ensure your device has adequate VRAM.

### How does the policy server handle authentication?

The server authenticates with HuggingFace using the `HF_TOKEN` environment variable to download the Cosmos3-Nano-Policy-DROID checkpoint. You must export this token before launching the Docker container, and the container mounts your local HuggingFace cache to avoid repeated downloads.

### Can I use Cosmos 3 policy prediction with real robots?

Yes. While the examples use RoboLab simulation, the architecture supports real hardware. The client receives action chunks (trajectories and gripper states) from the server and can forward these to physical robot controllers instead of simulation environments. Ensure your client implementation handles the hardware interface and safety protocols.

### What is the difference between the policy server and video generation models?

The policy server runs the Cosmos3-Nano-Policy-DROID model specifically trained for robotic manipulation, outputting actionable trajectories. Video generation models in the Cosmos suite produce future observations without necessarily emitting control signals. The policy model combines both capabilities, predicting future observations alongside executable actions for closed-loop control.