How to Use Cosmos 3 for Robot Action Prediction: Complete Policy Server Setup Guide

Cosmos 3 provides a policy server that runs the Cosmos3-Nano-Policy-DROID model, streaming predicted robot actions to a client that executes them in simulation or real environments.

The NVIDIA Cosmos repository enables end-to-end robot manipulation through a 16-billion parameter Vision-Language policy model. This guide explains how to deploy the policy server from the cosmos-framework package and connect it to a simulation client for closed-loop robot control.

Architecture Overview

The Cosmos 3 robot action prediction system consists of three coordinated components working across Docker containers.

Policy Server

The server loads the Cosmos3-Nano-Policy-DROID checkpoint from HuggingFace and exposes a TCP endpoint on port 8000. It receives observation-instruction requests and returns continuous action tokens including trajectory predictions and gripper states. The main implementation resides in cosmos_framework.scripts.action_policy_server_robolab within the cosmos-framework Docker image.

Simulation Client

The RoboLab client runs the robot environment and forwards observations to the policy server. It applies returned action chunks to the simulator. The entry point is policies/cosmos3/run.py in the RoboLab repository.

Docker Orchestration

GPU-enabled containers isolate dependencies for both server and client. The configuration handles CUDA drivers, HuggingFace cache sharing, and networking between components.

Prerequisites and Environment Setup

Before launching the policy server, ensure you have the following:

  • HuggingFace access token (HF_TOKEN) to download the Cosmos3-Nano-Policy-DROID model
  • CUDA 13 drivers for the cu130-train dependency group (substitute with cu128-train for older systems)
  • Docker with NVIDIA runtime support (--runtime nvidia)

Export your token before building containers:

export HF_TOKEN=<your_hf_token>

Step-by-Step Implementation

Building the Policy Server

Clone the framework repository and build the GPU-enabled image:

git clone https://github.com/NVIDIA/cosmos-framework.git
cd cosmos-framework
docker build -t cosmos-framework:latest .

Launching the Server Container

Run the container with GPU access and mount the HuggingFace cache to avoid redundant downloads:

docker run -it \
  -e HF_HOME=/workspace/.cache/huggingface \
  -e HF_TOKEN=$HF_TOKEN \
  --net host \
  --rm \
  --runtime nvidia \
  -v .:/workspace \
  -v $HOME/.cache/huggingface:/root/.cache/huggingface \
  cosmos-framework:latest \
  bash -c '
    uv sync --all-extras --group=cu130-train --group=policy-server && \
    python -m cosmos_framework.scripts.action_policy_server_robolab --port 8000
  '

The server initializes on port 8000 and loads the model into GPU memory.

Setting Up the RoboLab Client

From a separate terminal, clone and build the RoboLab environment:

git clone https://github.com/NVlabs/RoboLab.git
cd RoboLab
./docker/build_docker.sh latest
./docker/run_docker.sh latest

Running Interactive Policy Evaluation

Inside the RoboLab container, execute a policy-driven task with visualization:

python policies/cosmos3/run.py --task BananaInBowlTask

The client connects to localhost:8000, streams camera images and robot state to the server, and applies the returned action chunks in real-time.

Headless Batch Evaluation

For evaluation across multiple environments without a GUI:

python policies/cosmos3/run.py \
  --task BananaInBowlTask \
  --num-envs 10 \
  --headless

This runs 10 parallel simulations, prints per-environment metrics, and saves rollout videos to policies/cosmos3/results/.

Key Implementation Details

Data Flow

The per-step loop follows this pattern:

  1. Client captures current RGB camera image, robot joint positions, and gripper state
  2. Client transmits observations plus high-level instruction (e.g., "pick the banana") to the policy server
  3. Server processes inputs through the 16B Vision-Language Transformer
  4. Server returns a predicted action chunk (typically 10-30 frames of future trajectory)
  5. Client executes actions in the simulator or forwards them to physical hardware

The model jointly processes image and text inputs to emit future-observation + action sequences, supporting both open-loop video generation and closed-loop robotic manipulation.

Reference Documentation

For detailed configuration options, consult the cookbooks in the NVIDIA Cosmos repository:

Summary

  • Cosmos 3 robot action prediction relies on a client-server architecture separating the policy model from the execution environment
  • The policy server runs via cosmos_framework.scripts.action_policy_server_robolab and requires HF_TOKEN for model access
  • Simulation clients use policies/cosmos3/run.py to connect to the server and execute actions in environments like BananaInBowlTask
  • Docker containers with --runtime nvidia provide isolated GPU access for both components
  • The system supports interactive visualization and headless batch evaluation via --num-envs and --headless flags

Frequently Asked Questions

What hardware requirements does the Cosmos 3 policy server need?

The server requires an NVIDIA GPU with CUDA 13 support for optimal performance with the cu130-train dependency group. If your system uses older drivers, substitute with the cu128-train group during the uv sync step. The 16-billion parameter model requires significant GPU memory; ensure your device has adequate VRAM.

How does the policy server handle authentication?

The server authenticates with HuggingFace using the HF_TOKEN environment variable to download the Cosmos3-Nano-Policy-DROID checkpoint. You must export this token before launching the Docker container, and the container mounts your local HuggingFace cache to avoid repeated downloads.

Can I use Cosmos 3 policy prediction with real robots?

Yes. While the examples use RoboLab simulation, the architecture supports real hardware. The client receives action chunks (trajectories and gripper states) from the server and can forward these to physical robot controllers instead of simulation environments. Ensure your client implementation handles the hardware interface and safety protocols.

What is the difference between the policy server and video generation models?

The policy server runs the Cosmos3-Nano-Policy-DROID model specifically trained for robotic manipulation, outputting actionable trajectories. Video generation models in the Cosmos suite produce future observations without necessarily emitting control signals. The policy model combines both capabilities, predicting future observations alongside executable actions for closed-loop control.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →