How to Use Cosmos 3 for Robot Action Prediction: Complete Policy Server Setup Guide
Cosmos 3 provides a policy server that runs the Cosmos3-Nano-Policy-DROID model, streaming predicted robot actions to a client that executes them in simulation or real environments.
The NVIDIA Cosmos repository enables end-to-end robot manipulation through a 16-billion parameter Vision-Language policy model. This guide explains how to deploy the policy server from the cosmos-framework package and connect it to a simulation client for closed-loop robot control.
Architecture Overview
The Cosmos 3 robot action prediction system consists of three coordinated components working across Docker containers.
Policy Server
The server loads the Cosmos3-Nano-Policy-DROID checkpoint from HuggingFace and exposes a TCP endpoint on port 8000. It receives observation-instruction requests and returns continuous action tokens including trajectory predictions and gripper states. The main implementation resides in cosmos_framework.scripts.action_policy_server_robolab within the cosmos-framework Docker image.
Simulation Client
The RoboLab client runs the robot environment and forwards observations to the policy server. It applies returned action chunks to the simulator. The entry point is policies/cosmos3/run.py in the RoboLab repository.
Docker Orchestration
GPU-enabled containers isolate dependencies for both server and client. The configuration handles CUDA drivers, HuggingFace cache sharing, and networking between components.
Prerequisites and Environment Setup
Before launching the policy server, ensure you have the following:
- HuggingFace access token (
HF_TOKEN) to download the Cosmos3-Nano-Policy-DROID model - CUDA 13 drivers for the
cu130-traindependency group (substitute withcu128-trainfor older systems) - Docker with NVIDIA runtime support (
--runtime nvidia)
Export your token before building containers:
export HF_TOKEN=<your_hf_token>
Step-by-Step Implementation
Building the Policy Server
Clone the framework repository and build the GPU-enabled image:
git clone https://github.com/NVIDIA/cosmos-framework.git
cd cosmos-framework
docker build -t cosmos-framework:latest .
Launching the Server Container
Run the container with GPU access and mount the HuggingFace cache to avoid redundant downloads:
docker run -it \
-e HF_HOME=/workspace/.cache/huggingface \
-e HF_TOKEN=$HF_TOKEN \
--net host \
--rm \
--runtime nvidia \
-v .:/workspace \
-v $HOME/.cache/huggingface:/root/.cache/huggingface \
cosmos-framework:latest \
bash -c '
uv sync --all-extras --group=cu130-train --group=policy-server && \
python -m cosmos_framework.scripts.action_policy_server_robolab --port 8000
'
The server initializes on port 8000 and loads the model into GPU memory.
Setting Up the RoboLab Client
From a separate terminal, clone and build the RoboLab environment:
git clone https://github.com/NVlabs/RoboLab.git
cd RoboLab
./docker/build_docker.sh latest
./docker/run_docker.sh latest
Running Interactive Policy Evaluation
Inside the RoboLab container, execute a policy-driven task with visualization:
python policies/cosmos3/run.py --task BananaInBowlTask
The client connects to localhost:8000, streams camera images and robot state to the server, and applies the returned action chunks in real-time.
Headless Batch Evaluation
For evaluation across multiple environments without a GUI:
python policies/cosmos3/run.py \
--task BananaInBowlTask \
--num-envs 10 \
--headless
This runs 10 parallel simulations, prints per-environment metrics, and saves rollout videos to policies/cosmos3/results/.
Key Implementation Details
Data Flow
The per-step loop follows this pattern:
- Client captures current RGB camera image, robot joint positions, and gripper state
- Client transmits observations plus high-level instruction (e.g., "pick the banana") to the policy server
- Server processes inputs through the 16B Vision-Language Transformer
- Server returns a predicted action chunk (typically 10-30 frames of future trajectory)
- Client executes actions in the simulator or forwards them to physical hardware
The model jointly processes image and text inputs to emit future-observation + action sequences, supporting both open-loop video generation and closed-loop robotic manipulation.
Reference Documentation
For detailed configuration options, consult the cookbooks in the NVIDIA Cosmos repository:
cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.mdprovides the complete server launch tutorialcookbooks/cosmos3/generator/action/README.mdcovers action-generation recipesinference_benchmarks.mdcontains throughput and latency metrics for the policy model
Summary
- Cosmos 3 robot action prediction relies on a client-server architecture separating the policy model from the execution environment
- The policy server runs via
cosmos_framework.scripts.action_policy_server_robolaband requiresHF_TOKENfor model access - Simulation clients use
policies/cosmos3/run.pyto connect to the server and execute actions in environments likeBananaInBowlTask - Docker containers with
--runtime nvidiaprovide isolated GPU access for both components - The system supports interactive visualization and headless batch evaluation via
--num-envsand--headlessflags
Frequently Asked Questions
What hardware requirements does the Cosmos 3 policy server need?
The server requires an NVIDIA GPU with CUDA 13 support for optimal performance with the cu130-train dependency group. If your system uses older drivers, substitute with the cu128-train group during the uv sync step. The 16-billion parameter model requires significant GPU memory; ensure your device has adequate VRAM.
How does the policy server handle authentication?
The server authenticates with HuggingFace using the HF_TOKEN environment variable to download the Cosmos3-Nano-Policy-DROID checkpoint. You must export this token before launching the Docker container, and the container mounts your local HuggingFace cache to avoid repeated downloads.
Can I use Cosmos 3 policy prediction with real robots?
Yes. While the examples use RoboLab simulation, the architecture supports real hardware. The client receives action chunks (trajectories and gripper states) from the server and can forward these to physical robot controllers instead of simulation environments. Ensure your client implementation handles the hardware interface and safety protocols.
What is the difference between the policy server and video generation models?
The policy server runs the Cosmos3-Nano-Policy-DROID model specifically trained for robotic manipulation, outputting actionable trajectories. Video generation models in the Cosmos suite produce future observations without necessarily emitting control signals. The policy model combines both capabilities, predicting future observations alongside executable actions for closed-loop control.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →