How to Integrate Cosmos 3 with Custom Robotics Frameworks Using Action Inputs
Cosmos 3 exposes an OpenAI-compatible REST endpoint that streams action tokens (joint commands) when running as an action-policy server, allowing any robotics framework using HTTP or the OpenAI client library to consume trajectories without modifying the underlying model code.
Cosmos 3, NVIDIA’s multimodal foundation model, functions as both a vision reasoner and an action generator for robotics pipelines. When integrating Cosmos 3 with custom robotics frameworks using action inputs, you deploy the model as an action-policy server that communicates via standard HTTP requests, then bridge the streamed JSON responses to your robot’s native controller. This approach decouples the inference engine from the execution environment, enabling seamless integration with ROS 2, PyBullet, or proprietary C++ stacks.
Architecture Overview
The integration relies on two decoupled components: a server that hosts the Cosmos 3-Nano model and a client that translates streamed actions into robot commands.
Action-Policy Server
The server runs inside a Docker container (NIM) or a local Python process (vLLM-Omni) and exposes an OpenAI-compatible endpoint at http://localhost:8000/v1/chat/completions. As implemented in cosmos_framework.scripts.action_policy_server_robolab, the server loads the Cosmos 3-Nano reasoning weights and accepts requests containing vision inputs (images or videos) paired with text prompts requesting action prediction. It streams back JSON objects containing the action array—typically a 9-dimensional vector representing joint positions for a 7-DOF arm plus gripper and base.
Robot Client
The client is any Python-based robotics framework that knows how to read the JSON-L spec and call the OpenAI client library. Reference implementations are provided in cookbooks/cosmos3/generator/action/run_id_with_cosmos_framework.ipynb (inverse-dynamics) and cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb (forward-dynamics). The client parses the streamed action field and converts the list of floats into the robot’s native control format, such as ROS 2 JointTrajectory messages or PyBullet motor commands.
Preparing the Request Payload
The core of the integration is the OpenAI-compatible messages payload. You must include a system prompt instructing the model to return machine-readable JSON, followed by a user message containing the vision input and task description.
{
"model": "nvidia/cosmos3-nano-reasoner",
"messages": [
{
"role": "system",
"content": "You are a robot controller. Output the next action as a JSON list of 9 floats under the key `action`."
},
{
"role": "user",
"content": [
{
"type": "video_url",
"video_url": {"url": "file:///path/to/scene.mp4"}
},
{
"type": "text",
"text": "Predict the next joint trajectory for the robot."
}
]
}
],
"stream": true,
"extra_body": {
"media_io_kwargs": {"video": {"fps": 4.0}}
}
}
The extra_body.media_io_kwargs field controls video decoding parameters, such as frame sampling rate. The server expects the vision input to be a local file path or URL, and the text prompt to explicitly request action prediction.
Deploying the Action-Policy Server
To serve the model locally, pull and run the NIM container as documented in cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md.
export CONTAINER_NAME="nvidia-cosmos3-reasoner"
export IMG_NAME="nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0"
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
docker run -it --rm --name=$CONTAINER_NAME \
--runtime=nvidia --gpus all \
--shm-size=32GB \
-e NGC_API_KEY=$NGC_API_KEY \
-e NIM_MODEL_SIZE=nano \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-u $(id -u) -p 8000:8000 $IMG_NAME
After startup, the server listens on http://127.0.0.1:8000/v1. The NGC_API_KEY is required only for pulling the container; the key is never stored in the repository source code.
Building the Client Integration
Below is a minimal Python client that connects to the local server, streams the action tokens, and reconstructs the trajectory. This pattern works with any framework that can run Python and make HTTP requests.
from openai import OpenAI
import json
import pathlib
# Point the OpenAI client at the local Cosmos 3 server
client = OpenAI(
base_url="http://127.0.0.1:8000/v1",
api_key="not-used" # Ignored by the local server
)
# Configure the vision input path
vision_path = pathlib.Path("cookbooks/cosmos3/generator/action/assets/videos/av_0.mp4")
request = {
"model": "nvidia/cosmos3-nano-reasoner",
"messages": [
{
"role": "system",
"content": "You are a robot controller. Return the next joint trajectory as a JSON list `action`."
},
{
"role": "user",
"content": [
{
"type": "video_url",
"video_url": {"url": f"file://{vision_path.resolve()}"}
},
{
"type": "text",
"text": "Predict the next robot action."
}
]
}
],
"max_tokens": 256,
"stream": True,
"extra_body": {
"media_io_kwargs": {"video": {"fps": 4.0}}
}
}
# Stream the response and accumulate action tokens
response = client.chat.completions.create(**request)
action = []
for chunk in response:
content = chunk.choices[0].message.content
if "action" in content:
payload = json.loads(content)
action.extend(payload["action"])
# `action` is now a list of floats ready for the robot controller
print("Predicted action trajectory:", action)
Framework-Specific Hooks
- ROS 2: Convert the
actionlist into asensor_msgs/JointStateortrajectory_msgs/JointTrajectoryand publish to/arm_controller/command. - PyBullet: Feed the action values directly into
p.setJointMotorControlArrayto update joint positions. - RoboLab: Replace the networking layer in
action_policy_server_robolabwith your transport (ZeroMQ, gRPC) while preserving the JSON payload structure.
Because the contract is plain JSON over HTTP, you can embed the request inside a ROS 2 service call or a C++ bridge without modifying the server code.
Understanding the Action-Input Specification
The action-input spec defines the contract between the client and the model. For inverse-dynamics (predicting actions from video), create a JSONL file where each line contains:
{
"vision_path": "cookbooks/cosmos3/generator/action/assets/videos/av_0.mp4",
"action_path": null,
"action_chunk_size": 60,
"raw_action_dim": 9,
"num_frames": 61
}
- vision_path: Path to the input video or image.
- action_path: Set to
nullfor inverse-dynamics; for forward-dynamics (video generation), provide a JSON file containing the action trajectory. - action_chunk_size: Number of future frames the model should predict.
- raw_action_dim: Dimensionality of the action vector (default 9 for 7-DOF arms plus gripper and base).
- num_frames: Total frames to process (typically chunk size plus one initial frame).
The notebooks run_id_with_cosmos_framework.ipynb and run_fd_with_cosmos_framework.ipynb demonstrate how to consume this spec for batch inference or real-time streaming.
Summary
- Cosmos 3 acts as an action generator by exposing an OpenAI-compatible endpoint at
http://localhost:8000/v1/chat/completions. - The server runs via NIM Docker containers or vLLM-Omni, loading the
nvidia/cosmos3-nano-reasonermodel. - Clients send vision inputs and receive streamed JSON containing
actionarrays, which map to joint-space commands. - The action-input spec (
vision_path,action_path,raw_action_dim, etc.) defines the inference parameters for inverse-dynamics and forward-dynamics tasks. - Integration is framework-agnostic: any language with an OpenAI client library can consume the endpoint, and the raw action tokens can be translated to ROS 2, PyBullet, or custom controllers.
Frequently Asked Questions
What is the difference between inverse and forward dynamics in Cosmos 3?
Inverse-dynamics predicts the next robot action (joint trajectory) given a visual observation, which is the typical mode for closed-loop control. Forward-dynamics generates a video sequence given an action trajectory, useful for planning and simulation. The notebooks run_id_with_cosmos_framework.ipynb and run_fd_with_cosmos_framework.ipynb provide reference implementations for each mode, differing only in whether action_path is omitted or provided in the JSONL spec.
Can I integrate Cosmos 3 with ROS 2?
Yes. Because the server uses a standard HTTP REST API, you can wrap the Python client code in a ROS 2 node that publishes JointTrajectory messages on your robot’s command topic. The JSON action array maps directly to the points field of the trajectory message, requiring no modification to the Cosmos 3 inference server.
What hardware requirements are needed for the action policy server?
The server requires an NVIDIA GPU with sufficient VRAM to run the Cosmos 3-Nano model (exact VRAM depends on batch size and precision). When deploying via the NIM container, you must provide an NGC_API_KEY to pull the image, though the key is not stored in the repository. The client side can run on any CPU-only machine that can reach the server over the network.
How do I format action tokens for a custom robot?
The model outputs a flat JSON list of floats under the key action. The dimensionality is controlled by the raw_action_dim parameter in your request (default 9). For a custom robot, ensure your client maps these indices to the correct joints (e.g., indices 0-6 for the arm, 7 for the gripper, 8 for the mobile base) before sending commands to the hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →