# How to Integrate Cosmos 3 with Custom Robotics Frameworks Using Action Inputs

> Integrate Cosmos 3 with custom robotics frameworks using action inputs. Leverage its OpenAI-compatible REST endpoint to stream joint commands to any HTTP-enabled framework.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-06

---

**Cosmos 3 exposes an OpenAI-compatible REST endpoint that streams action tokens (joint commands) when running as an action-policy server, allowing any robotics framework using HTTP or the OpenAI client library to consume trajectories without modifying the underlying model code.**

Cosmos 3, NVIDIA’s multimodal foundation model, functions as both a vision reasoner and an action generator for robotics pipelines. When integrating Cosmos 3 with custom robotics frameworks using action inputs, you deploy the model as an **action-policy server** that communicates via standard HTTP requests, then bridge the streamed JSON responses to your robot’s native controller. This approach decouples the inference engine from the execution environment, enabling seamless integration with ROS 2, PyBullet, or proprietary C++ stacks.

## Architecture Overview

The integration relies on two decoupled components: a server that hosts the Cosmos 3-Nano model and a client that translates streamed actions into robot commands.

### Action-Policy Server

The server runs inside a Docker container (NIM) or a local Python process (vLLM-Omni) and exposes an OpenAI-compatible endpoint at `http://localhost:8000/v1/chat/completions`. As implemented in `cosmos_framework.scripts.action_policy_server_robolab`, the server loads the Cosmos 3-Nano reasoning weights and accepts requests containing vision inputs (images or videos) paired with text prompts requesting action prediction. It streams back JSON objects containing the `action` array—typically a 9-dimensional vector representing joint positions for a 7-DOF arm plus gripper and base.

### Robot Client

The client is any Python-based robotics framework that knows how to read the JSON-L spec and call the OpenAI client library. Reference implementations are provided in `cookbooks/cosmos3/generator/action/run_id_with_cosmos_framework.ipynb` (inverse-dynamics) and `cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb` (forward-dynamics). The client parses the streamed `action` field and converts the list of floats into the robot’s native control format, such as ROS 2 `JointTrajectory` messages or PyBullet motor commands.

## Preparing the Request Payload

The core of the integration is the OpenAI-compatible `messages` payload. You must include a system prompt instructing the model to return machine-readable JSON, followed by a user message containing the vision input and task description.

```json
{
  "model": "nvidia/cosmos3-nano-reasoner",
  "messages": [
    {
      "role": "system",
      "content": "You are a robot controller. Output the next action as a JSON list of 9 floats under the key `action`."
    },
    {
      "role": "user",
      "content": [
        {
          "type": "video_url",
          "video_url": {"url": "file:///path/to/scene.mp4"}
        },
        {
          "type": "text",
          "text": "Predict the next joint trajectory for the robot."
        }
      ]
    }
  ],
  "stream": true,
  "extra_body": {
    "media_io_kwargs": {"video": {"fps": 4.0}}
  }
}

```

The `extra_body.media_io_kwargs` field controls video decoding parameters, such as frame sampling rate. The server expects the vision input to be a local file path or URL, and the text prompt to explicitly request action prediction.

## Deploying the Action-Policy Server

To serve the model locally, pull and run the NIM container as documented in [`cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md).

```bash
export CONTAINER_NAME="nvidia-cosmos3-reasoner"
export IMG_NAME="nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0"
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"

docker run -it --rm --name=$CONTAINER_NAME \
  --runtime=nvidia --gpus all \
  --shm-size=32GB \
  -e NGC_API_KEY=$NGC_API_KEY \
  -e NIM_MODEL_SIZE=nano \
  -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
  -u $(id -u) -p 8000:8000 $IMG_NAME

```

After startup, the server listens on `http://127.0.0.1:8000/v1`. The `NGC_API_KEY` is required only for pulling the container; the key is never stored in the repository source code.

## Building the Client Integration

Below is a minimal Python client that connects to the local server, streams the action tokens, and reconstructs the trajectory. This pattern works with any framework that can run Python and make HTTP requests.

```python
from openai import OpenAI
import json
import pathlib

# Point the OpenAI client at the local Cosmos 3 server

client = OpenAI(
    base_url="http://127.0.0.1:8000/v1",
    api_key="not-used"  # Ignored by the local server

)

# Configure the vision input path

vision_path = pathlib.Path("cookbooks/cosmos3/generator/action/assets/videos/av_0.mp4")

request = {
    "model": "nvidia/cosmos3-nano-reasoner",
    "messages": [
        {
            "role": "system",
            "content": "You are a robot controller. Return the next joint trajectory as a JSON list `action`."
        },
        {
            "role": "user",
            "content": [
                {
                    "type": "video_url",
                    "video_url": {"url": f"file://{vision_path.resolve()}"}
                },
                {
                    "type": "text",
                    "text": "Predict the next robot action."
                }
            ]
        }
    ],
    "max_tokens": 256,
    "stream": True,
    "extra_body": {
        "media_io_kwargs": {"video": {"fps": 4.0}}
    }
}

# Stream the response and accumulate action tokens

response = client.chat.completions.create(**request)
action = []

for chunk in response:
    content = chunk.choices[0].message.content
    if "action" in content:
        payload = json.loads(content)
        action.extend(payload["action"])

# `action` is now a list of floats ready for the robot controller

print("Predicted action trajectory:", action)

```

### Framework-Specific Hooks

- **ROS 2**: Convert the `action` list into a `sensor_msgs/JointState` or `trajectory_msgs/JointTrajectory` and publish to `/arm_controller/command`.
- **PyBullet**: Feed the action values directly into `p.setJointMotorControlArray` to update joint positions.
- **RoboLab**: Replace the networking layer in `action_policy_server_robolab` with your transport (ZeroMQ, gRPC) while preserving the JSON payload structure.

Because the contract is plain JSON over HTTP, you can embed the request inside a ROS 2 service call or a C++ bridge without modifying the server code.

## Understanding the Action-Input Specification

The **action-input spec** defines the contract between the client and the model. For inverse-dynamics (predicting actions from video), create a JSONL file where each line contains:

```json
{
  "vision_path": "cookbooks/cosmos3/generator/action/assets/videos/av_0.mp4",
  "action_path": null,
  "action_chunk_size": 60,
  "raw_action_dim": 9,
  "num_frames": 61
}

```

- **vision_path**: Path to the input video or image.
- **action_path**: Set to `null` for inverse-dynamics; for forward-dynamics (video generation), provide a JSON file containing the action trajectory.
- **action_chunk_size**: Number of future frames the model should predict.
- **raw_action_dim**: Dimensionality of the action vector (default 9 for 7-DOF arms plus gripper and base).
- **num_frames**: Total frames to process (typically chunk size plus one initial frame).

The notebooks `run_id_with_cosmos_framework.ipynb` and `run_fd_with_cosmos_framework.ipynb` demonstrate how to consume this spec for batch inference or real-time streaming.

## Summary

- Cosmos 3 acts as an **action generator** by exposing an OpenAI-compatible endpoint at `http://localhost:8000/v1/chat/completions`.
- The server runs via **NIM Docker containers** or **vLLM-Omni**, loading the `nvidia/cosmos3-nano-reasoner` model.
- Clients send vision inputs and receive streamed JSON containing `action` arrays, which map to joint-space commands.
- The **action-input spec** (`vision_path`, `action_path`, `raw_action_dim`, etc.) defines the inference parameters for inverse-dynamics and forward-dynamics tasks.
- Integration is framework-agnostic: any language with an OpenAI client library can consume the endpoint, and the raw action tokens can be translated to ROS 2, PyBullet, or custom controllers.

## Frequently Asked Questions

### What is the difference between inverse and forward dynamics in Cosmos 3?

**Inverse-dynamics** predicts the next robot action (joint trajectory) given a visual observation, which is the typical mode for closed-loop control. **Forward-dynamics** generates a video sequence given an action trajectory, useful for planning and simulation. The notebooks `run_id_with_cosmos_framework.ipynb` and `run_fd_with_cosmos_framework.ipynb` provide reference implementations for each mode, differing only in whether `action_path` is omitted or provided in the JSONL spec.

### Can I integrate Cosmos 3 with ROS 2?

Yes. Because the server uses a standard HTTP REST API, you can wrap the Python client code in a ROS 2 node that publishes `JointTrajectory` messages on your robot’s command topic. The JSON `action` array maps directly to the `points` field of the trajectory message, requiring no modification to the Cosmos 3 inference server.

### What hardware requirements are needed for the action policy server?

The server requires an NVIDIA GPU with sufficient VRAM to run the Cosmos 3-Nano model (exact VRAM depends on batch size and precision). When deploying via the NIM container, you must provide an `NGC_API_KEY` to pull the image, though the key is not stored in the repository. The client side can run on any CPU-only machine that can reach the server over the network.

### How do I format action tokens for a custom robot?

The model outputs a flat JSON list of floats under the key `action`. The dimensionality is controlled by the `raw_action_dim` parameter in your request (default 9). For a custom robot, ensure your client maps these indices to the correct joints (e.g., indices 0-6 for the arm, 7 for the gripper, 8 for the mobile base) before sending commands to the hardware.