# Implementing 2D Grounding and Bounding Box Localization with Cosmos 3

> Implement 2D grounding and bounding box localization with Cosmos 3. Our Reasoner surface converts image-text prompts into JSON bounding boxes using a Mixture-of-Transformers.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Cosmos 3's Reasoner surface converts an image-plus-text prompt into JSON bounding boxes by processing multimodal tokens through a Mixture-of-Transformers architecture that outputs spatial coordinates and labels.**

The NVIDIA Cosmos repository provides a unified vision-language foundation model capable of **2D grounding and bounding box localization** without requiring separate detection heads or post-processing pipelines. According to the source code in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) (lines 153-154), the model accepts an image and natural language prompt, then returns a structured JSON list containing `x`, `y`, `width`, `height`, and `label` fields for each detected object.

## How Cosmos 3 2D Grounding Works

The grounding pipeline operates through five distinct stages that share the model's core transformer backbone.

### The Mixture-of-Transformers Architecture

Cosmos 3 utilizes a **Mixture-of-Transformers (MoT)** architecture that unifies reasoning and generation tasks within a single transformer core. As documented in the model architecture section (lines 70-75), the backbone employs **3-D rotary position embeddings (mRoPE)** to encode spatial (x-y) and temporal (frame) axes simultaneously. This allows the same transformer weights to handle both causal self-attention for reasoning and diffusion for generation, eliminating the need for task-specific model forks.

### From Image Tokens to JSON Bounding Boxes

The internal pipeline flow follows this sequence:

1. **Pre-processing** – Input images are resized to model-expected resolutions: 720p (1280×720), 480p (832×480), or 256p (320×192) as specified in the Input & Output documentation (lines 106-107).

2. **Tokenisation** – `AutoProcessor.from_pretrained` creates a multimodal token stream combining vision encoder outputs for the image and text encoder outputs for the prompt.

3. **Fusion** – The unified transformer processes the combined token sequence through multimodal attention blocks.

4. **Box Prediction Head** – A lightweight linear layer reads the final hidden states of image tokens and projects them to bounding box coordinates and optional class labels.

5. **Post-processing** – Raw coordinates are scaled to the original image dimensions and serialized as JSON.

## Implementing 2D Grounding with the Cosmos 3 Reasoner

You can access the grounding capability through three primary interfaces: direct Python inference with Transformers, REST API calls via vLLM-Omni, or OpenAI-compatible clients.

### Python Implementation with Hugging Face Transformers

The following implementation uses `Cosmos3OmniForConditionalGeneration` and `AutoProcessor` to perform local inference:

```python
from pathlib import Path
import torch
from transformers import AutoProcessor, Cosmos3OmniForConditionalGeneration

# Model checkpoint (Nano is the lightweight 16B version)

model_id = "nvidia/Cosmos3-Nano"

# Load image and processor

image_path = Path("cookbooks/cosmos3/reasoner/assets/grounding_2d.png")
processor = AutoProcessor.from_pretrained(model_id)

# Build a chat-style request: image + grounding prompt

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "path": str(image_path)},
            {"type": "text", "text": "Find all red apples in the picture and return their boxes."},
        ],
    }
]

# Tokenise for the model

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt",
    return_dict=True,
).to("cuda", torch.bfloat16)

# Load the model (device-map can shard across GPUs)

model = Cosmos3OmniForConditionalGeneration.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
)

# Generate the JSON response

generated = model.generate(**inputs, max_new_tokens=256)
output = processor.batch_decode(generated, skip_special_tokens=True)[0]

print("Grounding JSON:", output)

```

The processor automatically handles visual tokenisation and applies the correct resolution (720p by default). No additional box head code is required—the model's built-in decoder emits the JSON directly.

### REST API Implementation with cURL and vLLM-Omni

For server-based deployments, use the vLLM-Omni endpoint with base64-encoded images:

```bash

# Encode the image as a base-64 data URI (Linux example)

IMAGE_B64=$(base64 -w 0 cookbooks/cosmos3/reasoner/assets/grounding_2d.png)
DATA_URI="data:image/png;base64,$IMAGE_B64"

curl -sS -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "nvidia/cosmos3-nano-reasoner",
        "messages": [
          {"role": "system", "content": "You are a helpful assistant."},
          {"role": "user", "content": [
              {"type": "image_url", "image_url": {"url": "'"$DATA_URI"'"}},
              {"type": "text", "text": "Locate all blue boxes and output their coordinates as JSON."}
          ]}
        ],
        "max_tokens": 256,
        "stream": false,
        "extra_body": {
          "media_io_kwargs": {"image": {"size": {"shortest_edge": 720}}}
        }
      }'

```

The response contains a `content` field with the JSON-encoded boxes. The `media_io_kwargs` parameter controls image resolution when the default 720p is unsuitable (lines 85-90).

### OpenAI-Compatible Client Integration

You can also use standard OpenAI clients by pointing to a local Cosmos 3 endpoint:

```python
from openai import OpenAI
import base64

# Load image and create a data URI

with open("cookbooks/cosmos3/reasoner/assets/grounding_2d.png", "rb") as f:
    img_uri = "data:image/png;base64," + base64.b64encode(f.read()).decode()

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")

resp = client.chat.completions.create(
    model="nvidia/cosmos3-nano-reasoner",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": [
            {"type": "image_url", "image_url": {"url": img_uri}},
            {"type": "text", "text": "Return the bounding boxes of every green object."}
        ]}
    ],
    max_tokens=256,
    stream=False,
    extra_body={"media_io_kwargs": {"image": {"size": {"shortest_edge": 720}}}},
)

print(resp.choices[0].message.content)   # → JSON list of boxes

```

## Key Source Files and Reference Materials

Understanding the implementation requires referencing these specific files in the NVIDIA Cosmos repository:

- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** (lines 153-154): Documents the 2D grounding API contract and JSON output format.
- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** (lines 70-75): Details the MoT architecture and mRoPE position embeddings.
- **`cookbooks/cosmos3/reasoner/run_with_cosmos_framework.ipynb`**: End-to-end notebook demonstrating Reasoner calls via the Cosmos Framework entrypoint and JSON parsing.
- **[`cosmos_framework/scripts/inference.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/inference.py)**: Implements the generic CLI that wraps processor loading and chat template construction.
- **`cookbooks/cosmos3/reasoner/assets/grounding_2d.png`**: Sample image used for testing grounding pipelines.

## Input Resolution and Preprocessing Requirements

Cosmos 3 expects specific input resolutions for optimal performance. The `media_io_kwargs` parameter accepts a `size` dictionary with a `shortest_edge` key mapping to these supported values:

- **720p**: 1280×720 pixels (default)
- **480p**: 832×480 pixels
- **256p**: 320×192 pixels

The processor automatically handles resizing and normalization, but providing images at these native resolutions minimizes preprocessing artifacts and improves bounding box accuracy.

## Summary

- **Cosmos 3 implements 2D grounding** through its Reasoner surface, which shares the MoT backbone with captioning and generation tasks.
- **Input format**: Image plus text prompt; **Output format**: JSON list with `x`, `y`, `width`, `height`, and `label` fields.
- **Three implementation paths**: Direct Python with Transformers (`Cosmos3OmniForConditionalGeneration`), REST API via vLLM-Omni, or OpenAI-compatible clients.
- **Resolution options**: 720p (default), 480p, or 256p, controlled via `media_io_kwargs` in API calls or processor configuration.
- **Key technical components**: Vision encoder, text encoder, multimodal attention fusion, and a lightweight box-prediction head projecting to JSON coordinates.

## Frequently Asked Questions

### What JSON format does Cosmos 3 return for bounding boxes?

Cosmos 3 returns a JSON array where each object contains five fields: `x` and `y` for the top-left corner coordinates, `width` and `height` for the box dimensions, and `label` for the object class. These coordinates are automatically scaled to the original image dimensions during post-processing.

### Can I use Cosmos 3 for 2D grounding without installing the full training framework?

Yes. You can deploy the model using **vLLM-Omni** or **NVIDIA NIM** containers, which expose OpenAI-compatible REST APIs. The `nvidia/Cosmos3-Nano` checkpoint (16B parameters) runs efficiently on consumer GPUs when using `device_map="auto"` for layer sharding.

### How does the Mixture-of-Transformers architecture benefit grounding tasks?

The MoT architecture allows Cosmos 3 to use the same transformer weights for both understanding the visual scene (reasoning) and generating the structured JSON output (generation). This eliminates the need for separate detection heads and enables zero-shot grounding through natural language prompts rather than predefined class lists.

### What is the difference between the Reasoner and standard generation modes?

The **Reasoner** mode activates causal self-attention pathways optimized for analysis and structured output (like bounding boxes), while standard generation modes use diffusion pathways for video synthesis. Both share the MoT backbone, but the Reasoner projects final hidden states through a box-prediction head rather than a video decoder.