Implementing 2D Grounding and Bounding Box Localization with Cosmos 3
Cosmos 3's Reasoner surface converts an image-plus-text prompt into JSON bounding boxes by processing multimodal tokens through a Mixture-of-Transformers architecture that outputs spatial coordinates and labels.
The NVIDIA Cosmos repository provides a unified vision-language foundation model capable of 2D grounding and bounding box localization without requiring separate detection heads or post-processing pipelines. According to the source code in README.md (lines 153-154), the model accepts an image and natural language prompt, then returns a structured JSON list containing x, y, width, height, and label fields for each detected object.
How Cosmos 3 2D Grounding Works
The grounding pipeline operates through five distinct stages that share the model's core transformer backbone.
The Mixture-of-Transformers Architecture
Cosmos 3 utilizes a Mixture-of-Transformers (MoT) architecture that unifies reasoning and generation tasks within a single transformer core. As documented in the model architecture section (lines 70-75), the backbone employs 3-D rotary position embeddings (mRoPE) to encode spatial (x-y) and temporal (frame) axes simultaneously. This allows the same transformer weights to handle both causal self-attention for reasoning and diffusion for generation, eliminating the need for task-specific model forks.
From Image Tokens to JSON Bounding Boxes
The internal pipeline flow follows this sequence:
-
Pre-processing – Input images are resized to model-expected resolutions: 720p (1280×720), 480p (832×480), or 256p (320×192) as specified in the Input & Output documentation (lines 106-107).
-
Tokenisation –
AutoProcessor.from_pretrainedcreates a multimodal token stream combining vision encoder outputs for the image and text encoder outputs for the prompt. -
Fusion – The unified transformer processes the combined token sequence through multimodal attention blocks.
-
Box Prediction Head – A lightweight linear layer reads the final hidden states of image tokens and projects them to bounding box coordinates and optional class labels.
-
Post-processing – Raw coordinates are scaled to the original image dimensions and serialized as JSON.
Implementing 2D Grounding with the Cosmos 3 Reasoner
You can access the grounding capability through three primary interfaces: direct Python inference with Transformers, REST API calls via vLLM-Omni, or OpenAI-compatible clients.
Python Implementation with Hugging Face Transformers
The following implementation uses Cosmos3OmniForConditionalGeneration and AutoProcessor to perform local inference:
from pathlib import Path
import torch
from transformers import AutoProcessor, Cosmos3OmniForConditionalGeneration
# Model checkpoint (Nano is the lightweight 16B version)
model_id = "nvidia/Cosmos3-Nano"
# Load image and processor
image_path = Path("cookbooks/cosmos3/reasoner/assets/grounding_2d.png")
processor = AutoProcessor.from_pretrained(model_id)
# Build a chat-style request: image + grounding prompt
messages = [
{
"role": "user",
"content": [
{"type": "image", "path": str(image_path)},
{"type": "text", "text": "Find all red apples in the picture and return their boxes."},
],
}
]
# Tokenise for the model
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
).to("cuda", torch.bfloat16)
# Load the model (device-map can shard across GPUs)
model = Cosmos3OmniForConditionalGeneration.from_pretrained(
model_id,
dtype=torch.bfloat16,
device_map="auto",
)
# Generate the JSON response
generated = model.generate(**inputs, max_new_tokens=256)
output = processor.batch_decode(generated, skip_special_tokens=True)[0]
print("Grounding JSON:", output)
The processor automatically handles visual tokenisation and applies the correct resolution (720p by default). No additional box head code is required—the model's built-in decoder emits the JSON directly.
REST API Implementation with cURL and vLLM-Omni
For server-based deployments, use the vLLM-Omni endpoint with base64-encoded images:
# Encode the image as a base-64 data URI (Linux example)
IMAGE_B64=$(base64 -w 0 cookbooks/cosmos3/reasoner/assets/grounding_2d.png)
DATA_URI="data:image/png;base64,$IMAGE_B64"
curl -sS -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/cosmos3-nano-reasoner",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "'"$DATA_URI"'"}},
{"type": "text", "text": "Locate all blue boxes and output their coordinates as JSON."}
]}
],
"max_tokens": 256,
"stream": false,
"extra_body": {
"media_io_kwargs": {"image": {"size": {"shortest_edge": 720}}}
}
}'
The response contains a content field with the JSON-encoded boxes. The media_io_kwargs parameter controls image resolution when the default 720p is unsuitable (lines 85-90).
OpenAI-Compatible Client Integration
You can also use standard OpenAI clients by pointing to a local Cosmos 3 endpoint:
from openai import OpenAI
import base64
# Load image and create a data URI
with open("cookbooks/cosmos3/reasoner/assets/grounding_2d.png", "rb") as f:
img_uri = "data:image/png;base64," + base64.b64encode(f.read()).decode()
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
resp = client.chat.completions.create(
model="nvidia/cosmos3-nano-reasoner",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": img_uri}},
{"type": "text", "text": "Return the bounding boxes of every green object."}
]}
],
max_tokens=256,
stream=False,
extra_body={"media_io_kwargs": {"image": {"size": {"shortest_edge": 720}}}},
)
print(resp.choices[0].message.content) # → JSON list of boxes
Key Source Files and Reference Materials
Understanding the implementation requires referencing these specific files in the NVIDIA Cosmos repository:
README.md(lines 153-154): Documents the 2D grounding API contract and JSON output format.README.md(lines 70-75): Details the MoT architecture and mRoPE position embeddings.cookbooks/cosmos3/reasoner/run_with_cosmos_framework.ipynb: End-to-end notebook demonstrating Reasoner calls via the Cosmos Framework entrypoint and JSON parsing.cosmos_framework/scripts/inference.py: Implements the generic CLI that wraps processor loading and chat template construction.cookbooks/cosmos3/reasoner/assets/grounding_2d.png: Sample image used for testing grounding pipelines.
Input Resolution and Preprocessing Requirements
Cosmos 3 expects specific input resolutions for optimal performance. The media_io_kwargs parameter accepts a size dictionary with a shortest_edge key mapping to these supported values:
- 720p: 1280×720 pixels (default)
- 480p: 832×480 pixels
- 256p: 320×192 pixels
The processor automatically handles resizing and normalization, but providing images at these native resolutions minimizes preprocessing artifacts and improves bounding box accuracy.
Summary
- Cosmos 3 implements 2D grounding through its Reasoner surface, which shares the MoT backbone with captioning and generation tasks.
- Input format: Image plus text prompt; Output format: JSON list with
x,y,width,height, andlabelfields. - Three implementation paths: Direct Python with Transformers (
Cosmos3OmniForConditionalGeneration), REST API via vLLM-Omni, or OpenAI-compatible clients. - Resolution options: 720p (default), 480p, or 256p, controlled via
media_io_kwargsin API calls or processor configuration. - Key technical components: Vision encoder, text encoder, multimodal attention fusion, and a lightweight box-prediction head projecting to JSON coordinates.
Frequently Asked Questions
What JSON format does Cosmos 3 return for bounding boxes?
Cosmos 3 returns a JSON array where each object contains five fields: x and y for the top-left corner coordinates, width and height for the box dimensions, and label for the object class. These coordinates are automatically scaled to the original image dimensions during post-processing.
Can I use Cosmos 3 for 2D grounding without installing the full training framework?
Yes. You can deploy the model using vLLM-Omni or NVIDIA NIM containers, which expose OpenAI-compatible REST APIs. The nvidia/Cosmos3-Nano checkpoint (16B parameters) runs efficiently on consumer GPUs when using device_map="auto" for layer sharding.
How does the Mixture-of-Transformers architecture benefit grounding tasks?
The MoT architecture allows Cosmos 3 to use the same transformer weights for both understanding the visual scene (reasoning) and generating the structured JSON output (generation). This eliminates the need for separate detection heads and enables zero-shot grounding through natural language prompts rather than predefined class lists.
What is the difference between the Reasoner and standard generation modes?
The Reasoner mode activates causal self-attention pathways optimized for analysis and structured output (like bounding boxes), while standard generation modes use diffusion pathways for video synthesis. Both share the MoT backbone, but the Reasoner projects final hidden states through a box-prediction head rather than a video decoder.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →