# How to Use the Describe Anything Feature for Attribute Captioning in NVIDIA Cosmos

> Learn how to use the Describe Anything feature in NVIDIA Cosmos to generate structured JSON captions for image subjects. This guide explains the capability JSON file and inference module for attribute captioning.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-12

---

**The Describe Anything feature in NVIDIA Cosmos generates structured JSON captions for marked subjects in images by creating a capability JSON file and running the `cosmos_framework.scripts.inference` module.**

The NVIDIA Cosmos repository includes a multimodal Reasoner capability called "Describe Anything" that enables detailed attribute captioning for visual subjects. This feature processes marked images through a structured prompt interface defined in JSON configuration files and returns machine-readable descriptions suitable for downstream planning or visualization pipelines.

## What Is the Describe Anything Capability

The **Describe Anything** capability is a built-in Reasoner prompt that instructs multimodal models to generate detailed captions for every marked subject in an image. According to the source code in [`cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md) (lines 94-100), the prompt enforces a machine-readable output format requiring the model to return a JSON array of objects. Each object contains three fields: `subject_id`, `category`, and `caption`, making the output trivial to parse for downstream consumption.

The system relies on a **capability JSON file** that specifies both the natural-language instruction and the path to the target image. This file is consumed by the inference script at runtime to construct the multimodal request.

## Prerequisites and Environment Setup

Before running the Describe Anything feature, you must install the Cosmos Framework package and configure the required environment variables. The inference script expects specific directory paths to locate inputs and write outputs.

Set up your environment from the repository root:

```bash

# Create and activate virtual environment

python -m venv .venv
source .venv/bin/activate

# Install the Cosmos Framework package

pip install -e .

```

Export the required environment variables that the inference script reads:

```bash
export COSMOS3_INPUT_DIR=$PWD
export COSMOS3_OUTPUT_ROOT=$PWD/outputs

```

## Creating the Capability Configuration

The capability JSON file tells the inference engine which prompt to send and which image to analyze. You must create this file before running inference.

Create [`capabilities/describe_marked_subjects.json`](https://github.com/NVIDIA/cosmos/blob/main/capabilities/describe_marked_subjects.json) with the following structure:

```json
{
  "name": "describe_marked_subjects",
  "prompt": "Please caption the notable attributes in the provided image. List and describe all marked subjects in the image with their categories and detailed captions using a json with keyword \"subject_id\", \"category\" and \"caption\".",
  "vision_path": "cookbooks/cosmos3/reasoner/assets/describe_anything.png"
}

```

### Key JSON Fields

- **name**: Identifies the capability for logging and output directory naming
- **prompt**: The exact natural-language instruction sent to the model (as documented in [`reasoner_prompt_guide.md`](https://github.com/NVIDIA/cosmos/blob/main/reasoner_prompt_guide.md))
- **vision_path**: Relative or absolute path to the image file containing marked subjects

## Running the Inference Script

Once the capability JSON exists and environment variables are set, launch the inference entry point. The script `cosmos_framework.scripts.inference` parses the JSON, builds a multimodal request with the message structure `[{role:"user", content:[{type:"image_url",...}, {type:"text",...}]}]`, and sends it to the configured LLM (such as Cosmos-3-Super or a locally-hosted vLLM instance).

Execute the inference run:

```bash
.venv/bin/python -m cosmos_framework.scripts.inference \
    -i "$COSMOS3_INPUT_DIR/capabilities/describe_marked_subjects.json" \
    -o "$COSMOS3_OUTPUT_ROOT/cosmos_framework_describe_marked_subjects"

```

The script processes the image and prompt, then writes the model's response to the specified output directory.

## Inspecting the Output

The inference script generates structured output files in your designated output directory. The primary results appear in [`reasoner_text.txt`](https://github.com/NVIDIA/cosmos/blob/main/reasoner_text.txt) within the run-specific subdirectory.

Check the generated captions:

```bash
cat $COSMOS3_OUTPUT_ROOT/cosmos_framework_describe_marked_subjects/reasoner_text.txt

```

The output contains a JSON array similar to:

```json
[
  {"subject_id":"2","category":"ski","caption":"The ski is white with black bindings and shows signs of light use on the base."},
  {"subject_id":"4","category":"ski parka","caption":"This turquoise ski parka features a hood with fur trim and multiple zippered pockets."},
  {"subject_id":"5","category":"ski parka","caption":"The black ski parka has reflective strips and a reinforced waterproof shell."}
]

```

This structured format allows automated systems to extract specific attributes by `subject_id` or filter by `category` for downstream processing.

## Using Custom Images

To caption your own images instead of the demo asset, modify the `vision_path` field in your capability JSON file. The inference pipeline accepts any image path accessible from the execution environment.

Update the JSON configuration:

```json
{
  "name": "describe_marked_subjects",
  "prompt": "Please caption the notable attributes in the provided image. List and describe all marked subjects in the image with their categories and detailed captions using a json with keyword \"subject_id\", \"category\" and \"caption\".",
  "vision_path": "my_images/my_robot_scene.png"
}

```

All other steps remain identical. The model will analyze the new image and return subject-specific captions following the same JSON schema.

## Summary

- **Describe Anything** is a Reasoner prompt in NVIDIA Cosmos that generates structured attribute captions for marked image subjects.
- Configuration requires creating a JSON file with the prompt text and image path, stored by default in [`capabilities/describe_marked_subjects.json`](https://github.com/NVIDIA/cosmos/blob/main/capabilities/describe_marked_subjects.json).
- The inference entry point `cosmos_framework.scripts.inference` consumes this JSON and outputs results to [`reasoner_text.txt`](https://github.com/NVIDIA/cosmos/blob/main/reasoner_text.txt) in your specified output directory.
- Output follows a strict JSON schema with `subject_id`, `category`, and `caption` fields for easy programmatic parsing.
- You can substitute the demo image `cookbooks/cosmos3/reasoner/assets/describe_anything.png` with any custom image by updating the `vision_path` parameter.

## Frequently Asked Questions

### What output format does the Describe Anything feature return?

The feature returns a JSON array where each object contains three string fields: `subject_id` (identifier for the marked subject), `category` (object classification), and `caption` (detailed attribute description). This structure is enforced by the prompt template defined in [`cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md), ensuring machine-readable output that requires no additional parsing logic.

### Can I use custom images with the Describe Anything feature?

Yes. Simply modify the `vision_path` field in your capability JSON file to point to any image file accessible from your execution environment. The inference script `cosmos_framework.scripts.inference` loads the specified image and processes it using the same multimodal pipeline, returning subject-specific captions for your custom visual content.

### Where is the inference script located in the NVIDIA Cosmos repository?

The inference script is implemented as a module in the Cosmos Framework package at [`cosmos_framework/scripts/inference.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_framework/scripts/inference.py). After installing the repository with `pip install -e .`, you invoke it via the module syntax: `python -m cosmos_framework.scripts.inference`. This script handles the complete pipeline from JSON parsing to model interaction and result serialization.

### How do I modify the prompt for different captioning styles?

Edit the `prompt` field in your capability JSON file ([`capabilities/describe_marked_subjects.json`](https://github.com/NVIDIA/cosmos/blob/main/capabilities/describe_marked_subjects.json)). While the default prompt enforces specific JSON keywords (`subject_id`, `category`, `caption`) for structured output, you can adjust the descriptive instructions to emphasize different attributes such as color, texture, size, or spatial relationships while maintaining the required output format.