How to Use the Describe Anything Feature for Attribute Captioning in NVIDIA Cosmos

The Describe Anything feature in NVIDIA Cosmos generates structured JSON captions for marked subjects in images by creating a capability JSON file and running the cosmos_framework.scripts.inference module.

The NVIDIA Cosmos repository includes a multimodal Reasoner capability called "Describe Anything" that enables detailed attribute captioning for visual subjects. This feature processes marked images through a structured prompt interface defined in JSON configuration files and returns machine-readable descriptions suitable for downstream planning or visualization pipelines.

What Is the Describe Anything Capability

The Describe Anything capability is a built-in Reasoner prompt that instructs multimodal models to generate detailed captions for every marked subject in an image. According to the source code in cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md (lines 94-100), the prompt enforces a machine-readable output format requiring the model to return a JSON array of objects. Each object contains three fields: subject_id, category, and caption, making the output trivial to parse for downstream consumption.

The system relies on a capability JSON file that specifies both the natural-language instruction and the path to the target image. This file is consumed by the inference script at runtime to construct the multimodal request.

Prerequisites and Environment Setup

Before running the Describe Anything feature, you must install the Cosmos Framework package and configure the required environment variables. The inference script expects specific directory paths to locate inputs and write outputs.

Set up your environment from the repository root:


# Create and activate virtual environment

python -m venv .venv
source .venv/bin/activate

# Install the Cosmos Framework package

pip install -e .

Export the required environment variables that the inference script reads:

export COSMOS3_INPUT_DIR=$PWD
export COSMOS3_OUTPUT_ROOT=$PWD/outputs

Creating the Capability Configuration

The capability JSON file tells the inference engine which prompt to send and which image to analyze. You must create this file before running inference.

Create capabilities/describe_marked_subjects.json with the following structure:

{
  "name": "describe_marked_subjects",
  "prompt": "Please caption the notable attributes in the provided image. List and describe all marked subjects in the image with their categories and detailed captions using a json with keyword \"subject_id\", \"category\" and \"caption\".",
  "vision_path": "cookbooks/cosmos3/reasoner/assets/describe_anything.png"
}

Key JSON Fields

  • name: Identifies the capability for logging and output directory naming
  • prompt: The exact natural-language instruction sent to the model (as documented in reasoner_prompt_guide.md)
  • vision_path: Relative or absolute path to the image file containing marked subjects

Running the Inference Script

Once the capability JSON exists and environment variables are set, launch the inference entry point. The script cosmos_framework.scripts.inference parses the JSON, builds a multimodal request with the message structure [{role:"user", content:[{type:"image_url",...}, {type:"text",...}]}], and sends it to the configured LLM (such as Cosmos-3-Super or a locally-hosted vLLM instance).

Execute the inference run:

.venv/bin/python -m cosmos_framework.scripts.inference \
    -i "$COSMOS3_INPUT_DIR/capabilities/describe_marked_subjects.json" \
    -o "$COSMOS3_OUTPUT_ROOT/cosmos_framework_describe_marked_subjects"

The script processes the image and prompt, then writes the model's response to the specified output directory.

Inspecting the Output

The inference script generates structured output files in your designated output directory. The primary results appear in reasoner_text.txt within the run-specific subdirectory.

Check the generated captions:

cat $COSMOS3_OUTPUT_ROOT/cosmos_framework_describe_marked_subjects/reasoner_text.txt

The output contains a JSON array similar to:

[
  {"subject_id":"2","category":"ski","caption":"The ski is white with black bindings and shows signs of light use on the base."},
  {"subject_id":"4","category":"ski parka","caption":"This turquoise ski parka features a hood with fur trim and multiple zippered pockets."},
  {"subject_id":"5","category":"ski parka","caption":"The black ski parka has reflective strips and a reinforced waterproof shell."}
]

This structured format allows automated systems to extract specific attributes by subject_id or filter by category for downstream processing.

Using Custom Images

To caption your own images instead of the demo asset, modify the vision_path field in your capability JSON file. The inference pipeline accepts any image path accessible from the execution environment.

Update the JSON configuration:

{
  "name": "describe_marked_subjects",
  "prompt": "Please caption the notable attributes in the provided image. List and describe all marked subjects in the image with their categories and detailed captions using a json with keyword \"subject_id\", \"category\" and \"caption\".",
  "vision_path": "my_images/my_robot_scene.png"
}

All other steps remain identical. The model will analyze the new image and return subject-specific captions following the same JSON schema.

Summary

  • Describe Anything is a Reasoner prompt in NVIDIA Cosmos that generates structured attribute captions for marked image subjects.
  • Configuration requires creating a JSON file with the prompt text and image path, stored by default in capabilities/describe_marked_subjects.json.
  • The inference entry point cosmos_framework.scripts.inference consumes this JSON and outputs results to reasoner_text.txt in your specified output directory.
  • Output follows a strict JSON schema with subject_id, category, and caption fields for easy programmatic parsing.
  • You can substitute the demo image cookbooks/cosmos3/reasoner/assets/describe_anything.png with any custom image by updating the vision_path parameter.

Frequently Asked Questions

What output format does the Describe Anything feature return?

The feature returns a JSON array where each object contains three string fields: subject_id (identifier for the marked subject), category (object classification), and caption (detailed attribute description). This structure is enforced by the prompt template defined in cookbooks/cosmos3/reasoner/reasoner_prompt_guide.md, ensuring machine-readable output that requires no additional parsing logic.

Can I use custom images with the Describe Anything feature?

Yes. Simply modify the vision_path field in your capability JSON file to point to any image file accessible from your execution environment. The inference script cosmos_framework.scripts.inference loads the specified image and processes it using the same multimodal pipeline, returning subject-specific captions for your custom visual content.

Where is the inference script located in the NVIDIA Cosmos repository?

The inference script is implemented as a module in the Cosmos Framework package at cosmos_framework/scripts/inference.py. After installing the repository with pip install -e ., you invoke it via the module syntax: python -m cosmos_framework.scripts.inference. This script handles the complete pipeline from JSON parsing to model interaction and result serialization.

How do I modify the prompt for different captioning styles?

Edit the prompt field in your capability JSON file (capabilities/describe_marked_subjects.json). While the default prompt enforces specific JSON keywords (subject_id, category, caption) for structured output, you can adjust the descriptive instructions to emphasize different attributes such as color, texture, size, or spatial relationships while maintaining the required output format.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →