# How Prompt Upsampling Works in NVIDIA Cosmos 3 Generator for Expanding Scene Descriptions

> Discover how prompt upsampling in NVIDIA Cosmos 3 expands scene descriptions by loading pre-computed LLM-expanded details for richer textual inputs during generation.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: deep-dive
- Published: 2026-06-14

---

**Prompt upsampling in Cosmos 3 works by loading pre-computed, LLM-expanded scene descriptions from JSON assets and cycling through them during generation to enrich the model's textual inputs.**

The NVIDIA Cosmos repository includes a sophisticated **prompt upsampling** mechanism designed to expand terse scene descriptions into detailed, varied prompts for video generation. This technique leverages offline-computed paraphrases stored in repository assets to provide the Cosmos 3 generator with richer textual context without requiring real-time LLM inference.

## The Prompt Upsampling Pipeline

The upsampling workflow follows a four-stage process that decouples expansion logic from generation inference:

1. **Base prompt definition** – Task-specific templates (image-to-video or video-to-video) establish the foundation descriptions.
2. **Offline expansion** – A language model pre-computes multiple rewritten variants of each base prompt, creating diverse linguistic expressions of the same scene.
3. **Asset storage** – These expansions are serialized into JSON files located in the Physics IQ evaluation assets directory.
4. **Runtime selection** – The generator loads these assets and cycles through the upsampled variants when producing video segments.

### Pre-Computed Prompt Storage

According to the Cosmos source code, the upsampled prompts reside in `evaluation/cosmos3/Physics_IQ/assets/`. The repository maintains two primary collections:

- [`i2v_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/i2v_prompts.json) – Contains **198 upsampled** image-to-video prompts
- [`v2v_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/v2v_prompts.json) – Contains **198 upsampled** video-to-video prompts

## Loading Upsampled Prompts in Python

The generator notebook implements prompt selection through simple modulo indexing against the pre-loaded JSON arrays. This ensures deterministic cycling through the 198 available variants for each task.

```python
import json
from pathlib import Path

# Load the upsampled I2V (image-to-video) prompts

i2v_path = Path(__file__).parent.parent / "evaluation" / "cosmos3" / "Physics_IQ" / "assets" / "i2v_prompts.json"
with i2v_path.open() as f:
    i2v_prompts = json.load(f)          # 198 upsampled prompts

# Load the upsampled V2V (video-to-video) prompts

v2v_path = Path(__file__).parent.parent / "evaluation" / "cosmos3" / "Physics_IQ" / "assets" / "v2v_prompts.json"
with v2v_path.open() as f:
    v2v_prompts = json.load(f)          # 198 upsampled prompts

def get_prompt(task: str, idx: int) -> str:
    """Return an upsampled prompt for the given task and index."""
    if task == "i2v":
        return i2v_prompts[idx % len(i2v_prompts)]
    elif task == "v2v":
        return v2v_prompts[idx % len(v2v_prompts)]
    else:
        raise ValueError("Unsupported task")

# Example: feed a prompt to the generator

scene_idx = 5
prompt = get_prompt("i2v", scene_idx)
generated_video = cosmos_generator.generate_video(prompt)  # ← pseudo‑API call

```

### Selection Logic

The `get_prompt()` function uses modulo arithmetic (`idx % len(i2v_prompts)`) to cycle through the 198 available upsampled descriptions. This approach guarantees that generation jobs receive varied textual inputs while maintaining reproducibility through deterministic indexing.

## Key Files in the Upsampling Workflow

Understanding the repository structure is essential for implementing custom prompt upsampling. The following files constitute the complete pipeline:

- **[`evaluation/cosmos3/Physics_IQ/assets/i2v_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/evaluation/cosmos3/Physics_IQ/assets/i2v_prompts.json)** – Stores 198 pre-expanded image-to-video scene descriptions
- **[`evaluation/cosmos3/Physics_IQ/assets/v2v_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/evaluation/cosmos3/Physics_IQ/assets/v2v_prompts.json)** – Stores 198 pre-expanded video-to-video scene descriptions  
- **`cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb`** – Jupyter notebook demonstrating generator integration with upsampled prompts
- **[`cookbooks/cosmos3/generator/transfer/preview_helpers.py`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/preview_helpers.py)** – Utility functions for previewing videos generated with upsampled prompts

## Summary

- Prompt upsampling in Cosmos 3 relies on **pre-computed JSON assets** rather than real-time LLM inference, eliminating latency during video generation.
- The repository stores **198 upsampled variants** for both image-to-video and video-to-video tasks in the Physics IQ evaluation assets.
- The generator uses **modulo indexing** to cycle through prompt variations deterministically based on scene indices.
- Implementation requires loading JSON assets from `evaluation/cosmos3/Physics_IQ/assets/` and selecting prompts by task type and index to feed enriched descriptions into the model.

## Frequently Asked Questions

### What is the purpose of prompt upsampling in Cosmos 3?

Prompt upsampling expands concise scene descriptions into detailed, linguistically diverse variants. By feeding the generator pre-expanded prompts from [`i2v_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/i2v_prompts.json) or [`v2v_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/v2v_prompts.json), the model receives richer textual cues that improve scene detail and variation in output videos.

### Where are the upsampled prompts stored in the repository?

The upsampled prompts are stored in the Physics IQ evaluation assets directory at `evaluation/cosmos3/Physics_IQ/assets/`. Specifically, [`i2v_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/i2v_prompts.json) contains image-to-video prompts and [`v2v_prompts.json`](https://github.com/NVIDIA/cosmos/blob/main/v2v_prompts.json) contains video-to-video prompts, each holding 198 pre-computed expansions.

### Can I generate new upsampled prompts instead of using the provided JSON files?

Yes, though the current implementation uses offline-computed prompts. You could generate new variants by applying a language model to paraphrase and expand base descriptions, then serialize them to the same JSON structure. The generator notebook expects the same list format found in the existing asset files.

### How does the generator select which upsampled prompt to use?

The selection mechanism uses modulo arithmetic to cycle through the 198 available prompts based on the scene index. As shown in `get_prompt()`, the function returns `i2v_prompts[idx % len(i2v_prompts)]`, ensuring deterministic rotation through the upsampled pool for consistent generation workflows.