How Prompt Upsampling Works in NVIDIA Cosmos 3 Generator for Expanding Scene Descriptions

Prompt upsampling in Cosmos 3 works by loading pre-computed, LLM-expanded scene descriptions from JSON assets and cycling through them during generation to enrich the model's textual inputs.

The NVIDIA Cosmos repository includes a sophisticated prompt upsampling mechanism designed to expand terse scene descriptions into detailed, varied prompts for video generation. This technique leverages offline-computed paraphrases stored in repository assets to provide the Cosmos 3 generator with richer textual context without requiring real-time LLM inference.

The Prompt Upsampling Pipeline

The upsampling workflow follows a four-stage process that decouples expansion logic from generation inference:

  1. Base prompt definition – Task-specific templates (image-to-video or video-to-video) establish the foundation descriptions.
  2. Offline expansion – A language model pre-computes multiple rewritten variants of each base prompt, creating diverse linguistic expressions of the same scene.
  3. Asset storage – These expansions are serialized into JSON files located in the Physics IQ evaluation assets directory.
  4. Runtime selection – The generator loads these assets and cycles through the upsampled variants when producing video segments.

Pre-Computed Prompt Storage

According to the Cosmos source code, the upsampled prompts reside in evaluation/cosmos3/Physics_IQ/assets/. The repository maintains two primary collections:

Loading Upsampled Prompts in Python

The generator notebook implements prompt selection through simple modulo indexing against the pre-loaded JSON arrays. This ensures deterministic cycling through the 198 available variants for each task.

import json
from pathlib import Path

# Load the upsampled I2V (image-to-video) prompts

i2v_path = Path(__file__).parent.parent / "evaluation" / "cosmos3" / "Physics_IQ" / "assets" / "i2v_prompts.json"
with i2v_path.open() as f:
    i2v_prompts = json.load(f)          # 198 upsampled prompts

# Load the upsampled V2V (video-to-video) prompts

v2v_path = Path(__file__).parent.parent / "evaluation" / "cosmos3" / "Physics_IQ" / "assets" / "v2v_prompts.json"
with v2v_path.open() as f:
    v2v_prompts = json.load(f)          # 198 upsampled prompts

def get_prompt(task: str, idx: int) -> str:
    """Return an upsampled prompt for the given task and index."""
    if task == "i2v":
        return i2v_prompts[idx % len(i2v_prompts)]
    elif task == "v2v":
        return v2v_prompts[idx % len(v2v_prompts)]
    else:
        raise ValueError("Unsupported task")

# Example: feed a prompt to the generator

scene_idx = 5
prompt = get_prompt("i2v", scene_idx)
generated_video = cosmos_generator.generate_video(prompt)  # ← pseudo‑API call

Selection Logic

The get_prompt() function uses modulo arithmetic (idx % len(i2v_prompts)) to cycle through the 198 available upsampled descriptions. This approach guarantees that generation jobs receive varied textual inputs while maintaining reproducibility through deterministic indexing.

Key Files in the Upsampling Workflow

Understanding the repository structure is essential for implementing custom prompt upsampling. The following files constitute the complete pipeline:

Summary

  • Prompt upsampling in Cosmos 3 relies on pre-computed JSON assets rather than real-time LLM inference, eliminating latency during video generation.
  • The repository stores 198 upsampled variants for both image-to-video and video-to-video tasks in the Physics IQ evaluation assets.
  • The generator uses modulo indexing to cycle through prompt variations deterministically based on scene indices.
  • Implementation requires loading JSON assets from evaluation/cosmos3/Physics_IQ/assets/ and selecting prompts by task type and index to feed enriched descriptions into the model.

Frequently Asked Questions

What is the purpose of prompt upsampling in Cosmos 3?

Prompt upsampling expands concise scene descriptions into detailed, linguistically diverse variants. By feeding the generator pre-expanded prompts from i2v_prompts.json or v2v_prompts.json, the model receives richer textual cues that improve scene detail and variation in output videos.

Where are the upsampled prompts stored in the repository?

The upsampled prompts are stored in the Physics IQ evaluation assets directory at evaluation/cosmos3/Physics_IQ/assets/. Specifically, i2v_prompts.json contains image-to-video prompts and v2v_prompts.json contains video-to-video prompts, each holding 198 pre-computed expansions.

Can I generate new upsampled prompts instead of using the provided JSON files?

Yes, though the current implementation uses offline-computed prompts. You could generate new variants by applying a language model to paraphrase and expand base descriptions, then serialize them to the same JSON structure. The generator notebook expects the same list format found in the existing asset files.

How does the generator select which upsampled prompt to use?

The selection mechanism uses modulo arithmetic to cycle through the 198 available prompts based on the scene index. As shown in get_prompt(), the function returns i2v_prompts[idx % len(i2v_prompts)], ensuring deterministic rotation through the upsampled pool for consistent generation workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →