# How Prompt Upsampling Works in the NVIDIA Cosmos 3 Generator: Default Parameters Explained

> Discover how NVIDIA Cosmos 3 uses prompt upsampling to expand short scene descriptions into detailed prompts. Learn about default parameters including max_tokens, temperature, top_p, and seed.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: deep-dive
- Published: 2026-06-06

---

**The Cosmos 3 generator automatically expands short scene descriptions into detailed, structured prompts using an internal autoregressive transformer with default sampling parameters—`max_tokens=20000`, `temperature=0.7`, `top_p=0.8`, and `seed=3407`—before passing the enhanced text to the diffusion transformer.**

The NVIDIA `cosmos` repository provides a multimodal generative framework that enhances brief user inputs through an automated pipeline called **prompt upsampling**. When you submit a concise text prompt to the Cosmos 3 generator—whether via the Diffusers integration or the vLLM‑Omni server—the system evaluates the input length and, if it falls below a heuristic threshold, invokes an internal upsampler to rewrite the text into a dense structured format optimized for high‑fidelity video, image, or audio synthesis.

## What Is Prompt Upsampling?

**Prompt upsampling** is a preprocessing stage that transforms a short, informal scene description into a richly structured prompt before the diffusion model executes. The upsampler treats your brief input as a seed and uses the same autoregressive transformer that powers the **Reasoner** surface to generate an expanded version that includes explicit cues for resolution, duration, style, and modality‑specific templates.

This process runs automatically within the generator runtime for both the Diffusers Python API and vLLM‑Omni backend integrations, ensuring consistency across different inference interfaces.

## Default Upsampling Parameters

The default sampling hyper‑parameters that control the length, randomness, and diversity of the expanded prompt are documented in the repository’s [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) at [lines 331‑342](https://github.com/NVIDIA/cosmos/blob/main/README.md#L331). When the upsampler invokes the language model, it uses the following values unless explicitly overridden in the request payload:

- **`max_tokens`**: `20000` — The maximum number of tokens the upsampler may generate.
- **`temperature`**: `0.7` — Controls the randomness of token selection.
- **`top_p`**: `0.8` — Nucleus sampling probability threshold.
- **`top_k`**: `20` — The number of highest‑probability tokens to consider at each step.
- **`repetition_penalty`**: `1.0` — Penalty factor for repeated phrases.
- **`presence_penalty`**: `1.5` — Penalty factor to encourage novel token introduction.
- **`seed`**: `3407` — Fixed random seed for reproducible upsampling.

These defaults are hardcoded into the example notebooks and server configurations unless you provide alternative values via the API.

## The Four‑Step Upsampling Pipeline

The prompt upsampling process follows a strict sequence implemented in the generator runtime:

1. **Input Validation** — The generator checks if the prompt length is below a heuristic threshold (roughly a few sentences). If so, it flags the input for upsampling.

2. **Language‑Model Generation** — The autoregressive transformer generates tokens up to the `max_tokens` limit using the sampling parameters above. Because the upsampler shares the same model weights as the Reasoner, it can incorporate context from any vision or action conditioning supplied alongside the text.

3. **Structured Formatting** — The raw model output is post‑processed into a **dense structured prompt** that wraps the expanded description with metadata templates specifying resolution, duration, and style.

4. **Diffusion Handoff** — The finalized prompt is passed to the diffusion transformer to produce the ultimate image, video, or audio output.

## Code Implementation: Diffusers and vLLM‑Omni

Below are minimal examples demonstrating how the default upsampling parameters apply automatically when using either the Diffusers Python API or a direct HTTP call to a vLLM‑Omni server.

### Diffusers Python API

In this example, loading the `Cosmos3OmniPipeline` triggers automatic prompt upsampling when the input text is brief. The defaults are applied internally unless you pass explicit sampling overrides.

```python
import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler

# Load the model (the upsampler runs automatically for short prompts)

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)
pipe.scheduler = UniPCMultistepScheduler.from_config(pipe.scheduler.config, flow_shift=10.0)

# Short prompt – the upsampler will expand it internally using the defaults above

prompt = "A robot picks up a red cube."

result = pipe(
    prompt=prompt,
    negative_prompt="",
    image=None,
    num_frames=189,
    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    enable_sound=False,
    add_resolution_template=False,
    add_duration_template=False,
    generator=torch.Generator(device="cuda").manual_seed(1234),  # optional seed

)

```

### vLLM‑Omni HTTP API

When calling the vLLM‑Omni server endpoint, the same upsampling defaults are active. You may omit the sampling fields to accept the defaults, or include them to override.

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt= A robot picks up a red cube." \
  --form-string "negative_prompt=" \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=35" \
  --form-string "guidance_scale=6.0" \
  --form-string "seed=3407" \        # default seed; can be omitted

  --form-string "extra_params={\"use_resolution_template\":false,\"use_duration_template\":false}" \
  -o cosmos3_output.mp4

```

## Key Source Files

Understanding the implementation details requires consulting the following files in the `NVIDIA/cosmos` repository:

- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** (lines 331‑342) — Documents the default upsampling parameters and their semantic roles in the generation pipeline.
- **`cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb`** — Demonstrates the Diffusers integration where upsampling occurs automatically for short prompts.
- **`cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb`** — Shows how to interact with the vLLM‑Omni server and override default sampling values.

## Summary

- **Prompt upsampling** automatically expands brief inputs into dense structured prompts using the same autoregressive transformer that powers the Reasoner component.
- **Default parameters** include `max_tokens=20000`, `temperature=0.7`, `top_p=0.8`, `top_k=20`, `repetition_penalty=1.0`, `presence_penalty=1.5`, and `seed=3407`.
- The pipeline executes **four distinct stages**: input validation, language‑model generation, structured formatting, and diffusion handoff.
- You can **override defaults** by passing explicit sampling fields in either the Diffusers API or the vLLM‑Omni HTTP request payload.
- The feature is **multimodal‑aware**, incorporating vision and action conditioning into the expanded prompt to maintain consistency with all inputs.

## Frequently Asked Questions

### Can I disable prompt upsampling in the Cosmos 3 generator?

No, the generator does not expose a direct flag to disable upsampling entirely. However, supplying a sufficiently detailed prompt that exceeds the internal heuristic length threshold will prevent the upsampler from triggering, causing the generator to use your original text verbatim.

### How does prompt upsampling affect inference latency?

The upsampling step adds a language‑model forward pass to the overall pipeline. With a `max_tokens` limit of 20,000 and typical prompt lengths well below that ceiling, the overhead is usually less than a few seconds on modern GPU hardware, though this varies based on the actual token count generated.

### Can I use different sampling parameters for upsampling?

Yes. While the system applies the defaults documented in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) lines 331‑342, you can override any parameter—such as `temperature`, `top_p`, or `seed`—by including it in your request payload. In the Diffusers API, pass these as keyword arguments; in vLLM‑Omni, include them as form fields or JSON keys in the POST request.

### What determines whether a prompt gets upsampled?

The generator applies a heuristic threshold that roughly corresponds to "a few sentences." If your input falls below this length, the system automatically routes it through the autoregressive upsampler. The exact token count threshold is internal to the runtime and may adjust based on the presence of multimodal conditioning inputs.