Cosmos 3 Resolution, Aspect Ratio, and Frame Rate Settings: Complete Configuration Guide

Cosmos 3 supports three resolution tiers (256p, 480p, 720p), five aspect ratios (16:9, 4:3, 1:1, 3:4, 9:16), and four frame rates (10fps, 16fps, 24fps, 30fps), with defaults of 480p, 16:9, and 24fps for standard generation tasks.

NVIDIA Cosmos encodes specific resolution, aspect ratio, and frame rate tokens directly into its tokenizer and diffusion pipeline. Selecting from these supported values ensures optimal inference performance without costly token padding or server-side resampling that degrades generation quality, according to the NVIDIA/cosmos source code.

Supported Resolution, Aspect Ratio, and Frame Rate Values

Cosmos 3’s generator surface accepts a strict set of configuration values baked into the model architecture. These settings are defined in the repository’s README.md (lines 92–95) and map directly to the tokenizer’s discrete tokens.

Setting Supported Values Default
Resolution tier 256p, 480p, 720p 480p
Aspect ratio 16:9, 4:3, 1:1, 3:4, 9:16 16:9
Frame rate 10fps, 16fps, 24fps, 30fps 24fps

The generator always expects a size argument formatted as <width>x<height> (e.g., 1280x720). Supplying unsupported width-height combinations forces the server to up-sample or down-sample inputs internally, which reduces output fidelity.

The NVIDIA Cosmos repository provides specific guidance for mapping these parameters to common generation scenarios. The following recommendations balance visual quality, inference speed, and hardware utilization.

Use case Resolution Aspect ratio Frame rate Rationale
Text-to-image (single-frame) 480p 16:9 24fps (token only) Fast inference with sufficient pixels for visual tasks; FPS token is passed but irrelevant for static images.
Text-to-video (standard) 480p 16:9 24fps Matches cinematic standards and the UniPCMultistepScheduler default configuration.
High-fidelity video 720p 16:9 24fps Higher spatial resolution improves detail; positional embeddings are calibrated for 1280×720 inputs.
Fast prototyping 256p 1:1 or 4:3 10fps Minimal memory footprint for quick sanity checks and iterative development.
Portrait/mobile output 480p 9:16 or 3:4 24fps Native aspect-ratio tokens generate vertically-oriented video without post-processing cropping.

Configuring Settings in the Diffusion Pipeline

When instantiating the pipeline via Cosmos3OmniPipeline.from_pretrained, you pass height, width, and fps arguments that map directly to the supported tokens above. The default configuration in the Cosmos 3 cookbooks uses 720p and 24fps to demonstrate high-fidelity paths, as shown in cookbooks/cosmos3/generator/audiovisual/README.md.

High-Fidelity Video Generation (720p)

Use this configuration for demo or offline rendering when visual detail is critical. This example uses the 720p tier with 1280×1920 dimensions for 16:9 output at cinematic frame rates.

import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler
from diffusers.utils import export_to_video

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)
pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, flow_shift=10.0
)

result = pipe(
    prompt="A mobile robot navigates a warehouse aisle and stops at a shelf.",
    negative_prompt="",
    image=None,
    num_frames=189,               # ~7.9 seconds at 24 fps

    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    enable_sound=False,
    generator=torch.Generator(device="cuda").manual_seed(1234),
)

export_to_video(result.video, "cosmos3_t2v.mp4", fps=24, macro_block_size=1)

Standard Text-to-Video (480p)

For most production workloads, accept the 480p defaults. This balances generation speed and quality without requiring explicit height and width overrides beyond the standard pipeline initialization.

Fast Prototyping (256p)

Reduce compute requirements for debugging or rapid iteration by selecting the smallest resolution tier and lowest frame rate.


# 256p square video at 10 fps for rapid testing

result = pipe(
    prompt="A robot arm grasps a cube.",
    height=192,      # 256p supports 320×192; 192×192 gives 1:1 aspect

    width=192,
    fps=10,
    num_frames=50,   # 5 seconds at 10 fps

    num_inference_steps=20,
)

REST API Configuration (vLLM-Omni)

When using the vLLM-Omni server, encode both resolution and aspect ratio in the size parameter.

curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=A small warehouse robot moves a blue box across a clean floor." \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "guidance_scale=6.0" \
  -o cosmos3_t2v_output.mp4

Key Reference Files

File Relevance
README.md (lines 92–95) Enumerates the exact resolution tiers, aspect ratios, and frame rates supported by the tokenizer.
cookbooks/cosmos3/generator/audiovisual/README.md Demonstrates concrete usage patterns for 720p generation and API parameter formatting.
inference_benchmarks.md Provides latency tables across resolution tiers to guide hardware-specific configuration choices.

Summary

  • Default settings: Use 480p resolution, 16:9 aspect ratio, and 24fps for most text-to-image and text-to-video tasks.
  • High-fidelity output: Select 720p (1280×720) with 24fps when visual detail is critical, as positional embeddings are calibrated for this tier.
  • Rapid prototyping: Configure 256p with 10fps for low-resource testing and fast iteration cycles.
  • Aspect ratio flexibility: Choose 9:16 or 3:4 for portrait/mobile content to avoid post-processing; the model generates correctly oriented media natively.
  • Parameter formatting: Always provide dimensions as <width>x<height> (e.g., 1280x720) to prevent server-side resampling artifacts.

Frequently Asked Questions

What resolution settings does Cosmos 3 support?

Cosmos 3 supports three discrete resolution tiers: 256p, 480p, and 720p. These values correspond to specific pixel dimensions (such as 1280×720 for 720p at 16:9) that are hard-coded into the model’s tokenizer tokens, as documented in README.md.

Can I use custom frame rates outside the supported list?

No, you must select from 10fps, 16fps, 24fps, or 30fps. The diffusion pipeline’s UniPCMultistepScheduler and tokenizer are calibrated for these specific temporal tokens, and using unsupported values triggers internal resampling that degrades motion consistency.

How do I generate portrait-oriented video with Cosmos 3?

Set the aspect ratio to 9:16 or 3:4 and provide corresponding dimensions (e.g., 720x1280 for 9:16 at 720p). The model accepts aspect-ratio tokens directly, generating vertically-oriented video without requiring post-processing rotation or cropping.

What happens if I specify an unsupported resolution or aspect ratio?

The server will either reject the request or internally up-sample/down-sample your input to the nearest supported tier. According to the NVIDIA Cosmos source code, this forced resampling reduces generation quality and increases inference overhead, so always use supported <width>x<height> combinations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →