How to Use Image-to-Video Generation with Conditioning Frames in Cosmos 3

Cosmos 3 generates video from a conditioning image by encoding the frame through an autoregressive transformer and denoising latent video tokens via a diffusion transformer, exposed through the Cosmos3OmniPipeline, vLLM-Omni HTTP API, or Cosmos Framework CLI.

Image-to-video generation with conditioning frames in Cosmos 3 enables you to animate a still image into a full motion sequence using NVIDIA's diffusion-based transformer architecture. The nvidia/Cosmos3-Nano and nvidia/Cosmos3-Super checkpoints implement this through a unified Omni pipeline that encodes the conditioning image before denoising latent video tokens. According to the NVIDIA/cosmos source code, you can run this workflow through the Diffusers Python API, an OpenAI-compatible vLLM-Omni server, or the Cosmos Framework CLI entrypoint.

Architecture of Image-to-Video Generation in Cosmos 3

Cosmos 3 operates in diffusion mode. An autoregressive transformer first encodes the conditioning image, and a diffusion transformer denoises a sequence of latent video tokens until a complete MP4 is produced. When an image is supplied through any of the supported interfaces, the model treats frame 0 as that image and synthesizes the remaining frames.

Core Components

  • Cosmos 3-Omni pipeline – The unified Diffusers pipeline (Cosmos3OmniPipeline) houses the transformer, tokenizers, and schedulers. It loads the checkpoint and runs the diffusion process.
  • input_reference field – Holds the conditioning image (or video for other modes) in API requests. When an image is supplied, the model treats frame 0 as that image.
  • size, num_frames, and fps – Control output resolution, length, and frame rate. By default, Cosmos 3 generates 189 frames at 24 FPS (approximately 8 seconds).
  • extra_params – Optional field that allows you to disable guardrails, tweak template use, or set per-request options.

Diffusers Image-to-Video Pipeline with Conditioning Frames

The pure-Python Diffusers workflow loads Cosmos3OmniPipeline and passes the conditioning image directly to the image parameter of the pipe call. Lines 35-55 of the main README.md illustrate this workflow.

import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler
from diffusers.utils import export_to_video
from PIL import Image

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, flow_shift=10.0
)

# Load the conditioning image

cond_img = Image.open("robot_start.png").convert("RGB")

result = pipe(
    prompt="A robot arm picks up a cup and places it on a table.",
    image=cond_img,                     # Conditioning frame

    num_frames=189,                     # Default length

    height=720, width=1280,             # Resolution tier

    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    generator=torch.Generator(device="cuda").manual_seed(42),
)

export_to_video(result.video, "robot_i2v.mp4", fps=24)

The image argument accepts a PIL Image object, which the pipeline encodes as frame 0. The scheduler is explicitly set to UniPCMultistepScheduler with flow_shift=10.0 for optimal denoising. Outputs are saved via export_to_video.

vLLM-Omni Image-to-Video API with Conditioning Frames

For OpenAI-compatible server deployments, vLLM-Omni exposes a synchronous POST /v1/videos/sync endpoint. You send the conditioning image in the input_reference multipart field, and the server returns MP4 bytes synchronously. According to the vLLM-Omni notes on conditioning image resolution in README.md, the pipeline forces the top-level size to match the conditioning image resolution to avoid reflection or padding.

curl -sS -X POST http://localhost:8000/v1/videos/sync \
  -F "prompt=A robot arm lifts a block and moves it onto a platform." \
  -F "size=1280x720" \
  -F "num_frames=189" \
  -F "fps=24" \
  -F "guidance_scale=6.0" \
  -F "input_reference=@robot_start.png" \
  -o robot_i2v_vllm.mp4

The input_reference parameter uploads the conditioning image and maps it to frame 0 in the generated sequence.

Cosmos Framework CLI for Image-to-Video Generation

The Cosmos Framework provides a CLI entrypoint through cosmos_framework.scripts.inference that forwards the same JSON payload to either the Diffusers or vLLM-Omni backend. Use the --image flag to pass the conditioning frame.

cosmos_framework \
  --model nvidia/Cosmos3-Nano \
  --task generator \
  --prompt "A small drone flies through a forest and circles a tree." \
  --image robot_start.png \
  --num_frames 189 \
  --resolution 720p \
  --output robot_i2v_framework.mp4

This command is documented in the generator cookbook at cookbooks/cosmos3/generator/audiovisual/README.md.

Key Source Files

The following files in the NVIDIA/cosmos repository define the exact contracts and examples for image-to-video generation with conditioning frames:

  • README.md (repo root) – Describes the image-to-video mode, request fields, and example code.
  • cookbooks/cosmos3/generator/audiovisual/README.md – Details generator usage, including the image-to-video flag and sound options.
  • cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb – Notebook running the Diffusers pipeline end-to-end.
  • cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb – Notebook showing the vLLM-Omni HTTP request.
  • cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb – Notebook using the Cosmos Framework CLI.

Summary

  • Cosmos 3 uses a diffusion transformer to generate video conditioned on a single image, treating the input as frame 0.
  • The Cosmos3OmniPipeline accepts a conditioning image via the image parameter in Diffusers.
  • For vLLM-Omni, upload the image via the input_reference multipart field to POST /v1/videos/sync.
  • The Cosmos Framework CLI uses the --image flag to pass conditioning frames through cosmos_framework.
  • Default generation produces 189 frames at 24 FPS; match the output resolution to the conditioning image dimensions to prevent padding.

Frequently Asked Questions

What checkpoints support image-to-video generation in Cosmos 3?

The NVIDIA/cosmos repository provides nvidia/Cosmos3-Nano and nvidia/Cosmos3-Super checkpoints. Both load into the Cosmos3OmniPipeline and support single-image conditioning for video generation.

How does the model handle the conditioning frame during generation?

The autoregressive transformer encodes the supplied conditioning image as frame 0. A diffusion transformer then denoises the remaining latent video tokens, synthesizing the subsequent frames without reflecting or padding the input.

Can I adjust the video length and frame rate?

Yes. Set num_frames and fps in the Diffusers pipeline, or include num_frames and fps fields in the vLLM-Omni request. The default is 189 frames at 24 FPS, producing approximately 8 seconds of video.

Does the vLLM-Omni server use the same endpoint for image-to-video and text-to-video?

Yes. The synchronous video endpoint at POST /v1/videos/sync handles both modalities. When you include the input_reference field with an image, the server treats the request as image-to-video generation instead of pure text-to-video.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →