# How to Use Image-to-Video Generation with Conditioning Frames in Cosmos 3

> Learn to generate video from images using conditioning frames in Cosmos 3 through Pipelines, APIs, or CLI. Explore advanced image-to-video techniques with NVIDIA's Cosmos 3 framework.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-05

---

**Cosmos 3 generates video from a conditioning image by encoding the frame through an autoregressive transformer and denoising latent video tokens via a diffusion transformer, exposed through the `Cosmos3OmniPipeline`, vLLM-Omni HTTP API, or Cosmos Framework CLI.**

Image-to-video generation with conditioning frames in Cosmos 3 enables you to animate a still image into a full motion sequence using NVIDIA's diffusion-based transformer architecture. The `nvidia/Cosmos3-Nano` and `nvidia/Cosmos3-Super` checkpoints implement this through a unified Omni pipeline that encodes the conditioning image before denoising latent video tokens. According to the NVIDIA/cosmos source code, you can run this workflow through the Diffusers Python API, an OpenAI-compatible vLLM-Omni server, or the Cosmos Framework CLI entrypoint.

## Architecture of Image-to-Video Generation in Cosmos 3

Cosmos 3 operates in **diffusion mode**. An autoregressive transformer first encodes the conditioning image, and a diffusion transformer denoises a sequence of latent video tokens until a complete MP4 is produced. When an image is supplied through any of the supported interfaces, the model treats frame 0 as that image and synthesizes the remaining frames.

### Core Components

- **Cosmos 3-Omni pipeline** – The unified Diffusers pipeline (`Cosmos3OmniPipeline`) houses the transformer, tokenizers, and schedulers. It loads the checkpoint and runs the diffusion process.
- **`input_reference` field** – Holds the conditioning image (or video for other modes) in API requests. When an image is supplied, the model treats frame 0 as that image.
- **`size`, `num_frames`, and `fps`** – Control output resolution, length, and frame rate. By default, Cosmos 3 generates **189 frames at 24 FPS** (approximately 8 seconds).
- **`extra_params`** – Optional field that allows you to disable guardrails, tweak template use, or set per-request options.

## Diffusers Image-to-Video Pipeline with Conditioning Frames

The pure-Python Diffusers workflow loads `Cosmos3OmniPipeline` and passes the conditioning image directly to the `image` parameter of the `pipe` call. Lines 35-55 of the main [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) illustrate this workflow.

```python
import torch
from diffusers import Cosmos3OmniPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler
from diffusers.utils import export_to_video
from PIL import Image

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

pipe.scheduler = UniPCMultistepScheduler.from_config(
    pipe.scheduler.config, flow_shift=10.0
)

# Load the conditioning image

cond_img = Image.open("robot_start.png").convert("RGB")

result = pipe(
    prompt="A robot arm picks up a cup and places it on a table.",
    image=cond_img,                     # Conditioning frame

    num_frames=189,                     # Default length

    height=720, width=1280,             # Resolution tier

    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    generator=torch.Generator(device="cuda").manual_seed(42),
)

export_to_video(result.video, "robot_i2v.mp4", fps=24)

```

The `image` argument accepts a PIL `Image` object, which the pipeline encodes as frame 0. The scheduler is explicitly set to `UniPCMultistepScheduler` with `flow_shift=10.0` for optimal denoising. Outputs are saved via `export_to_video`.

## vLLM-Omni Image-to-Video API with Conditioning Frames

For OpenAI-compatible server deployments, vLLM-Omni exposes a synchronous `POST /v1/videos/sync` endpoint. You send the conditioning image in the `input_reference` multipart field, and the server returns MP4 bytes synchronously. According to the vLLM-Omni notes on conditioning image resolution in [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md), the pipeline forces the top-level `size` to match the conditioning image resolution to avoid reflection or padding.

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  -F "prompt=A robot arm lifts a block and moves it onto a platform." \
  -F "size=1280x720" \
  -F "num_frames=189" \
  -F "fps=24" \
  -F "guidance_scale=6.0" \
  -F "input_reference=@robot_start.png" \
  -o robot_i2v_vllm.mp4

```

The `input_reference` parameter uploads the conditioning image and maps it to frame 0 in the generated sequence.

## Cosmos Framework CLI for Image-to-Video Generation

The Cosmos Framework provides a CLI entrypoint through `cosmos_framework.scripts.inference` that forwards the same JSON payload to either the Diffusers or vLLM-Omni backend. Use the `--image` flag to pass the conditioning frame.

```bash
cosmos_framework \
  --model nvidia/Cosmos3-Nano \
  --task generator \
  --prompt "A small drone flies through a forest and circles a tree." \
  --image robot_start.png \
  --num_frames 189 \
  --resolution 720p \
  --output robot_i2v_framework.mp4

```

This command is documented in the generator cookbook at [`cookbooks/cosmos3/generator/audiovisual/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/README.md).

## Key Source Files

The following files in the NVIDIA/cosmos repository define the exact contracts and examples for image-to-video generation with conditioning frames:

- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** (repo root) – Describes the image-to-video mode, request fields, and example code.
- **[`cookbooks/cosmos3/generator/audiovisual/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/README.md)** – Details generator usage, including the image-to-video flag and sound options.
- **`cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb`** – Notebook running the Diffusers pipeline end-to-end.
- **`cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb`** – Notebook showing the vLLM-Omni HTTP request.
- **`cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb`** – Notebook using the Cosmos Framework CLI.

## Summary

- Cosmos 3 uses a diffusion transformer to generate video conditioned on a single image, treating the input as frame 0.
- The `Cosmos3OmniPipeline` accepts a conditioning image via the `image` parameter in Diffusers.
- For vLLM-Omni, upload the image via the `input_reference` multipart field to `POST /v1/videos/sync`.
- The Cosmos Framework CLI uses the `--image` flag to pass conditioning frames through `cosmos_framework`.
- Default generation produces **189 frames at 24 FPS**; match the output resolution to the conditioning image dimensions to prevent padding.

## Frequently Asked Questions

### What checkpoints support image-to-video generation in Cosmos 3?

The NVIDIA/cosmos repository provides `nvidia/Cosmos3-Nano` and `nvidia/Cosmos3-Super` checkpoints. Both load into the `Cosmos3OmniPipeline` and support single-image conditioning for video generation.

### How does the model handle the conditioning frame during generation?

The autoregressive transformer encodes the supplied conditioning image as frame 0. A diffusion transformer then denoises the remaining latent video tokens, synthesizing the subsequent frames without reflecting or padding the input.

### Can I adjust the video length and frame rate?

Yes. Set `num_frames` and `fps` in the Diffusers pipeline, or include `num_frames` and `fps` fields in the vLLM-Omni request. The default is 189 frames at 24 FPS, producing approximately 8 seconds of video.

### Does the vLLM-Omni server use the same endpoint for image-to-video and text-to-video?

Yes. The synchronous video endpoint at `POST /v1/videos/sync` handles both modalities. When you include the `input_reference` field with an image, the server treats the request as image-to-video generation instead of pure text-to-video.