# How to Use Video-to-Video Transformation with Prompt-Guided Conditioning in vLLM-Omni

> Learn video-to-video transformation with prompt-guided conditioning in vLLM-Omni. Regenerate frames using POST requests to the sync endpoint.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: tutorial
- Published: 2026-06-06

---

**Submit a `POST` request to the `/v1/videos/sync` endpoint with an input video, text prompt, and conditioning flags (`condition_frame_indexes_vision` or `condition_video_keep`) to regenerate specific frames while preserving others as visual anchors.**

The NVIDIA Cosmos repository implements a unified Mixture-of-Transformers (MoT) architecture capable of multimodal reasoning and generation. When operating in Generator mode, the vLLM-Omni server exposes this capability through an OpenAI-compatible HTTP API that accepts source video conditioning for precise frame-level editing. This approach allows you to transform video content by keeping selected frames unchanged while the diffusion process regenerates the remaining segments according to your textual instructions.

## Architecture Overview

Cosmos 3 employs a **single, unified Mixture-of-Transformers (MoT)** backbone that processes all modalities through shared transformer stacks. The architecture separates reasoning from generation through two distinct paths:

- **Causal Stack**: Processes text, image, and video tokens for multimodal reasoning
- **Diffusion Stack**: Denoises latent video-frame tokens to produce the final output

The **3-D Rotary Position Embedding (mRoPE)** encodes spatial (x, y) and temporal (t) axes, ensuring consistent motion across regenerated frames. When running in video-to-video mode, the model copies selected source frames into the latent sequence, masks the remaining positions, and runs the diffusion transformer over the full sequence while the prompt steers the new content.

## Conditioning Parameters for Video-to-Video

Video-to-video transformation relies on specific keys within the `extra_params` field to control which source frames serve as conditioning anchors. These parameters determine how the model blends existing visual content with generated output.

**`condition_frame_indexes_vision`**: A list of zero-based frame indices from the uploaded source video that remain **preserved unchanged** in the output. These frames become clean conditioning signals that anchor the diffusion process.

**`condition_video_keep`**: A boolean flag that, when set to `true`, instructs the model to preserve the **entire** source video as conditioning. When `false`, only the frames specified in `condition_frame_indexes_vision` are kept, allowing selective regeneration of specific temporal segments.

According to the [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) in the NVIDIA Cosmos repository, these parameters enable precise control over the generation boundary, letting you mix source content with new generation (e.g., keeping a robotic arm's initial motion while transforming the background).

## Implementation Methods

### Using cURL for Direct API Access

The simplest approach uses standard HTTP multipart requests to the vLLM-Omni server. Upload your source video via the `input_reference` field and specify the conditioning frames in `extra_params`:

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=Add a flowing stream of water behind the robot." \
  --form-string "size=1280x720" \
  --form-string "num_frames=189" \
  --form-string "fps=24" \
  --form-string "num_inference_steps=35" \
  --form-string "guidance_scale=6.0" \
  --form-string 'extra_params={"condition_frame_indexes_vision":[0,1,2],"condition_video_keep":false}' \
  -F "input_reference=@/path/to/source_video.mp4" \
  -o transformed_video.mp4

```

This command preserves the first three frames (`[0,1,2]`) and regenerates the remainder according to the prompt, saving the result as `transformed_video.mp4`.

### Python OpenAI Client Integration

For programmatic access, use the OpenAI Python client to interact with the vLLM-Omni server. First upload your video file, then reference it by ID in the generation request:

```python
import openai
import json

client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="skip")

with open("source_video.mp4", "rb") as f:
    response = client.files.create(
        file=f,
        purpose="video"
    )
    video_id = response.id

payload = {
    "model": "nvidia/Cosmos3-Nano",
    "prompt": "Turn the room into a sunny garden with birds flying.",
    "size": "1280x720",
    "num_frames": 189,
    "fps": 24,
    "num_inference_steps": 35,
    "guidance_scale": 6.0,
    "input_reference": video_id,
    "extra_params": json.dumps({
        "condition_frame_indexes_vision": [0, 1, 2],
        "condition_video_keep": False
    })
}

result = client.videos.sync(**payload)
with open("garden_transform.mp4", "wb") as out:
    out.write(result.content)

```

The `input_reference` field accepts either a file ID (as shown above) or a direct multipart upload. The `extra_params` JSON string controls which source frames are preserved during the diffusion process.

### Diffusers Pipeline Interface

For HuggingFace Diffusers users, the `Cosmos3OmniPipeline` abstracts the HTTP API while maintaining the same conditioning logic:

```python
from diffusers import Cosmos3OmniPipeline
import torch

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.float16,
    device="cuda"
)

output = pipe(
    prompt="Replace the background with a night sky.",
    input_reference="source_video.mp4",
    condition_frame_indexes_vision=[0, 1, 2],
    condition_video_keep=False,
    num_frames=189,
    fps=24,
    guidance_scale=6.0,
    num_inference_steps=35,
)

output.video.save("night_sky_transform.mp4")

```

This pipeline forwards the conditioning parameters to the underlying vLLM-Omni server automatically.

## Key Source Files and References

The following files in the NVIDIA Cosmos repository contain the concrete API specifications and implementation details:

- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)**: Documents the main vLLM-Omni integration, including the `extra_params` table for video-to-video conditioning (lines 90-92)
- **[`cookbooks/cosmos3/generator/audiovisual/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/README.md)**: Provides quick-start notebooks for audiovisual generation workflows
- **`cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb`**: Full notebook demonstrating server launch and generation calls
- **[`inference_benchmarks.md`](https://github.com/NVIDIA/cosmos/blob/main/inference_benchmarks.md)**: Performance benchmarks covering text-to-image, text-to-video, image-to-video, and video-to-video modalities

## Summary

- Video-to-video transformation in vLLM-Omni uses the standard `POST /v1/videos/sync` endpoint with an `input_reference` field for source video upload
- **Prompt-guided conditioning** combines textual instructions with visual anchors via `condition_frame_indexes_vision` (selective frames) or `condition_video_keep` (full video)
- The Cosmos 3 MoT architecture processes all modalities through a shared transformer, using the diffusion stack to denoise latent representations while preserving conditioning frames
- Implementation options include direct cURL requests, the OpenAI Python client, and the HuggingFace Diffusers pipeline
- Frame indices are zero-based; preserved frames act as clean conditioning signals that maintain temporal consistency with generated content

## Frequently Asked Questions

### What is the difference between `condition_frame_indexes_vision` and `condition_video_keep`?

The `condition_frame_indexes_vision` parameter accepts a list of specific frame indices (e.g., `[0, 1, 2]`) that should remain unchanged in the output, while the diffusion model regenerates all other frames. Setting `condition_video_keep: true` overrides this behavior by preserving the entire source video as conditioning, effectively treating the full input as a visual anchor rather than a partial one.

### Can I use the asynchronous endpoint for video-to-video transformation?

The analysis and documentation in the NVIDIA Cosmos repository specifically reference the synchronous `POST /v1/videos/sync` endpoint for video-to-video operations. While the vLLM-Omni server may support async patterns for other modalities, the video-to-video implementation with prompt-guided conditioning is designed and benchmarked around the synchronous API that returns raw MP4 bytes directly.

### How do I determine which frame indices to preserve for optimal results?

Select frames that establish the critical motion, subject position, or scene context you want to maintain. For robotic motion or character animation, preserving the initial 2-3 frames (`[0, 1, 2]`) typically provides sufficient conditioning for the model to maintain temporal consistency. The preserved frames should contain clear visual information that guides the diffusion process for the regenerated segments.

### What video formats does the `input_reference` parameter accept?

The vLLM-Omni server accepts standard video formats including MP4, as demonstrated in the code examples using `source_video.mp4`. When using the OpenAI client, you upload the file first to obtain a file ID, or pass a direct file path when using the Diffusers pipeline interface. The server processes the input into latent tokens compatible with the Cosmos 3 tokenizer before applying the conditioning logic.