How to Use Video-to-Video Transformation with Prompt-Guided Conditioning in vLLM-Omni
Submit a POST request to the /v1/videos/sync endpoint with an input video, text prompt, and conditioning flags (condition_frame_indexes_vision or condition_video_keep) to regenerate specific frames while preserving others as visual anchors.
The NVIDIA Cosmos repository implements a unified Mixture-of-Transformers (MoT) architecture capable of multimodal reasoning and generation. When operating in Generator mode, the vLLM-Omni server exposes this capability through an OpenAI-compatible HTTP API that accepts source video conditioning for precise frame-level editing. This approach allows you to transform video content by keeping selected frames unchanged while the diffusion process regenerates the remaining segments according to your textual instructions.
Architecture Overview
Cosmos 3 employs a single, unified Mixture-of-Transformers (MoT) backbone that processes all modalities through shared transformer stacks. The architecture separates reasoning from generation through two distinct paths:
- Causal Stack: Processes text, image, and video tokens for multimodal reasoning
- Diffusion Stack: Denoises latent video-frame tokens to produce the final output
The 3-D Rotary Position Embedding (mRoPE) encodes spatial (x, y) and temporal (t) axes, ensuring consistent motion across regenerated frames. When running in video-to-video mode, the model copies selected source frames into the latent sequence, masks the remaining positions, and runs the diffusion transformer over the full sequence while the prompt steers the new content.
Conditioning Parameters for Video-to-Video
Video-to-video transformation relies on specific keys within the extra_params field to control which source frames serve as conditioning anchors. These parameters determine how the model blends existing visual content with generated output.
condition_frame_indexes_vision: A list of zero-based frame indices from the uploaded source video that remain preserved unchanged in the output. These frames become clean conditioning signals that anchor the diffusion process.
condition_video_keep: A boolean flag that, when set to true, instructs the model to preserve the entire source video as conditioning. When false, only the frames specified in condition_frame_indexes_vision are kept, allowing selective regeneration of specific temporal segments.
According to the README.md in the NVIDIA Cosmos repository, these parameters enable precise control over the generation boundary, letting you mix source content with new generation (e.g., keeping a robotic arm's initial motion while transforming the background).
Implementation Methods
Using cURL for Direct API Access
The simplest approach uses standard HTTP multipart requests to the vLLM-Omni server. Upload your source video via the input_reference field and specify the conditioning frames in extra_params:
curl -sS -X POST http://localhost:8000/v1/videos/sync \
--form-string "prompt=Add a flowing stream of water behind the robot." \
--form-string "size=1280x720" \
--form-string "num_frames=189" \
--form-string "fps=24" \
--form-string "num_inference_steps=35" \
--form-string "guidance_scale=6.0" \
--form-string 'extra_params={"condition_frame_indexes_vision":[0,1,2],"condition_video_keep":false}' \
-F "input_reference=@/path/to/source_video.mp4" \
-o transformed_video.mp4
This command preserves the first three frames ([0,1,2]) and regenerates the remainder according to the prompt, saving the result as transformed_video.mp4.
Python OpenAI Client Integration
For programmatic access, use the OpenAI Python client to interact with the vLLM-Omni server. First upload your video file, then reference it by ID in the generation request:
import openai
import json
client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="skip")
with open("source_video.mp4", "rb") as f:
response = client.files.create(
file=f,
purpose="video"
)
video_id = response.id
payload = {
"model": "nvidia/Cosmos3-Nano",
"prompt": "Turn the room into a sunny garden with birds flying.",
"size": "1280x720",
"num_frames": 189,
"fps": 24,
"num_inference_steps": 35,
"guidance_scale": 6.0,
"input_reference": video_id,
"extra_params": json.dumps({
"condition_frame_indexes_vision": [0, 1, 2],
"condition_video_keep": False
})
}
result = client.videos.sync(**payload)
with open("garden_transform.mp4", "wb") as out:
out.write(result.content)
The input_reference field accepts either a file ID (as shown above) or a direct multipart upload. The extra_params JSON string controls which source frames are preserved during the diffusion process.
Diffusers Pipeline Interface
For HuggingFace Diffusers users, the Cosmos3OmniPipeline abstracts the HTTP API while maintaining the same conditioning logic:
from diffusers import Cosmos3OmniPipeline
import torch
pipe = Cosmos3OmniPipeline.from_pretrained(
"nvidia/Cosmos3-Nano",
torch_dtype=torch.float16,
device="cuda"
)
output = pipe(
prompt="Replace the background with a night sky.",
input_reference="source_video.mp4",
condition_frame_indexes_vision=[0, 1, 2],
condition_video_keep=False,
num_frames=189,
fps=24,
guidance_scale=6.0,
num_inference_steps=35,
)
output.video.save("night_sky_transform.mp4")
This pipeline forwards the conditioning parameters to the underlying vLLM-Omni server automatically.
Key Source Files and References
The following files in the NVIDIA Cosmos repository contain the concrete API specifications and implementation details:
README.md: Documents the main vLLM-Omni integration, including theextra_paramstable for video-to-video conditioning (lines 90-92)cookbooks/cosmos3/generator/audiovisual/README.md: Provides quick-start notebooks for audiovisual generation workflowscookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb: Full notebook demonstrating server launch and generation callsinference_benchmarks.md: Performance benchmarks covering text-to-image, text-to-video, image-to-video, and video-to-video modalities
Summary
- Video-to-video transformation in vLLM-Omni uses the standard
POST /v1/videos/syncendpoint with aninput_referencefield for source video upload - Prompt-guided conditioning combines textual instructions with visual anchors via
condition_frame_indexes_vision(selective frames) orcondition_video_keep(full video) - The Cosmos 3 MoT architecture processes all modalities through a shared transformer, using the diffusion stack to denoise latent representations while preserving conditioning frames
- Implementation options include direct cURL requests, the OpenAI Python client, and the HuggingFace Diffusers pipeline
- Frame indices are zero-based; preserved frames act as clean conditioning signals that maintain temporal consistency with generated content
Frequently Asked Questions
What is the difference between condition_frame_indexes_vision and condition_video_keep?
The condition_frame_indexes_vision parameter accepts a list of specific frame indices (e.g., [0, 1, 2]) that should remain unchanged in the output, while the diffusion model regenerates all other frames. Setting condition_video_keep: true overrides this behavior by preserving the entire source video as conditioning, effectively treating the full input as a visual anchor rather than a partial one.
Can I use the asynchronous endpoint for video-to-video transformation?
The analysis and documentation in the NVIDIA Cosmos repository specifically reference the synchronous POST /v1/videos/sync endpoint for video-to-video operations. While the vLLM-Omni server may support async patterns for other modalities, the video-to-video implementation with prompt-guided conditioning is designed and benchmarked around the synchronous API that returns raw MP4 bytes directly.
How do I determine which frame indices to preserve for optimal results?
Select frames that establish the critical motion, subject position, or scene context you want to maintain. For robotic motion or character animation, preserving the initial 2-3 frames ([0, 1, 2]) typically provides sufficient conditioning for the model to maintain temporal consistency. The preserved frames should contain clear visual information that guides the diffusion process for the regenerated segments.
What video formats does the input_reference parameter accept?
The vLLM-Omni server accepts standard video formats including MP4, as demonstrated in the code examples using source_video.mp4. When using the OpenAI client, you upload the file first to obtain a file ID, or pass a direct file path when using the Diffusers pipeline interface. The server processes the input into latent tokens compatible with the Cosmos 3 tokenizer before applying the conditioning logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →