How to Implement Video Transfer with Edge Detection, Blur, Depth Estimation, and Segmentation Controls in NVIDIA Cosmos
NVIDIA Cosmos 3 enables video-to-video generation conditioned on spatial control maps—edge, blur, depth, segmentation, and world-scenario—by running inference specs through the Cosmos Framework with the cosmos_framework.scripts.inference entry point.
This guide walks through the open-source implementation of video transfer (also called video-to-video) in the NVIDIA Cosmos repository. You will learn how to generate temporally coherent videos guided by auxiliary control maps, using JSON specifications and helper utilities located in cookbooks/cosmos3/generator/transfer/.
Understanding Video Transfer Controls
Cosmos 3 supports five distinct spatial-map modalities that act as conditioning signals for the diffusion pipeline. Each control type directs the model to preserve specific structural or semantic properties from a source video while generating new content.
The available controls are:
- Edge – Canny edge maps extracted from source frames, preserving structural boundaries.
- Blur – Gaussian-blurred reference frames that guide composition without sharp detail constraints.
- Depth – Monocular depth estimates providing per-pixel distance information for 3D-aware generation.
- Seg – Semantic segmentation maps defining object classes and scene layout.
- WSM – World-scenario maps for custom spatial layouts and environmental control.
Each modality corresponds to a specific JSON configuration file in cookbooks/cosmos3/generator/transfer/specs/, such as edge.json for Canny edge guidance or depth.json for depth-map conditioned transfer.
Prerequisites and Environment Setup
The workflow requires the Cosmos Framework, a separate repository that provides the inference engine, along with standard media processing tools.
Install the framework using uv for isolated Python management:
# Create managed Python 3.13 environment
uv venv --python 3.13 --seed --managed-python
source .venv/bin/activate
# Install Cosmos Framework
uv pip install "cosmos-framework @ git+https://github.com/NVIDIA/cosmos-framework.git"
# Install ffmpeg helper (optional; ensures codec availability)
uv pip install imageio-ffmpeg
The framework provides the cosmos_framework.scripts.inference module, which parses specifications and initializes the Cosmos3OmniPipeline with the UniPCMultistepScheduler.
Configuring Control Specifications
Control parameters reside in JSON specification files that define the diffusion pipeline configuration. Each spec sets model_mode: "video2video" alongside frame count, temporal sampling, and control-specific guidance parameters.
Key specification files include:
cookbooks/cosmos3/generator/transfer/specs/edge.json– Configures Canny edge thresholds and edge-preservation strength.cookbooks/cosmos3/generator/transfer/specs/blur.json– Defines Gaussian kernel parameters for blur-based guidance.cookbooks/cosmos3/generator/transfer/specs/depth.json– Specifies depth-map normalization and conditioning scale.cookbooks/cosmos3/generator/transfer/specs/seg.json– Sets segmentation label mapping and semantic guidance weights.cookbooks/cosmos3/generator/transfer/specs/wsm.json– Configures world-scenario map interpretation.
Each JSON file contains the path to its corresponding control video (e.g., control_edge.mp4, control_depth.mp4) and generation parameters such as control_guidance, resolution, and diffusion steps.
Running Video Transfer Inference
Execute inference by pointing the framework to your chosen specification file. The script loads the default checkpoint (nvidia/Cosmos3-Nano) and processes the control video as an additional conditioning token stream via pipeline.control_video.
Run edge-conditioned transfer:
python -m cosmos_framework.scripts.inference \
--spec $(pwd)/cookbooks/cosmos3/generator/transfer/specs/edge.json
For depth estimation control:
python -m cosmos_framework.scripts.inference \
--spec $(pwd)/cookbooks/cosmos3/generator/transfer/specs/depth.json
The framework handles checkpoint loading, scheduler initialization, and the video-to-video diffusion process, outputting a generated vision.mp4 file.
Previewing Results with Helper Utilities
After inference completes, inspect results using the preview_helpers.py module. This utility generates low-bit-rate H.264 previews (CRF 28) suitable for notebook embedding or CLI inspection without re-encoding full-resolution outputs.
Preview the edge control result:
from cookbooks.cosmos3.generator.transfer.preview_helpers import preview_transfer
preview_transfer('edge')
The preview_transfer() function locates the specification root, retrieves the corresponding control video and generated output, and constructs a side-by-side comparison using ffmpeg. For interactive development, the complete notebook at cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb automates environment setup, inference execution, and inline visualization.
Summary
- Video transfer in Cosmos 3 uses spatial control maps (edge, blur, depth, segmentation, world-scenario) to condition video-to-video diffusion.
- Specification files in
cookbooks/cosmos3/generator/transfer/specs/define parameters for each control modality, including guidance strength and input video paths. - Inference runs via
python -m cosmos_framework.scripts.inference --spec <path>, utilizing theCosmos3OmniPipelinewithUniPCMultistepScheduler. - Preview generation uses
preview_helpers.preview_transfer()to create lightweight H.264 comparisons of control inputs and generated outputs. - Scaling to larger checkpoints like
Cosmos3-Superrequires only changing the model identifier in the inference command.
Frequently Asked Questions
What is the difference between video transfer and standard video generation in Cosmos 3?
Video transfer (video-to-video) conditions generation on both a source video and a control map, preserving structural or semantic elements from the input. Standard video generation typically uses text prompts or image inputs without temporal consistency constraints from a reference video sequence.
How do I switch between different control modalities like edge detection and depth estimation?
Change the --spec argument to point to the desired JSON configuration file. For edge detection use edge.json, for depth estimation use depth.json, and similarly for blur, segmentation, or world-scenario controls. Each spec automatically configures the pipeline to process the corresponding control video format.
Can I use custom control videos instead of the provided examples?
Yes. Modify the control video path inside the JSON specification file (e.g., control_edge.mp4 in edge.json) to reference your custom video. Ensure the video dimensions and frame count match the specification parameters, and that the visual content corresponds to the expected modality (Canny edges for edge control, depth maps for depth control, etc.).
Where is the inference code that processes these control videos?
The core inference logic resides in the Cosmos Framework repository, specifically within cosmos_framework.scripts.inference. This script builds the Cosmos3OmniPipeline, loads the checkpoint, and passes control videos to pipeline.control_video as additional conditioning tokens. The NVIDIA Cosmos repository provides the cookbook specifications and preview helpers that interface with this framework.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →