Implementing Video-to-Video Transformation with Prompt Guidance in Cosmos 3

Cosmos 3 enables video-to-video generation through a declarative JSON specification that combines textual prompts with spatial control signals like edge maps, depth, and segmentation, orchestrated via the cosmos_framework.scripts.inference CLI.

The NVIDIA Cosmos repository provides a comprehensive framework for generative video AI, including the Cosmos 3 video-to-video pipeline. This system transforms source videos into target clips using text prompts alongside spatial control hints, requiring no model code modifications—only configuration through JSON spec files and the Cosmos Framework inference engine.

Architecture of the Cosmos 3 Video-to-Video Pipeline

The Cosmos 3 video-to-video system operates through a separation between the inference engine and the declarative configuration. The Cosmos Framework (hosted externally and referenced via the COSMOS_FRAMEWORK environment variable) supplies the inference engine and model weights, while the repository contains the JSON specs, control assets, and helper utilities.

Core Components

The pipeline relies on four primary components defined in the repository structure:

  • Spec JSON: Declarative configuration files located at cookbooks/cosmos3/generator/transfer/specs/*.json (e.g., edge.json, multi_control.json) that define the prompt paths, control hints, and generation hyper-parameters.
  • Prompt Files: Textual conditioning stored in assets/*/prompt.json files (e.g., assets/edge/prompt.json) containing positive captions, with optional negative prompts in assets/negative_prompt.json.
  • Control Videos: Pre-computed spatial hints (edge maps, depth maps, segmentation) stored as assets/*/control_*.mp4 files (e.g., assets/edge/control_edge.mp4), or derived on-the-fly from raw source videos via vision_path.
  • Inference Driver: The cosmos_framework.scripts.inference CLI entry point that parses specifications, loads the Cosmos 3 model (Nano or Super variants), and executes generation.

Prompt Guidance Flow

The model processes conditioning through a multi-stream architecture:

  1. Text Tokenization: The prompt file specified in prompt_path is tokenized and embedded by the model's text encoder.
  2. Control Concatenation: Control hint blocks (e.g., "edge", "depth", "blur") supply video frames that are concatenated with latent conditioning at each timestep.
  3. Guidance Scaling: The control_guidance parameter (default 1.5) scales the influence of all active control streams, while per-hint weight values (default 1.0) distribute influence among multiple controls.
  4. Classifier-Free Guidance: The guidance parameter applies standard CFG scaling on top of the prompt and control conditioning to produce the final output.

Configuring the Generation Spec

The JSON specification serves as the single source of truth for video-to-video transformation parameters.

Single-Control Configuration

A minimal spec for edge-map guidance resides at cookbooks/cosmos3/generator/transfer/specs/edge.json:

{
  "name": "transfer_edge",
  "model_mode": "video2video",
  "resolution": "720",
  "aspect_ratio": "16,9",
  "num_frames": 121,
  "fps": 30,
  "guidance": 3.0,
  "control_guidance": 1.5,
  "negative_prompt_file": "../assets/negative_prompt.json",
  "prompt_path": "../assets/edge/prompt.json",
  "edge": {
    "control_path": "../assets/edge/control_edge.mp4",
    "preset_edge_threshold": "medium"
  }
}

Key fields include:

  • model_mode: Must be set to "video2video" to trigger the transformation pipeline.
  • guidance: Classifier-free guidance scale for the textual prompt (typically 3.0).
  • control_guidance: Global scale applied to all spatial control streams.
  • prompt_path: Path to a JSON file containing the positive caption.
  • Control block: Each hint type accepts a control_path (pre-computed video) or vision_path (raw source) plus hint-specific parameters like preset_edge_threshold.

Multi-Control Configuration

For combining multiple spatial signals, the spec includes separate blocks with per-hint weights:

{
  "name": "transfer_multi_control",
  "model_mode": "video2video",
  "resolution": "720",
  "aspect_ratio": "16,9",
  "num_frames": 121,
  "fps": 30,
  "guidance": 3.0,
  "control_guidance": 1.5,
  "negative_prompt_file": "../assets/negative_prompt.json",
  "prompt_path": "../assets/multi_control/prompt.json",
  "vision_path": "https://example.com/robot_pouring.mp4",
  "edge": {
    "weight": 0.75,
    "preset_edge_threshold": "medium"
  },
  "blur": {
    "weight": 0.25,
    "preset_blur_strength": "medium"
  }
}

When vision_path is provided instead of control_path, the framework computes the control signals on-the-fly using the specified presets.

Running Video-to-Video Generation

Execution requires setting the COSMOS_FRAMEWORK environment variable to point to the external Cosmos Framework package, then invoking the inference CLI.

CLI Execution

Set up the environment and run the transfer using the Cosmos 3 Nano checkpoint:

export COSMOS_FRAMEWORK=/path/to/cosmos-framework
export TRANSFER_ROOT=$(pwd)/cookbooks/cosmos3/generator/transfer

CUDA_VISIBLE_DEVICES=0 \
.venv/bin/python -m cosmos_framework.scripts.inference \
  --parallelism-preset=latency \
  -i "$TRANSFER_ROOT/specs/edge.json" \
  -o "$TRANSFER_ROOT/outputs/Cosmos3-Nano/" \
  --checkpoint-path Cosmos3-Nano \
  --seed 2026

Replace edge.json with blur.json, depth.json, seg.json, wsm.json, or multi_control.json to utilize different control modalities. For multi-GPU execution with the Super checkpoint, adjust --checkpoint-path to Cosmos3-Super and modify the parallelism preset accordingly.

Notebook Workflow

The repository includes a self-contained notebook at cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb that bundles environment setup, Hugging Face token handling, and execution. Run it headlessly with:

cd cookbooks/cosmos3/generator/transfer
jupyter execute run_video_transfer_with_cosmos_framework.ipynb

Prompt Guidance and Control Parameters

Fine-tuning the relationship between textual and spatial conditioning involves three key parameters specified in the JSON spec:

  • control_guidance: A global multiplier (default 1.5) affecting the strength of all control signals relative to the noise prediction.
  • weight: Per-hint multipliers (default 1.0) that distribute the control_guidance budget among active controls in multi-hint scenarios.
  • emphasize_control_in_prompt: A boolean flag that biases the model toward control-derived appearance when set to true.

The text encoder processes prompt files exactly as in standard text-to-video generation, meaning you can utilize complex prompts with weighting syntax while the control streams handle spatial consistency.

Utility Functions for Validation

The repository provides cookbooks/cosmos3/generator/transfer/preview_helpers.py for input validation and output preview:

from pathlib import Path
from cookbooks.cosmos3.generator.transfer.preview_helpers import preview_video

spec_path = Path("specs/edge.json")
output_video = Path("outputs/Cosmos3-Nano/transfer_edge/vision.mp4")

# Validates control files exist and displays a thumbnail

preview_video(output_video)

This helper raises descriptive errors if control videos or prompt files are missing before execution begins.

Summary

  • The Cosmos 3 video-to-video pipeline is driven by JSON specs in cookbooks/cosmos3/generator/transfer/specs/ that configure prompts, control hints, and guidance parameters.
  • model_mode must be set to "video2video" to enable transformation rather than text-to-video generation.
  • Control guidance is managed through the global control_guidance scale (default 1.5) and per-hint weight values for multi-control scenarios.
  • Execution occurs via python -m cosmos_framework.scripts.inference with the COSMOS_FRAMEWORK environment variable pointing to the external framework package.
  • Pre-computed control videos can be replaced with raw vision_path URLs to generate controls on-the-fly using preset thresholds.

Frequently Asked Questions

What is the difference between guidance and control_guidance in Cosmos 3?

The guidance parameter implements standard classifier-free guidance for the textual prompt, typically set to 3.0, while control_guidance specifically scales the influence of spatial control signals like edge or depth maps. The latter defaults to 1.5 and applies globally to all active control streams configured in the spec.

Can I run video-to-video transformation without pre-computed control videos?

Yes. By specifying a vision_path in your spec JSON instead of control_path, the Cosmos Framework automatically derives control signals on-the-fly from the source video. You must include the appropriate preset parameters (e.g., preset_edge_threshold or preset_blur_strength) to configure how these controls are computed.

How do I combine multiple control signals like edge and depth?

Create a multi-control spec JSON that includes multiple hint blocks (e.g., "edge" and "blur") with individual weight values summing to 1.0. The framework blends these signals according to their weights while applying the global control_guidance scale to the combined conditioning.

Which model checkpoints support video-to-video generation?

The Cosmos 3 pipeline supports both Cosmos3-Nano for single-GPU inference and Cosmos3-Super for high-quality multi-GPU generation. Specify your chosen checkpoint using the --checkpoint-path argument when invoking the inference script.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →