# Implementing Video-to-Video Transformation with Prompt Guidance in Cosmos 3

> Learn how to implement video-to-video transformation with prompt guidance in Cosmos 3. This guide covers using textual prompts and spatial controls for powerful video generation.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Cosmos 3 enables video-to-video generation through a declarative JSON specification that combines textual prompts with spatial control signals like edge maps, depth, and segmentation, orchestrated via the `cosmos_framework.scripts.inference` CLI.**

The NVIDIA Cosmos repository provides a comprehensive framework for generative video AI, including the Cosmos 3 video-to-video pipeline. This system transforms source videos into target clips using text prompts alongside spatial control hints, requiring no model code modifications—only configuration through JSON spec files and the Cosmos Framework inference engine.

## Architecture of the Cosmos 3 Video-to-Video Pipeline

The Cosmos 3 video-to-video system operates through a separation between the inference engine and the declarative configuration. The **Cosmos Framework** (hosted externally and referenced via the `COSMOS_FRAMEWORK` environment variable) supplies the inference engine and model weights, while the repository contains the JSON specs, control assets, and helper utilities.

### Core Components

The pipeline relies on four primary components defined in the repository structure:

- **Spec JSON**: Declarative configuration files located at `cookbooks/cosmos3/generator/transfer/specs/*.json` (e.g., [`edge.json`](https://github.com/NVIDIA/cosmos/blob/main/edge.json), [`multi_control.json`](https://github.com/NVIDIA/cosmos/blob/main/multi_control.json)) that define the prompt paths, control hints, and generation hyper-parameters.
- **Prompt Files**: Textual conditioning stored in `assets/*/prompt.json` files (e.g., [`assets/edge/prompt.json`](https://github.com/NVIDIA/cosmos/blob/main/assets/edge/prompt.json)) containing positive captions, with optional negative prompts in [`assets/negative_prompt.json`](https://github.com/NVIDIA/cosmos/blob/main/assets/negative_prompt.json).
- **Control Videos**: Pre-computed spatial hints (edge maps, depth maps, segmentation) stored as `assets/*/control_*.mp4` files (e.g., `assets/edge/control_edge.mp4`), or derived on-the-fly from raw source videos via `vision_path`.
- **Inference Driver**: The `cosmos_framework.scripts.inference` CLI entry point that parses specifications, loads the Cosmos 3 model (Nano or Super variants), and executes generation.

### Prompt Guidance Flow

The model processes conditioning through a multi-stream architecture:

1. **Text Tokenization**: The prompt file specified in `prompt_path` is tokenized and embedded by the model's text encoder.
2. **Control Concatenation**: Control hint blocks (e.g., `"edge"`, `"depth"`, `"blur"`) supply video frames that are concatenated with latent conditioning at each timestep.
3. **Guidance Scaling**: The `control_guidance` parameter (default **1.5**) scales the influence of all active control streams, while per-hint **`weight`** values (default **1.0**) distribute influence among multiple controls.
4. **Classifier-Free Guidance**: The `guidance` parameter applies standard CFG scaling on top of the prompt and control conditioning to produce the final output.

## Configuring the Generation Spec

The JSON specification serves as the single source of truth for video-to-video transformation parameters.

### Single-Control Configuration

A minimal spec for edge-map guidance resides at [`cookbooks/cosmos3/generator/transfer/specs/edge.json`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/specs/edge.json):

```json
{
  "name": "transfer_edge",
  "model_mode": "video2video",
  "resolution": "720",
  "aspect_ratio": "16,9",
  "num_frames": 121,
  "fps": 30,
  "guidance": 3.0,
  "control_guidance": 1.5,
  "negative_prompt_file": "../assets/negative_prompt.json",
  "prompt_path": "../assets/edge/prompt.json",
  "edge": {
    "control_path": "../assets/edge/control_edge.mp4",
    "preset_edge_threshold": "medium"
  }
}

```

Key fields include:
- **`model_mode`**: Must be set to `"video2video"` to trigger the transformation pipeline.
- **`guidance`**: Classifier-free guidance scale for the textual prompt (typically 3.0).
- **`control_guidance`**: Global scale applied to all spatial control streams.
- **`prompt_path`**: Path to a JSON file containing the positive caption.
- **Control block**: Each hint type accepts a `control_path` (pre-computed video) or `vision_path` (raw source) plus hint-specific parameters like `preset_edge_threshold`.

### Multi-Control Configuration

For combining multiple spatial signals, the spec includes separate blocks with per-hint weights:

```json
{
  "name": "transfer_multi_control",
  "model_mode": "video2video",
  "resolution": "720",
  "aspect_ratio": "16,9",
  "num_frames": 121,
  "fps": 30,
  "guidance": 3.0,
  "control_guidance": 1.5,
  "negative_prompt_file": "../assets/negative_prompt.json",
  "prompt_path": "../assets/multi_control/prompt.json",
  "vision_path": "https://example.com/robot_pouring.mp4",
  "edge": {
    "weight": 0.75,
    "preset_edge_threshold": "medium"
  },
  "blur": {
    "weight": 0.25,
    "preset_blur_strength": "medium"
  }
}

```

When `vision_path` is provided instead of `control_path`, the framework computes the control signals on-the-fly using the specified presets.

## Running Video-to-Video Generation

Execution requires setting the `COSMOS_FRAMEWORK` environment variable to point to the external Cosmos Framework package, then invoking the inference CLI.

### CLI Execution

Set up the environment and run the transfer using the Cosmos 3 Nano checkpoint:

```bash
export COSMOS_FRAMEWORK=/path/to/cosmos-framework
export TRANSFER_ROOT=$(pwd)/cookbooks/cosmos3/generator/transfer

CUDA_VISIBLE_DEVICES=0 \
.venv/bin/python -m cosmos_framework.scripts.inference \
  --parallelism-preset=latency \
  -i "$TRANSFER_ROOT/specs/edge.json" \
  -o "$TRANSFER_ROOT/outputs/Cosmos3-Nano/" \
  --checkpoint-path Cosmos3-Nano \
  --seed 2026

```

Replace [`edge.json`](https://github.com/NVIDIA/cosmos/blob/main/edge.json) with [`blur.json`](https://github.com/NVIDIA/cosmos/blob/main/blur.json), [`depth.json`](https://github.com/NVIDIA/cosmos/blob/main/depth.json), [`seg.json`](https://github.com/NVIDIA/cosmos/blob/main/seg.json), [`wsm.json`](https://github.com/NVIDIA/cosmos/blob/main/wsm.json), or [`multi_control.json`](https://github.com/NVIDIA/cosmos/blob/main/multi_control.json) to utilize different control modalities. For multi-GPU execution with the Super checkpoint, adjust `--checkpoint-path` to `Cosmos3-Super` and modify the parallelism preset accordingly.

### Notebook Workflow

The repository includes a self-contained notebook at `cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb` that bundles environment setup, Hugging Face token handling, and execution. Run it headlessly with:

```bash
cd cookbooks/cosmos3/generator/transfer
jupyter execute run_video_transfer_with_cosmos_framework.ipynb

```

## Prompt Guidance and Control Parameters

Fine-tuning the relationship between textual and spatial conditioning involves three key parameters specified in the JSON spec:

- **`control_guidance`**: A global multiplier (default 1.5) affecting the strength of all control signals relative to the noise prediction.
- **`weight`**: Per-hint multipliers (default 1.0) that distribute the `control_guidance` budget among active controls in multi-hint scenarios.
- **`emphasize_control_in_prompt`**: A boolean flag that biases the model toward control-derived appearance when set to `true`.

The text encoder processes prompt files exactly as in standard text-to-video generation, meaning you can utilize complex prompts with weighting syntax while the control streams handle spatial consistency.

## Utility Functions for Validation

The repository provides [`cookbooks/cosmos3/generator/transfer/preview_helpers.py`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/preview_helpers.py) for input validation and output preview:

```python
from pathlib import Path
from cookbooks.cosmos3.generator.transfer.preview_helpers import preview_video

spec_path = Path("specs/edge.json")
output_video = Path("outputs/Cosmos3-Nano/transfer_edge/vision.mp4")

# Validates control files exist and displays a thumbnail

preview_video(output_video)

```

This helper raises descriptive errors if control videos or prompt files are missing before execution begins.

## Summary

- The Cosmos 3 video-to-video pipeline is driven by JSON specs in `cookbooks/cosmos3/generator/transfer/specs/` that configure prompts, control hints, and guidance parameters.
- **`model_mode`** must be set to `"video2video"` to enable transformation rather than text-to-video generation.
- **Control guidance** is managed through the global `control_guidance` scale (default 1.5) and per-hint `weight` values for multi-control scenarios.
- Execution occurs via `python -m cosmos_framework.scripts.inference` with the `COSMOS_FRAMEWORK` environment variable pointing to the external framework package.
- Pre-computed control videos can be replaced with raw `vision_path` URLs to generate controls on-the-fly using preset thresholds.

## Frequently Asked Questions

### What is the difference between `guidance` and `control_guidance` in Cosmos 3?

The **`guidance`** parameter implements standard classifier-free guidance for the textual prompt, typically set to 3.0, while **`control_guidance`** specifically scales the influence of spatial control signals like edge or depth maps. The latter defaults to 1.5 and applies globally to all active control streams configured in the spec.

### Can I run video-to-video transformation without pre-computed control videos?

Yes. By specifying a **`vision_path`** in your spec JSON instead of `control_path`, the Cosmos Framework automatically derives control signals on-the-fly from the source video. You must include the appropriate preset parameters (e.g., `preset_edge_threshold` or `preset_blur_strength`) to configure how these controls are computed.

### How do I combine multiple control signals like edge and depth?

Create a multi-control spec JSON that includes multiple hint blocks (e.g., `"edge"` and `"blur"`) with individual **`weight`** values summing to 1.0. The framework blends these signals according to their weights while applying the global `control_guidance` scale to the combined conditioning.

### Which model checkpoints support video-to-video generation?

The Cosmos 3 pipeline supports both **Cosmos3-Nano** for single-GPU inference and **Cosmos3-Super** for high-quality multi-GPU generation. Specify your chosen checkpoint using the `--checkpoint-path` argument when invoking the inference script.