How Portrait Source Detection and Scaling Works in video‑use

video‑use automatically detects portrait orientation by probing source dimensions with ffprobe and applies conditional FFmpeg scaling filters that fix either height or width to preserve aspect ratio during extraction.

The browser-use/video-use repository processes video clips from diverse sources, including vertical smartphone footage and traditional landscape media. Understanding how portrait source detection and scaling works ensures the pipeline automatically adapts to any input aspect ratio while maintaining consistent output resolution and file sizes across draft and final renders.

Detecting Portrait Orientation with ffprobe

Portrait detection relies on the is_portrait_source helper located in helpers/render.py (lines 134–145). This function executes an ffprobe subprocess to query the video stream’s intrinsic dimensions and compares the height against the width.


# helpers/render.py

def is_portrait_source(video: Path) -> bool:
    """Return True if the video's height > width (portrait / vertical)."""
    try:
        out = subprocess.run(
            ["ffprobe", "-v", "error", "-select_streams", "v:0",
             "-show_entries", "stream=width,height",
             "-of", "csv=p=0", str(video)],
            capture_output=True, text=True, check=True,
        )
        w, h = map(int, out.stdout.strip().split(","))
        return h > w
    except Exception:
        return False

The function parses the comma-separated output from ffprobe’s csv=p=0 format, converts the values to integers, and returns True only when h > w. If ffprobe fails or returns malformed data, the function safely defaults to False, treating the source as landscape to avoid incorrect scaling assumptions.

Adaptive Scaling Logic in extract_segment

Inside extract_segment (lines 73–78 of helpers/render.py), the code branches based on the boolean returned by is_portrait_source. The goal is to ensure the longer side of the output matches the target resolution (1280 px for draft mode, 1920 px for final mode) while letting FFmpeg calculate the complementary dimension to preserve the original aspect ratio.

  • Portrait sources receive a filter fixing the height: scale=-2:1280 (draft) or scale=-2:1920 (final).
  • Landscape sources receive a filter fixing the width: scale=1280:-2 (draft) or scale=1920:-2 (final).

# helpers/render.py (excerpt)

portrait = is_portrait_source(source)
if draft:
    scale = "scale=-2:1280" if portrait else "scale=1280:-2"
else:
    scale = "scale=-2:1920" if portrait else "scale=1920:-2"

The -2 value is a FFmpeg convention that instructs the scaler to calculate the respective dimension automatically while rounding to the nearest multiple of 2, ensuring codec compatibility without manual arithmetic.

Assembling the FFmpeg Filter Chain

The selected scale string is integrated into a dynamic filter chain built within extract_segment. The pipeline first checks for HDR content using is_hdr_source. If present, tone-mapping filters are prepended to vf_parts. The scaling filter is then appended, followed by any optional color-grading filters.

vf_parts: list[str] = []
if is_hdr_source(source):
    vf_parts.append(TONEMAP_CHAIN)
vf_parts.append(scale)          # ← portrait-aware scaling

if grade_filter:
    vf_parts.append(grade_filter)
vf = ",".join(vf_parts)

The final comma-separated vf string is passed to FFmpeg’s -vf argument, ensuring the video is processed in the correct order: tone mapping, resolution scaling, and color grading.

Practical Code Examples

Extract a Portrait Clip in Draft Mode

This example processes a vertical video where height > width. Because draft=True, the function selects scale=-2:1280, fixing the height to 1280 pixels.

from pathlib import Path
from helpers.render import extract_segment

src = Path("my-portrait-clip.mp4")   # height > width

start = 12.0                         # seconds

duration = 4.5
grade = ""                           # no colour grading

out = Path("tmp/seg_00.mp4")

# Draft = True → the code will pick scale=-2:1280 (height-fixed)

extract_segment(src, start, duration, grade, out, preview=False, draft=True)

Extract a Landscape Clip in Final Mode

For a traditional landscape source, extract_segment automatically uses scale=1920:-2, locking the width to 1920 pixels and calculating the height.

src = Path("my-landscape-clip.mp4")   # width > height

extract_segment(src, 30.0, 5.0, "", Path("tmp/seg_01.mp4"))

# The function picks scale=1920:-2 (width-fixed)

Batch Processing with the High-Level API

Use extract_all_segments to process an edit decision list (EDL) containing multiple sources of mixed orientations. The function invokes the detection and scaling logic automatically for each entry.

from helpers.render import extract_all_segments

edl = {
    "ranges": [{"source": "A", "start": 0, "end": 10}],
    "sources": {"A": "my-portrait-clip.mp4"},
}
edit_dir = Path("edit")
segments = extract_all_segments(edl, edit_dir, preview=False, draft=False)

# `segments` now contains the re-scaled MP4 ready for further processing.

Summary

  • Portrait detection is performed by is_portrait_source in helpers/render.py, which uses ffprobe to compare height and width values (lines 134–145).
  • Adaptive scaling switches between scale=-2:<height> for portrait and scale=<width>:-2 for landscape inside extract_segment (lines 73–78).
  • Resolution targets differ by mode: draft uses 1280 px while final uses 1920 px on the dominant axis.
  • Filter chain assembly dynamically prepends HDR tone mapping and appends color grading, with the scaling filter positioned between them.
  • The architecture guarantees consistent output dimensions regardless of whether the source is vertical or horizontal.

Frequently Asked Questions

How does video-use determine if a video is portrait or landscape?

The is_portrait_source function in helpers/render.py executes ffprobe with -show_entries stream=width,height and parses the CSV output. It returns True when the height integer exceeds the width integer, indicating a portrait or vertical video.

What does the -2 value mean in the FFmpeg scale filter?

The -2 placeholder instructs FFmpeg to automatically calculate the respective dimension (width or height) while ensuring the result is divisible by 2. This preserves the original aspect ratio without manual math and maintains codec compatibility.

Why does the scaling differ between draft and final modes?

Draft mode targets 1280 px on the dominant axis for faster processing and reduced file sizes suitable for previews. Final mode targets 1920 px for high-resolution delivery. The logic applies these values to the fixed side of the scale filter based on the detected orientation.

Can portrait detection handle HDR sources?

Yes, portrait detection and HDR handling operate independently. The extract_segment function checks for HDR via is_hdr_source and prepends tone-mapping filters to the filter chain before appending the orientation-specific scaling filter.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →