# How the video-use Render Pipeline Works: Extract → Concat → Composite → Normalize

> Understand the video-use render pipeline: extract, concat, composite, and normalize. Learn how this process generates final videos with color grading, overlays, and loudness normalization.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: internals
- Published: 2026-07-07

---

**TLDR:** The video-use render pipeline processes an Edit Decision List (EDL) through four sequential stages—**Extract**, **Concat**, **Composite**, and **Normalize**—to generate final videos with color grading, overlays, and broadcast-compliant loudness normalization.

The `video-use` tool from the `browser-use/video-use` repository implements this deterministic pipeline in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py). Each stage handles specific aspects of video production, from HDR tone-mapping to social-media loudness standards, ensuring that source footage transforms into platform-ready content through a predictable sequence of FFmpeg operations.

## Stage 1: Extract

The **Extract** stage creates self-contained clip files for every range defined in the EDL. This is the most computationally intensive phase, handling per-segment color grading, resolution scaling, and audio processing.

### Segment Processing and Color Grading

In [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py), the `extract_all_segments()` function (lines 14-62) iterates over `edl["ranges"]` and delegates to `extract_segment()` for individual clip generation. Each segment is processed with:

- **Resolution scaling** to 1080p (or 720p in draft mode)
- **Color grading** applied via `resolve_grade_filter` presets
- **HDR tone-mapping** to SDR if the source material contains HDR metadata
- **Audio fades** with 30ms padding at clip boundaries

The resulting FFmpeg command chains video filters (`-vf`) for tone-mapping, scaling, and grading, alongside audio filters (`-af`) for the fade-in/fade-out curves.

### Output Files

This stage produces discrete segment files named `seg_00_… .mp4`, `seg_01_… .mp4`, etc., which serve as lossless inputs for the next phase.

## Stage 2: Concat

The **Concat** stage assembles all extracted segments into a single base video without re-encoding, preserving quality and minimizing processing time.

### Lossless Assembly with FFmpeg

The `concat_segments()` function (lines 67-83) in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) generates a temporary [`_concat.txt`](https://github.com/browser-use/video-use/blob/main/_concat.txt) file listing all segment paths in sequence. It then invokes FFmpeg with the concat demuxer:

```bash
ffmpeg -f concat -c copy ...

```

This produces `base.mp4` (or `base_preview.mp4`/`base_draft.mp4` depending on mode) by copying codec streams directly—no generational loss occurs during this assembly.

## Stage 3: Composite

The **Composite** stage layers graphics, animations, and subtitles over the base video. If no overlays exist, the base file passes through unchanged.

### Overlay Positioning and the "Subtitles Last" Rule

Implemented in `build_final_composite()` within [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py), this function:

1. Adds each overlay as an additional input (`-i <overlay>`)
2. Offsets overlay PTS (Presentation Timestamp) to match the overlay's `start_in_output` timing
3. Chains overlays using `overlay=` filters
4. **Appends subtitles last** via the `subtitles=` filter, ensuring text renders above all visual elements

The function outputs to a temporary `*.prenorm.mp4` file when loudness normalization is enabled, or directly to the final destination if normalization is skipped.

## Stage 4: Normalize

The **Normalize** stage ensures audio meets social-media loudness standards using a two-pass loudness normalization process.

### Two-Pass Loudness Measurement

The `apply_loudnorm_two_pass()` function (lines 31-90) in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) implements the EBU R128 standard via FFmpeg's `loudnorm` filter:

- **Target:** -14 LUFS integrated loudness, -1 dBTP true peak, LRA 11
- **Pass 1:** `measure_loudness()` runs `ffmpeg … -af loudnorm=…:print_format=json` to analyze the audio stream and extract measured values (`measured_I`, `measured_TP`, etc.)
- **Pass 2:** Applies the `loudnorm` filter with measured parameters to produce `final.mp4`

### Draft and Preview Modes

In preview or draft mode, the pipeline skips the measurement pass and uses a faster single-pass approximation, trading precision for render speed.

## Orchestration and Entry Points

The `main()` function in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) orchestrates the video-use render pipeline execution:

1. Calls `extract_all_segments()` to generate per-range clips
2. Calls `concat_segments()` to build the base video
3. Optionally builds a master SRT via `build_master_srt()` if subtitles are required
4. Calls `build_final_composite()` to apply overlays and subtitles
5. Runs `apply_loudnorm_two_pass()` unless `--no-loudnorm` is specified

This deterministic flow ensures that every render follows the **extract → concat → composite → normalize** sequence, with each stage feeding artifact files into the next.

## Usage Examples

Control the pipeline behavior via command-line flags:

```bash

# Full quality render with subtitles and loudness normalization

python helpers/render.py edl.json -o final.mp4

# Preview render (faster, lower CRF)

python helpers/render.py edl.json -o preview.mp4 --preview

# Draft render (720p, ultra-fast, skip normalization)

python helpers/render.py edl.json -o draft.mp4 --draft --no-loudnorm

# Generate subtitles only without video processing

python helpers/render.py edl.json -o dummy.mp4 --build-subtitles

```

## Summary

- **Extract** cuts, grades, and scales individual EDL ranges into self-contained segments with 30ms audio fades and HDR tone-mapping.
- **Concat** losslessly joins segments using FFmpeg's concat demuxer to create a base video.
- **Composite** layers overlays and subtitles (always last) over the base video.
- **Normalize** applies two-pass EBU R128 loudness normalization (-14 LUFS) unless skipped via `--no-loudnorm`.

## Frequently Asked Questions

### What file formats does the video-use render pipeline support?

The pipeline uses FFmpeg as its processing engine, inheriting support for most container formats (MP4, MOV, MKV) and codecs (H.264, H.265, ProRes). The specific codec choice depends on your EDL configuration and whether draft mode (720p) or full quality (1080p) is selected.

### Can I skip the loudness normalization stage?

Yes. Pass the `--no-loudnorm` flag when running [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) to disable the normalization stage. In this case, `build_final_composite()` outputs directly to your destination file, and the audio maintains its original levels without the two-pass EBU R128 processing.

### How does the pipeline handle HDR source footage?

During the **Extract** stage, `extract_segment()` detects HDR metadata and applies tone-mapping filters via FFmpeg to convert HDR content to SDR. This ensures compatibility with standard delivery platforms while maintaining visual fidelity through the `resolve_grade_filter` color pipeline.

### What is the difference between preview and draft modes?

**Preview mode** enables faster encoding with a lower CRF (Constant Rate Factor) value for quick review. **Draft mode** additionally reduces resolution to 720p and uses ultra-fast encoding presets. Both modes use single-pass loudness approximation instead of the standard two-pass normalization, significantly reducing render times for iterative workflows.