# Per-Segment Extract and Lossless Concatenation Pipeline in video-use

> Learn about the video-use per segment extract and lossless concatenation pipeline. Preserve exact video quality using FFmpeg's concat demuxer for efficient processing.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: internals
- Published: 2026-07-03

---

**The video-use rendering engine processes footage by first extracting individual clips with color grading and audio fades, then concatenating them losslessly using FFmpeg's concat demuxer to preserve exact quality before final compositing.**

The `video-use` repository implements a three-phase rendering pipeline that separates clip processing from final assembly. At its core, the **per-segment extract and lossless concatenation pipeline** found in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) extracts each edit decision list (EDL) range individually, applies transformations, and stitches them together without re-encoding to maintain pristine quality.

## Per-Segment Extraction with `extract_segment`

The `extract_segment` function in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) (lines 152-190) handles individual clip extraction with precise frame-level control. This function processes each segment defined in the EDL through a series of FFmpeg filter chains to prepare footage for assembly.

### Fast Seeking and Input Handling

Accurate seeking is critical for frame-perfect extraction. The pipeline places the `-ss` flag **before** the input file to enable FFmpeg's fast seek mechanism, avoiding slow decoding of unwanted frames:

```python
"-ss", f"{seg_start:.3f}", "-i", str(source)

```

This positioning ensures the decoder jumps directly to the segment start time, significantly accelerating processing for long source files.

### HDR Tone Mapping and Portrait Scaling

The pipeline automatically detects HDR content and applies tone mapping when `is_hdr_source` returns true. The `TONEMAP_CHAIN` filter is prepended to the video filter graph to convert HDR to SDR while preserving color accuracy.

For mobile footage, the `is_portrait_source` check determines orientation and applies appropriate scaling:

- **Landscape**: `scale=1920:-2`
- **Portrait**: `scale=-2:1920`

This maintains 1920px on the longest edge while preserving aspect ratios.

### Color Grading and Audio Fades

The resolved `grade_filter` (whether preset, raw FFmpeg filter, or auto-grade output) is appended to the video filter chain via `vf_parts.append(grade_filter)`.

To prevent audio pops at clip boundaries, the pipeline applies 30-millisecond fade curves:

```python
af = f"afade=t=in:st=0:d=0.03,afade=t=out:st={fade_out_start:.3f}:d=0.03"

```

These fades are applied via the `-af` flag during the extraction pass.

### Quality Presets and Encoding

`extract_segment` selects from three quality ladders based on the render mode:

- **Draft**: 720p, ultrafast preset, CRF 28
- **Preview**: 1080p, medium preset, CRF 22  
- **Final**: 1080p, fast preset, CRF 20

The function encodes to `libx264` with `yuv420p` pixel format, ensuring compatibility with the subsequent lossless concatenation step.

## Lossless Concatenation via `concat_segments`

After all segments are extracted, the `concat_segments` function (lines 267-283 in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py)) assembles them without re-encoding. This preservation step is crucial for maintaining the exact quality established during extraction.

### Creating the Concatenation List

The function generates a temporary text file named [`_concat.txt`](https://github.com/browser-use/video-use/blob/main/_concat.txt) containing absolute paths to each extracted segment:

```python
concat_list.write_text("".join(f"file '{p.resolve()}'\n" for p in segment_paths))

```

This list follows FFmpeg's concat demuxer format, with one file entry per line.

### FFmpeg Copy Mode and Stream Preservation

The concatenation command uses `-c copy` to stream-copy both video and audio without decoding:

```bash
ffmpeg -f concat -safe 0 -i _concat.txt -c copy -movflags +faststart output.mp4

```

Because all extracted segments share identical codec parameters (`libx264`, `yuv420p`, and matching resolution), FFmpeg can stitch them instantaneously while preserving exact timestamps and quality. The `-movflags +faststart` option ensures web-ready playback by moving the moov atom to the file header.

After successful concatenation, the temporary list file is removed via `concat_list.unlink(missing_ok=True)`.

## Orchestrating the Full Pipeline

The `extract_all_segments` helper (lines 214-260) coordinates the extraction phase by iterating over every range in the EDL. It resolves per-segment grades (including auto-grading when `grade: "auto"` is specified) and invokes `extract_segment` for each clip.

### From EDL to Final Base Video

In practice, the pipeline executes as follows:

```bash

# Extract segments and concatenate (preview quality)

python helpers/render.py my_edl.json -o final.mp4 --preview

# Internal flow:

#   1. extract_all_segments → extract_segment (per-clip processing)

#   2. concat_segments → base_preview.mp4 (lossless assembly)

#   3. build_final_composite (overlays/subtitles)

```

When run without `--preview` or `--draft` flags, the pipeline uses final quality settings (1080p, CRF 20). The concatenated base video is stored as `base.mp4` in the EDL directory, ready for the final compositing stage where overlays and subtitles are applied.

## Summary

- **Per-segment extraction** in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) processes each EDL range individually with color grading, HDR tone mapping, and 30ms audio fades.
- **Fast seeking** is achieved by placing `-ss` before the input file, while portrait orientation triggers height-based scaling to maintain 1920px on the longest edge.
- **Lossless concatenation** uses FFmpeg's concat demuxer with `-c copy` to assemble clips without quality degradation, requiring identical codec parameters across all segments.
- **Quality ladders** provide three presets (draft, preview, final) with varying CRF values and resolutions to balance speed against quality during iteration.
- **Automatic cleanup** removes temporary concat list files after successful assembly, leaving only the final base video for compositing.

## Frequently Asked Questions

### What is the purpose of the 30ms audio fades in video-use?

The 30-millisecond fade-in and fade-out prevents audible pops and clicks at clip boundaries. These artifacts occur when audio waveforms start or end abruptly at non-zero crossings; the `afade` filter ramps the amplitude smoothly from and to silence over 30ms, ensuring clean transitions between concatenated segments.

### Why does video-use use `-c copy` instead of re-encoding during concatenation?

The `-c copy` flag performs stream copying without decoding and re-encoding, which preserves the exact bitstream quality established during the extraction phase. Since all segments are already encoded with identical parameters (`libx264`, `yuv420p`, matching resolution), FFmpeg can concatenate them instantly while avoiding generation loss and significantly reducing processing time.

### How does the pipeline handle HDR source footage?

When `is_hdr_source` detects HDR content, the pipeline prepends the `TONEMAP_CHAIN` filter to the video filter graph before color grading. This converts high dynamic range footage to standard dynamic range using tone mapping algorithms, ensuring consistent color reproduction across the final output without clipping highlights or crushing shadows.

### What is the difference between the draft, preview, and final quality presets?

The three presets trade quality for processing speed: **Draft** (720p, ultrafast, CRF 28) provides quick iteration for rough cuts; **Preview** (1080p, medium, CRF 22) offers near-final quality for client review; **Final** (1080p, fast, CRF 20) delivers broadcast-quality output with optimized encoding. Each preset adjusts both resolution and encoding parameters to match the intended use case.