# How video-use Self-Evaluates Rendered Output at Cut Boundaries

> Discover how video-use self-evaluates rendered output at cut boundaries. Learn about its self-evaluation loop and composite visualizations for LLM inspection.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: internals
- Published: 2026-07-01

---

**video-use validates every cut by running a self-evaluation loop after rendering, using [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py) to generate composite visualizations of ±1.5 s windows around each cut boundary for LLM inspection.**

The browser-use/video-use repository implements a rigorous quality assurance mechanism that automatically validates every cut boundary before presenting the final output to users. This self-evaluation process ensures that audio-visual transitions remain seamless by analyzing the rendered video rather than the source timeline, catching discontinuities that traditional editing previews might miss.

## The Self-Evaluation Pipeline

The self-evaluation workflow operates as a five-stage pipeline that executes immediately after the initial render completes. Each stage is designed to detect specific failure modes that could degrade the viewing experience.

### Step 1: Render the Full Edit

The process begins in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py), which extracts each segment defined in [`edl.json`](https://github.com/browser-use/video-use/blob/main/edl.json), applies the colour-grade, and concatenates the clips into a single output file at `edit/final.mp4`. Crucially, this stage adds **30 ms audio fades** at every cut point to prevent audible pops or clicks during transitions. The script also optionally overlays subtitles or graphics before writing the final rendered video.

### Step 2: Identify Cut Boundaries from the EDL

After rendering, video-use parses the EDL ([`edl.json`](https://github.com/browser-use/video-use/blob/main/edl.json)) to determine the exact timestamps of every cut boundary. The EDL contains a list of `[start, end]` ranges for each segment, providing precise temporal coordinates for the validation windows.

### Step 3: Generate Visual Diagnostics with timeline_view

For each cut point, video-use invokes [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py) on a short window (±≈1.5 s) surrounding the boundary. This helper performs several analytical operations:

- **`extract_frames`**: Captures evenly-spaced frames from the video window
- **`compute_envelope`**: Builds an audio RMS envelope to visualize volume levels
- **`words_in_range`**: Loads the word-level transcript for the segment
- **`find_silences`**: Identifies silence gaps for optimal cut alignment

The tool composites these elements into a single PNG image displaying a filmstrip, waveform, word labels, and shaded silence regions. This visualization serves as the input for the automated quality check.

### Step 4: LLM-Based Visual Validation

The generated composite image is passed to an LLM for visual inspection. The model examines the PNG for:

- **Visual discontinuities** such as flashes or jumps at the cut boundary
- **Audio pops** indicating missing or insufficient 30 ms fades
- **Hidden subtitles** that might appear abruptly and distract viewers

By reading this "visual layer," the LLM decides whether the cut passes quality standards or requires correction.

### Step 5: Iterative Correction

If the validation detects any issues, video-use automatically re-renders the affected segment. According to [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md), the system attempts up to **three passes** to correct the cut, repeating the self-evaluation after each render until the preview passes or the maximum attempt limit is exhausted.

## Manual Validation of Cut Boundaries

You can manually trigger the self-evaluation pipeline for specific cut points using the command-line tools provided in the repository.

### Validate a Single Cut Boundary

First render the complete edit, then inspect a specific window around a cut at 12.34 s:

```bash

# Render the full edit

python helpers/render.py edit/edl.json -o edit/final.mp4

# Check ±1.5 s window around the 12.34 s cut

python helpers/timeline_view.py edit/final.mp4 10.84 13.84 -o verify/cut_12.34.png

```

This produces `verify/cut_12.34.png`, which you can inspect for jumps or audio pops before finalizing.

### Programmatic Batch Validation

For automated testing of all cuts in an edit, use the following Python pattern:

```python
import json
from pathlib import Path
from subprocess import run

edl_path = Path("edit/edl.json")
rendered = Path("edit/final.mp4")
out_dir = Path("edit/verify")
out_dir.mkdir(parents=True, exist_ok=True)

edl = json.loads(edl_path.read_text())

for i, r in enumerate(edl["ranges"]):
    start = float(r["start"])
    end = float(r["end"])
    
    # Window of ±1.5 s around the cut (except first/last cut)

    win = 1.5
    win_start = max(0.0, start - win)
    win_end = end + win
    
    out_img = out_dir / f"cut_{i:02d}_{start:.2f}-{end:.2f}.png"
    
    run([
        "python", "helpers/timeline_view.py",
        str(rendered), f"{win_start:.3f}", f"{win_end:.3f}",
        "-o", str(out_img)
    ], check=True)
    
    # The LLM now reads `out_img` and decides whether the cut passes

```

## Key Implementation Files

The self-evaluation system relies on specific helper scripts and configuration files:

- **[`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py)**: Extracts segments, applies 30 ms audio fades, concatenates clips, and builds the final video at `edit/final.mp4`
- **[`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py)**: Generates the filmstrip + waveform + word-label composite visualization used for automated inspection
- **[`edit/edl.json`](https://github.com/browser-use/video-use/blob/main/edit/edl.json)**: Contains the `[start, end]` ranges defining cut boundaries
- **[`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md)**: Formalizes the "self-eval before showing the user" rule and documents the three-attempt limit for re-rendering
- **[`README.md`](https://github.com/browser-use/video-use/blob/main/README.md)**: Describes the overall self-evaluation concept and rationale for validating rendered output rather than source timelines

## Summary

- video-use automatically validates every cut boundary after rendering by analyzing the final output file rather than the source timeline
- The system uses [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py) to generate composite PNG visualizations of ±1.5 s windows around each cut, combining video frames, audio waveforms, and transcript data
- An LLM inspects these visualizations for discontinuities, audio pops, and subtitle issues before approving the cut
- Failed cuts trigger automatic re-rendering with up to three retry attempts according to the policy defined in [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md)
- Manual validation is possible using the [`timeline_view.py`](https://github.com/browser-use/video-use/blob/main/timeline_view.py) CLI tool with custom time windows

## Frequently Asked Questions

### What is the ±1.5 s window used for in video-use self-evaluation?

The ±1.5 s window provides sufficient context around each cut boundary to detect visual discontinuities and audio artifacts that might not be visible at the exact cut point. This window captures the tail of the outgoing clip and the head of the incoming clip, allowing the [`timeline_view.py`](https://github.com/browser-use/video-use/blob/main/timeline_view.py) helper to generate a composite showing the transition dynamics, including the 30 ms fade regions and any surrounding silence gaps.

### How does video-use prevent audio pops at cut boundaries?

The [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) script automatically applies **30 ms audio fades** at every cut point during the concatenation process. These fades smooth the transition between segments with different audio levels, preventing the abrupt amplitude changes that cause audible pops or clicks. The self-evaluation step then verifies these fades were applied correctly by examining the audio waveform in the generated timeline visualization.

### What happens if a cut fails the self-evaluation check?

If the LLM detects visual discontinuities, missing audio fades, or subtitle issues in the composite image, video-use flags the cut as failed and triggers a re-render of the affected segment. According to [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md), the system will attempt correction up to **three times** before giving up, ensuring that users receive only validated, seamless transitions or a clear notification that manual intervention is required.

### Which helper script generates the composite visualization for cut validation?

The [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py) script handles all composite generation for self-evaluation. It extracts frames via `extract_frames`, computes the audio RMS envelope through `compute_envelope`, loads word-level timing data with `words_in_range`, and identifies silent regions using `find_silences`. The script composites these layers into a single PNG that serves as the input for automated visual inspection.