How video-use Self-Evaluates Rendered Output at Cut Boundaries

video-use validates every cut by running a self-evaluation loop after rendering, using helpers/timeline_view.py to generate composite visualizations of ±1.5 s windows around each cut boundary for LLM inspection.

The browser-use/video-use repository implements a rigorous quality assurance mechanism that automatically validates every cut boundary before presenting the final output to users. This self-evaluation process ensures that audio-visual transitions remain seamless by analyzing the rendered video rather than the source timeline, catching discontinuities that traditional editing previews might miss.

The Self-Evaluation Pipeline

The self-evaluation workflow operates as a five-stage pipeline that executes immediately after the initial render completes. Each stage is designed to detect specific failure modes that could degrade the viewing experience.

Step 1: Render the Full Edit

The process begins in helpers/render.py, which extracts each segment defined in edl.json, applies the colour-grade, and concatenates the clips into a single output file at edit/final.mp4. Crucially, this stage adds 30 ms audio fades at every cut point to prevent audible pops or clicks during transitions. The script also optionally overlays subtitles or graphics before writing the final rendered video.

Step 2: Identify Cut Boundaries from the EDL

After rendering, video-use parses the EDL (edl.json) to determine the exact timestamps of every cut boundary. The EDL contains a list of [start, end] ranges for each segment, providing precise temporal coordinates for the validation windows.

Step 3: Generate Visual Diagnostics with timeline_view

For each cut point, video-use invokes helpers/timeline_view.py on a short window (±≈1.5 s) surrounding the boundary. This helper performs several analytical operations:

  • extract_frames: Captures evenly-spaced frames from the video window
  • compute_envelope: Builds an audio RMS envelope to visualize volume levels
  • words_in_range: Loads the word-level transcript for the segment
  • find_silences: Identifies silence gaps for optimal cut alignment

The tool composites these elements into a single PNG image displaying a filmstrip, waveform, word labels, and shaded silence regions. This visualization serves as the input for the automated quality check.

Step 4: LLM-Based Visual Validation

The generated composite image is passed to an LLM for visual inspection. The model examines the PNG for:

  • Visual discontinuities such as flashes or jumps at the cut boundary
  • Audio pops indicating missing or insufficient 30 ms fades
  • Hidden subtitles that might appear abruptly and distract viewers

By reading this "visual layer," the LLM decides whether the cut passes quality standards or requires correction.

Step 5: Iterative Correction

If the validation detects any issues, video-use automatically re-renders the affected segment. According to SKILL.md, the system attempts up to three passes to correct the cut, repeating the self-evaluation after each render until the preview passes or the maximum attempt limit is exhausted.

Manual Validation of Cut Boundaries

You can manually trigger the self-evaluation pipeline for specific cut points using the command-line tools provided in the repository.

Validate a Single Cut Boundary

First render the complete edit, then inspect a specific window around a cut at 12.34 s:


# Render the full edit

python helpers/render.py edit/edl.json -o edit/final.mp4

# Check ±1.5 s window around the 12.34 s cut

python helpers/timeline_view.py edit/final.mp4 10.84 13.84 -o verify/cut_12.34.png

This produces verify/cut_12.34.png, which you can inspect for jumps or audio pops before finalizing.

Programmatic Batch Validation

For automated testing of all cuts in an edit, use the following Python pattern:

import json
from pathlib import Path
from subprocess import run

edl_path = Path("edit/edl.json")
rendered = Path("edit/final.mp4")
out_dir = Path("edit/verify")
out_dir.mkdir(parents=True, exist_ok=True)

edl = json.loads(edl_path.read_text())

for i, r in enumerate(edl["ranges"]):
    start = float(r["start"])
    end = float(r["end"])
    
    # Window of ±1.5 s around the cut (except first/last cut)

    win = 1.5
    win_start = max(0.0, start - win)
    win_end = end + win
    
    out_img = out_dir / f"cut_{i:02d}_{start:.2f}-{end:.2f}.png"
    
    run([
        "python", "helpers/timeline_view.py",
        str(rendered), f"{win_start:.3f}", f"{win_end:.3f}",
        "-o", str(out_img)
    ], check=True)
    
    # The LLM now reads `out_img` and decides whether the cut passes

Key Implementation Files

The self-evaluation system relies on specific helper scripts and configuration files:

  • helpers/render.py: Extracts segments, applies 30 ms audio fades, concatenates clips, and builds the final video at edit/final.mp4
  • helpers/timeline_view.py: Generates the filmstrip + waveform + word-label composite visualization used for automated inspection
  • edit/edl.json: Contains the [start, end] ranges defining cut boundaries
  • SKILL.md: Formalizes the "self-eval before showing the user" rule and documents the three-attempt limit for re-rendering
  • README.md: Describes the overall self-evaluation concept and rationale for validating rendered output rather than source timelines

Summary

  • video-use automatically validates every cut boundary after rendering by analyzing the final output file rather than the source timeline
  • The system uses helpers/timeline_view.py to generate composite PNG visualizations of ±1.5 s windows around each cut, combining video frames, audio waveforms, and transcript data
  • An LLM inspects these visualizations for discontinuities, audio pops, and subtitle issues before approving the cut
  • Failed cuts trigger automatic re-rendering with up to three retry attempts according to the policy defined in SKILL.md
  • Manual validation is possible using the timeline_view.py CLI tool with custom time windows

Frequently Asked Questions

What is the ±1.5 s window used for in video-use self-evaluation?

The ±1.5 s window provides sufficient context around each cut boundary to detect visual discontinuities and audio artifacts that might not be visible at the exact cut point. This window captures the tail of the outgoing clip and the head of the incoming clip, allowing the timeline_view.py helper to generate a composite showing the transition dynamics, including the 30 ms fade regions and any surrounding silence gaps.

How does video-use prevent audio pops at cut boundaries?

The helpers/render.py script automatically applies 30 ms audio fades at every cut point during the concatenation process. These fades smooth the transition between segments with different audio levels, preventing the abrupt amplitude changes that cause audible pops or clicks. The self-evaluation step then verifies these fades were applied correctly by examining the audio waveform in the generated timeline visualization.

What happens if a cut fails the self-evaluation check?

If the LLM detects visual discontinuities, missing audio fades, or subtitle issues in the composite image, video-use flags the cut as failed and triggers a re-render of the affected segment. According to SKILL.md, the system will attempt correction up to three times before giving up, ensuring that users receive only validated, seamless transitions or a clear notification that manual intervention is required.

Which helper script generates the composite visualization for cut validation?

The helpers/timeline_view.py script handles all composite generation for self-evaluation. It extracts frames via extract_frames, computes the audio RMS envelope through compute_envelope, loads word-level timing data with words_in_range, and identifies silent regions using find_silences. The script composites these layers into a single PNG that serves as the input for automated visual inspection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →