What Components Make Up the Visual Composite in video-use? Filmstrip, Waveform, and Word Labels Explained

The visual composite in video-use is a multi-layer PNG generated by helpers/timeline_view.py that combines a header, filmstrip, waveform background, silence shading, audio envelope, word labels, time ruler, and optional legend.

The browser-use/video-use repository provides a utility to create data-rich visual summaries of video segments. At the heart of this feature is the visual composite, which layers several elements—including a filmstrip, waveform, and word labels—into a single image for quick analysis.

Visual Composite Layers in video-use

The composite is rendered in a specific order to ensure readability. Each layer serves a distinct purpose in representing the video segment.

Header and Metadata

The top of the image displays contextual metadata drawn with ImageDraw.text. This includes the video filename, the start and end times of the segment, and the total frame count. According to the source code in helpers/timeline_view.py, this header is drawn at the top of the canvas between lines 27 and 33.

Filmstrip Frames

The filmstrip provides a horizontal series of evenly spaced frames extracted from the video.

The extract_frames() function orchestrates this by calling ffmpeg to grab N frames (specified by --n-frames), returning JPEG paths. These frames are then loaded, resized to a uniform height, and pasted side-by-side. The frame extraction logic resides in lines 37–62, while the rendering and resizing occur around lines 35–53.

Waveform Background

Beneath the filmstrip, a dark rectangular bar marks the area where the audio envelope will be plotted. This simple background is filled with a color derived from the BG constant, defined at lines 64–66.

Silence Shading

To highlight non-speech segments, semi-transparent blue bands are drawn over periods of silence lasting 400 ms or longer.

The process involves words_in_range() to extract words from the transcript (lines 18–33) and find_silences() to compute gaps between them (lines 35–48). Each identified gap is then drawn as a rectangle on the canvas (lines 66–73).

Audio Envelope (Waveform)

The central waveform visualizes the RMS amplitude of the audio as a thin line with a filled polygon underneath.

The compute_envelope() function handles the heavy lifting: it extracts a mono 16 kHz PCM audio snippet using ffmpeg, computes a windowed RMS array, normalizes it to the range [0, 1], and returns the data for plotting. The computation is implemented in lines 68–111, with the actual drawing (top/bottom lines and translucent fill) occurring at lines 74–89.

Word Labels

Word labels provide textual context by displaying short, legible text (e.g., "hello") positioned just above the waveform at each word’s midpoint.

The rendering logic filters out very short words to maintain clarity. Each label is spaced at least approximately 28 pixels from the previous one to prevent crowding, and a tiny tick mark is drawn on the waveform to anchor the label. This loop is found at lines 92–110.

Time Ruler and Legend

A time ruler at the bottom provides navigational aid. A fixed number of ticks (n_ticks = 6) are drawn, each with a short line and a timestamp label formatted as "{t:.2f}s" (lines 113–121).

Finally, an optional legend is rendered only when silence gaps are detected in the segment, explaining that the shaded bands represent silent periods (lines 123–127).

How to Generate the Visual Composite

You can generate this composite via command-line or programmatically.

Command-Line Usage

Run the script directly from the terminal to produce a PNG for a specific time range:

python helpers/timeline_view.py \
    sample.mp4 12.5 17.5 \
    -o out.png \
    --n-frames 12 \
    --transcript transcripts/sample.json

This extracts 12 frames between 12.5 s and 17.5 s, draws the audio envelope, overlays the word labels from the supplied transcript, and saves the composite to out.png.

Programmatic Usage with Python

For batch processing or integration into larger workflows, import the render_timeline() function:

from pathlib import Path
from helpers.timeline_view import render_timeline

video = Path("sample.mp4")
start = 30.0          # seconds

end   = 45.0
output = Path("report/segment.png")
transcript = Path("edit/transcripts/sample.json")

render_timeline(
    video=video,
    start=start,
    end=end,
    out_path=output,
    n_frames=15,
    transcript=transcript,
)
print(f"Composite written to {output}")

This API approach is ideal for generating reports across multiple video segments automatically.

Summary

  • The visual composite in browser-use/video-use is built by helpers/timeline_view.py.
  • It consists of eight distinct layers: header, filmstrip, waveform background, silence shading, audio envelope, word labels, time ruler, and legend.
  • Frame extraction relies on ffmpeg via extract_frames(), while audio analysis uses compute_envelope() on mono 16 kHz PCM data.
  • Word labels are spaced with a minimum 28-pixel gap to ensure readability.
  • The composite can be generated via CLI or the render_timeline() Python API.

Frequently Asked Questions

What file generates the visual composite in video-use?

The helpers/timeline_view.py file contains the core implementation. It defines the render_timeline() function and all helper logic for assembling the filmstrip, waveform, and word labels into a single PNG image.

How does video-use extract frames for the filmstrip?

The extract_frames() function runs ffmpeg to grab a specified number of frames (--n-frames) from the video segment. It returns paths to JPEG files, which are then loaded, resized to a uniform height, and pasted side-by-side into the composite.

What audio format is used to compute the waveform envelope?

The compute_envelope() function extracts audio as a mono 16 kHz PCM snippet using ffmpeg. It then computes a windowed RMS (Root Mean Square) array to determine the amplitude envelope that is plotted as the waveform line.

How are word labels positioned to avoid overlap?

After filtering out very short words, the rendering loop enforces a minimum spacing of approximately 28 pixels between consecutive labels. If a word is too close to the previous one, it is skipped to prevent crowding, and a small tick mark is drawn on the waveform to indicate the word's position.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →