# What Components Make Up the Visual Composite in video-use? Filmstrip, Waveform, and Word Labels Explained

> Understand the visual composite in video use. Explore components like filmstrip, waveform, and word labels that create clear video timelines and data visualization.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: deep-dive
- Published: 2026-07-01

---

**The visual composite in video-use is a multi-layer PNG generated by [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py) that combines a header, filmstrip, waveform background, silence shading, audio envelope, word labels, time ruler, and optional legend.**

The `browser-use/video-use` repository provides a utility to create data-rich visual summaries of video segments. At the heart of this feature is the visual composite, which layers several elements—including a filmstrip, waveform, and word labels—into a single image for quick analysis.

## Visual Composite Layers in video-use

The composite is rendered in a specific order to ensure readability. Each layer serves a distinct purpose in representing the video segment.

### Header and Metadata

The top of the image displays contextual metadata drawn with `ImageDraw.text`. This includes the video filename, the start and end times of the segment, and the total frame count. According to the source code in [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py), this header is drawn at the top of the canvas between lines 27 and 33.

### Filmstrip Frames

The **filmstrip** provides a horizontal series of evenly spaced frames extracted from the video.

The `extract_frames()` function orchestrates this by calling `ffmpeg` to grab *N* frames (specified by `--n-frames`), returning JPEG paths. These frames are then loaded, resized to a uniform height, and pasted side-by-side. The frame extraction logic resides in lines 37–62, while the rendering and resizing occur around lines 35–53.

### Waveform Background

Beneath the filmstrip, a dark rectangular bar marks the area where the audio envelope will be plotted. This simple background is filled with a color derived from the `BG` constant, defined at lines 64–66.

### Silence Shading

To highlight non-speech segments, semi-transparent blue bands are drawn over periods of silence lasting 400 ms or longer.

The process involves `words_in_range()` to extract words from the transcript (lines 18–33) and `find_silences()` to compute gaps between them (lines 35–48). Each identified gap is then drawn as a rectangle on the canvas (lines 66–73).

### Audio Envelope (Waveform)

The central **waveform** visualizes the RMS amplitude of the audio as a thin line with a filled polygon underneath.

The `compute_envelope()` function handles the heavy lifting: it extracts a mono 16 kHz PCM audio snippet using `ffmpeg`, computes a windowed RMS array, normalizes it to the range [0, 1], and returns the data for plotting. The computation is implemented in lines 68–111, with the actual drawing (top/bottom lines and translucent fill) occurring at lines 74–89.

### Word Labels

**Word labels** provide textual context by displaying short, legible text (e.g., "hello") positioned just above the waveform at each word’s midpoint.

The rendering logic filters out very short words to maintain clarity. Each label is spaced at least approximately 28 pixels from the previous one to prevent crowding, and a tiny tick mark is drawn on the waveform to anchor the label. This loop is found at lines 92–110.

### Time Ruler and Legend

A **time ruler** at the bottom provides navigational aid. A fixed number of ticks (`n_ticks = 6`) are drawn, each with a short line and a timestamp label formatted as `"{t:.2f}s"` (lines 113–121).

Finally, an optional **legend** is rendered only when silence gaps are detected in the segment, explaining that the shaded bands represent silent periods (lines 123–127).

## How to Generate the Visual Composite

You can generate this composite via command-line or programmatically.

### Command-Line Usage

Run the script directly from the terminal to produce a PNG for a specific time range:

```bash
python helpers/timeline_view.py \
    sample.mp4 12.5 17.5 \
    -o out.png \
    --n-frames 12 \
    --transcript transcripts/sample.json

```

This extracts 12 frames between 12.5 s and 17.5 s, draws the audio envelope, overlays the word labels from the supplied transcript, and saves the composite to `out.png`.

### Programmatic Usage with Python

For batch processing or integration into larger workflows, import the `render_timeline()` function:

```python
from pathlib import Path
from helpers.timeline_view import render_timeline

video = Path("sample.mp4")
start = 30.0          # seconds

end   = 45.0
output = Path("report/segment.png")
transcript = Path("edit/transcripts/sample.json")

render_timeline(
    video=video,
    start=start,
    end=end,
    out_path=output,
    n_frames=15,
    transcript=transcript,
)
print(f"Composite written to {output}")

```

This API approach is ideal for generating reports across multiple video segments automatically.

## Summary

- The visual composite in `browser-use/video-use` is built by [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py).
- It consists of eight distinct layers: header, filmstrip, waveform background, silence shading, audio envelope, word labels, time ruler, and legend.
- Frame extraction relies on `ffmpeg` via `extract_frames()`, while audio analysis uses `compute_envelope()` on mono 16 kHz PCM data.
- Word labels are spaced with a minimum 28-pixel gap to ensure readability.
- The composite can be generated via CLI or the `render_timeline()` Python API.

## Frequently Asked Questions

### What file generates the visual composite in video-use?

The [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py) file contains the core implementation. It defines the `render_timeline()` function and all helper logic for assembling the filmstrip, waveform, and word labels into a single PNG image.

### How does video-use extract frames for the filmstrip?

The `extract_frames()` function runs `ffmpeg` to grab a specified number of frames (`--n-frames`) from the video segment. It returns paths to JPEG files, which are then loaded, resized to a uniform height, and pasted side-by-side into the composite.

### What audio format is used to compute the waveform envelope?

The `compute_envelope()` function extracts audio as a mono 16 kHz PCM snippet using `ffmpeg`. It then computes a windowed RMS (Root Mean Square) array to determine the amplitude envelope that is plotted as the waveform line.

### How are word labels positioned to avoid overlap?

After filtering out very short words, the rendering loop enforces a minimum spacing of approximately 28 pixels between consecutive labels. If a word is too close to the previous one, it is skipped to prevent crowding, and a small tick mark is drawn on the waveform to indicate the word's position.