# How timeline_view Generates Composite Visuals On Demand in browser-use/video-use

> Learn how timeline_view creates on-demand composite visuals using ffmpeg, numpy, and Pillow. Discover the deterministic Python pipeline for efficient video and audio rendering.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: internals
- Published: 2026-06-29

---

**The `timeline_view` helper generates on-demand "film-strip + waveform" composites by extracting video frames with ffmpeg, computing audio RMS envelopes with numpy, and rendering a unified PNG using Pillow—all within a deterministic Python pipeline that executes only when explicitly invoked.**

The `browser-use/video-use` repository provides a specialized command-line tool for creating detailed timeline visualizations from video files. The [`timeline_view.py`](https://github.com/browser-use/video-use/blob/main/timeline_view.py) module constructs these **composite visuals** on demand for exact time ranges, combining thumbnail strips, audio waveforms, transcript labels, and silence indicators into a single verification image.

## The Seven-Stage Rendering Pipeline

All processing occurs within [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py), operating independently of any background indexing service. The pipeline runs sequentially from extraction through final compositing.

### Frame Extraction with ffmpeg

The process begins with `extract_frames()`, which invokes ffmpeg once per frame to pull evenly-spaced JPEGs from the specified time window. Each frame scales to a default width of 320 pixels and stores temporarily in a transient directory before compositing begins.

### Audio Envelope Computation

Parallel to visual extraction, `compute_envelope()` processes the audio segment by dumping it to a 16 kHz mono WAV via ffmpeg. The function manually reads the PCM data and calculates windowed **RMS (Root Mean Square)** values across configurable windows of 2000 samples (default), producing a normalized `np.ndarray` of amplitude values for subsequent waveform rendering.

### Transcript Integration and Silence Detection

When provided with a JSON transcript path, `words_in_range()` parses the speech-to-text output and returns only words whose timestamps intersect the requested window. The companion `find_silences()` function analyzes gaps between consecutive words, recording intervals exceeding **400 milliseconds** (configurable threshold) for visual highlighting as semi-transparent blue bands.

### Font Initialization

The `load_font()` utility attempts to load system-wide monospaced fonts—including Consolas, Monaco, and DejaVu Sans Mono—and gracefully falls back to Pillow's default bitmap font if none are available. This ensures consistent text rendering for transcript labels across different operating systems.

### Canvas Composition and Rendering

The `render_timeline()` function orchestrates final assembly on a **1920×~500 pixel canvas** with a dark background (RGB 18,18,22).

**Film-Strip Layout:** Extracted frames open via Pillow, resize to a uniform height of **180 pixels** (`frame_h`), and position horizontally. The total width of all frames (`total_frame_w`) determines the scaling factor to fit within canvas margins while maintaining aspect ratios.

**Waveform Visualization:** A dark background rectangle hosts the audio visualization. The envelope array samples across the strip width; for each sample, the helper computes top and bottom Y-coordinates and draws two polylines in light blue. The region between them fills with a partly transparent polygon, creating a smooth "filled-wave" appearance.

**Silence Highlighting:** Using `time_to_x()` for coordinate conversion, the renderer draws semi-transparent blue rectangles over the waveform area corresponding to gaps detected by `find_silences()`.

**Transcript Labels:** Words longer than 50 milliseconds receive consideration. The label's center X-coordinate calculates from its timestamp; if no collision exists with the previous label, a short vertical tick renders on the waveform followed by the word text in a small monospaced font.

**Time Reference:** Six evenly-spaced timestamp markers and second-precision labels appear below the waveform for temporal orientation.

**Legend:** When silence intervals exist, an explanatory line draws near the canvas bottom indicating the meaning of blue shading.

### Optimized PNG Output

The final composite writes to the user-specified path using Pillow's `save(..., optimize=True)`, minimizing file size while preserving quality. Default output directs to an `edit/verify/` folder if no destination is explicitly provided.

## Command-Line Usage

Invoke the helper on-demand via CLI, specifying exact time ranges and optional transcript paths:

```bash
python helpers/timeline_view.py \
    path/to/video.mp4 12.0 18.5 \
    --n-frames 12 \
    --transcript path/to/transcript.json \
    -o output/timeline_12-18.png

```

**Arguments:**
- `video.mp4` — Source video file path
- `12.0` — Start time in seconds
- `18.5` — End time in seconds  
- `--n-frames 12` — Request twelve equally spaced frames across the range
- `--transcript` — Optional JSON transcript path for word labels and silence detection
- `-o` — Destination PNG file path

Execution prints progress messages through frame extraction, envelope calculation, and final rendering stages.

## Summary

- The `timeline_view` helper generates composite visuals on demand through a deterministic Python pipeline requiring only ffmpeg, numpy, and Pillow.
- `extract_frames()` pulls scaled JPEGs via ffmpeg while `compute_envelope()` calculates RMS audio amplitudes from 16 kHz mono PCM data using 2000-sample windows.
- Transcript parsing via `words_in_range()` and silence detection through `find_silences()` enable contextual annotations and gap highlighting at configurable thresholds.
- `render_timeline()` assembles frames, waveforms, silence bands, word labels, and time rulers onto a 1920-pixel wide canvas before outputting optimized PNG files.

## Frequently Asked Questions

### What dependencies does timeline_view require?

The implementation relies solely on standard Python libraries plus **ffmpeg** (external binary), **numpy** for numerical processing, and **Pillow** (PIL) for image manipulation and text rendering. No continuous background services or databases are necessary.

### Can I customize the number of frames in the film strip?

Yes, the `--n-frames` CLI argument controls frame density, passing directly to `extract_frames()` which calculates even temporal spacing across your specified time range. The default width scaling ensures the strip fits within the 1920-pixel canvas regardless of frame count.

### How does the silence detection threshold work?

The `find_silences()` function compares consecutive word timestamps from the transcript JSON, marking gaps exceeding **400 milliseconds** by default. These intervals render as semi-transparent blue rectangles over the waveform area, with the threshold adjustable via function parameters when calling programmatically.

### Why are images generated on-demand rather than cached?

The architecture intentionally rebuilds composites for exact time ranges via CLI invocation without background indexing, ensuring fresh renders that reflect current transcript states and parameters. This "on demand" approach avoids storage overhead from pre-generated assets while providing precise visualizations for arbitrary time windows.