How timeline_view Generates Composite Visuals On Demand in browser-use/video-use
The timeline_view helper generates on-demand "film-strip + waveform" composites by extracting video frames with ffmpeg, computing audio RMS envelopes with numpy, and rendering a unified PNG using Pillow—all within a deterministic Python pipeline that executes only when explicitly invoked.
The browser-use/video-use repository provides a specialized command-line tool for creating detailed timeline visualizations from video files. The timeline_view.py module constructs these composite visuals on demand for exact time ranges, combining thumbnail strips, audio waveforms, transcript labels, and silence indicators into a single verification image.
The Seven-Stage Rendering Pipeline
All processing occurs within helpers/timeline_view.py, operating independently of any background indexing service. The pipeline runs sequentially from extraction through final compositing.
Frame Extraction with ffmpeg
The process begins with extract_frames(), which invokes ffmpeg once per frame to pull evenly-spaced JPEGs from the specified time window. Each frame scales to a default width of 320 pixels and stores temporarily in a transient directory before compositing begins.
Audio Envelope Computation
Parallel to visual extraction, compute_envelope() processes the audio segment by dumping it to a 16 kHz mono WAV via ffmpeg. The function manually reads the PCM data and calculates windowed RMS (Root Mean Square) values across configurable windows of 2000 samples (default), producing a normalized np.ndarray of amplitude values for subsequent waveform rendering.
Transcript Integration and Silence Detection
When provided with a JSON transcript path, words_in_range() parses the speech-to-text output and returns only words whose timestamps intersect the requested window. The companion find_silences() function analyzes gaps between consecutive words, recording intervals exceeding 400 milliseconds (configurable threshold) for visual highlighting as semi-transparent blue bands.
Font Initialization
The load_font() utility attempts to load system-wide monospaced fonts—including Consolas, Monaco, and DejaVu Sans Mono—and gracefully falls back to Pillow's default bitmap font if none are available. This ensures consistent text rendering for transcript labels across different operating systems.
Canvas Composition and Rendering
The render_timeline() function orchestrates final assembly on a 1920×~500 pixel canvas with a dark background (RGB 18,18,22).
Film-Strip Layout: Extracted frames open via Pillow, resize to a uniform height of 180 pixels (frame_h), and position horizontally. The total width of all frames (total_frame_w) determines the scaling factor to fit within canvas margins while maintaining aspect ratios.
Waveform Visualization: A dark background rectangle hosts the audio visualization. The envelope array samples across the strip width; for each sample, the helper computes top and bottom Y-coordinates and draws two polylines in light blue. The region between them fills with a partly transparent polygon, creating a smooth "filled-wave" appearance.
Silence Highlighting: Using time_to_x() for coordinate conversion, the renderer draws semi-transparent blue rectangles over the waveform area corresponding to gaps detected by find_silences().
Transcript Labels: Words longer than 50 milliseconds receive consideration. The label's center X-coordinate calculates from its timestamp; if no collision exists with the previous label, a short vertical tick renders on the waveform followed by the word text in a small monospaced font.
Time Reference: Six evenly-spaced timestamp markers and second-precision labels appear below the waveform for temporal orientation.
Legend: When silence intervals exist, an explanatory line draws near the canvas bottom indicating the meaning of blue shading.
Optimized PNG Output
The final composite writes to the user-specified path using Pillow's save(..., optimize=True), minimizing file size while preserving quality. Default output directs to an edit/verify/ folder if no destination is explicitly provided.
Command-Line Usage
Invoke the helper on-demand via CLI, specifying exact time ranges and optional transcript paths:
python helpers/timeline_view.py \
path/to/video.mp4 12.0 18.5 \
--n-frames 12 \
--transcript path/to/transcript.json \
-o output/timeline_12-18.png
Arguments:
video.mp4— Source video file path12.0— Start time in seconds18.5— End time in seconds--n-frames 12— Request twelve equally spaced frames across the range--transcript— Optional JSON transcript path for word labels and silence detection-o— Destination PNG file path
Execution prints progress messages through frame extraction, envelope calculation, and final rendering stages.
Summary
- The
timeline_viewhelper generates composite visuals on demand through a deterministic Python pipeline requiring only ffmpeg, numpy, and Pillow. extract_frames()pulls scaled JPEGs via ffmpeg whilecompute_envelope()calculates RMS audio amplitudes from 16 kHz mono PCM data using 2000-sample windows.- Transcript parsing via
words_in_range()and silence detection throughfind_silences()enable contextual annotations and gap highlighting at configurable thresholds. render_timeline()assembles frames, waveforms, silence bands, word labels, and time rulers onto a 1920-pixel wide canvas before outputting optimized PNG files.
Frequently Asked Questions
What dependencies does timeline_view require?
The implementation relies solely on standard Python libraries plus ffmpeg (external binary), numpy for numerical processing, and Pillow (PIL) for image manipulation and text rendering. No continuous background services or databases are necessary.
Can I customize the number of frames in the film strip?
Yes, the --n-frames CLI argument controls frame density, passing directly to extract_frames() which calculates even temporal spacing across your specified time range. The default width scaling ensures the strip fits within the 1920-pixel canvas regardless of frame count.
How does the silence detection threshold work?
The find_silences() function compares consecutive word timestamps from the transcript JSON, marking gaps exceeding 400 milliseconds by default. These intervals render as semi-transparent blue rectangles over the waveform area, with the threshold adjustable via function parameters when calling programmatically.
Why are images generated on-demand rather than cached?
The architecture intentionally rebuilds composites for exact time ranges via CLI invocation without background indexing, ensuring fresh renders that reflect current transcript states and parameters. This "on demand" approach avoids storage overhead from pre-generated assets while providing precise visualizations for arbitrary time windows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →