Default Subtitle Chunk Sizes in video-use: How the 2-Word Algorithm Works

Default subtitle chunk sizes in video-use are 2-word chunks, with automatic breaks triggered by punctuation marks to ensure readable caption timing.

The browser-use/video-use repository generates automatic captions by grouping transcript words into fixed-size segments. According to the source code, the subtitle generation system defaults to 2-word chunks unless interrupted by punctuation boundaries, creating easily readable captions that synchronize with the video timeline.

How Subtitle Chunking Works in helpers/render.py

The subtitle generation logic resides in helpers/render.py within the build_master_srt function. This function processes word-level transcripts and groups them into discrete caption cues using a specific algorithmic approach.

The chunking mechanism operates on two primary conditions:

  • Length threshold: When the current word buffer reaches exactly 2 words, the system finalizes the chunk
  • Punctuation breaks: If a word ends with any character defined in PUNCT_BREAK, the chunk closes immediately regardless of word count

Each completed chunk is then converted to an uppercase string to form the final caption line.

The build_master_srt Function Implementation

Inside build_master_srt, the function iterates over the transcript words and maintains a temporary list called current that accumulates words until the chunking criteria are met.

As implemented in the source code, the algorithm appends each word to the current list and checks whether len(current) >= 2 or whether the current word ends with punctuation. When either condition evaluates to true, the function saves the accumulated words as a finalized chunk and resets the current buffer to empty.

The function's docstring explicitly documents this behavior as "2-word chunks (break on any punctuation in between)", confirming this is the intended default behavior rather than a configurable parameter.

Generating Subtitles with Default Settings

To generate subtitles using the default 2-word chunking algorithm, run the render script with the subtitles flag:

python helpers/render.py my_edl.json -o final.mp4 --build-subtitles

This command invokes build_master_srt internally, which receives the word-level timestamps from helpers/transcribe.py and applies the 2-word grouping logic. The resulting SRT file contains cues that display no more than two words at a time (or fewer if punctuation naturally breaks the phrase), ensuring viewers can read captions at a comfortable pace synchronized with the audio.

Summary

  • Default subtitle chunk sizes in video-use are hardcoded as 2-word chunks in the build_master_srt function
  • The chunking logic automatically breaks early when encountering punctuation defined in PUNCT_BREAK
  • All subtitle text is converted to uppercase during the rendering process
  • The helpers/render.py file contains the complete implementation, while helpers/transcribe.py supplies the word-level timing data
  • No command-line configuration exists to adjust the chunk size; modification requires editing the source code

Frequently Asked Questions

How does video-use determine when to break a subtitle chunk?

The system breaks a subtitle chunk when it accumulates 2 words or when the current word ends with punctuation. This logic is implemented in lines 43-55 of helpers/render.py, where the code checks if len(current) >= 2 or if the word terminates with a punctuation character from the PUNCT_BREAK set.

Can I change the default subtitle chunk size from 2 words to a different value?

Currently, the 2-word chunk size is hardcoded in the build_master_srt function within helpers/render.py. To use a different chunk size, you must modify the source code directly by changing the comparison value 2 in the length check to your desired word count.

Why does video-use convert all subtitles to uppercase?

The build_master_srt function transforms all caption text to uppercase as part of the chunk formatting process. This creates visual consistency across all generated subtitles and ensures maximum readability against varying video backgrounds, though the specific casing logic is applied after the 2-word chunking algorithm completes.

Where does video-use get the word timing data for subtitles?

The build_master_srt function receives word-level timestamps from helpers/transcribe.py, which handles the speech-to-text processing and alignment. This transcription module supplies the individual word entries that the render script then groups into the default 2-word chunks using the subtitle generation algorithm.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →