Silence Gap Detection Algorithm in Video-Use: How to Find Clean Cut Candidates

Video-Use detects silence gaps by scanning transcript timestamps for intervals between words that exceed a configurable threshold, returning reliable cut points as (start, end) tuples.

Video-Use, an open-source video editing toolkit from browser-use, implements a deterministic silence gap detection algorithm to identify optimal cutting points in video transcripts. This lightweight approach analyzes word-level timing data to find natural pauses in speech, making it ideal for automated video editing workflows that require precise, clean cuts without manual scrubbing.

How the Silence Gap Detection Algorithm Works

The core silence gap detection logic resides in helpers/timeline_view.py within the find_silences function (lines 35-48). The algorithm processes transcript data sequentially to identify pauses long enough to serve as clean cut candidates.

Input Processing and Filtering

The function accepts three parameters: a list of word-level entries, a segment time window [start, end], and a configurable threshold. Each transcript entry contains start, end, and type fields. The algorithm immediately filters out entries where type equals "spacing", as these represent non-verbal separators rather than spoken content.

Sequential Gap Analysis

The algorithm maintains a running prev_end marker initialized to the segment's start time. For each word in the filtered list:

  1. Clamp timestamps: The word's start time is clamped to the segment start using ws = max(start, w.get("start", start)) to ensure boundaries are respected.
  2. Measure gaps: If the difference between ws and prev_end is greater than or equal to the threshold (default 0.4 seconds), the algorithm records a silence tuple (prev_end, ws).
  3. Update pointer: The prev_end marker is updated to the later of the current prev_end and the word's end timestamp, ensuring overlapping words don't create false gaps.

Trailing Silence Handling

After processing all words, the algorithm performs a final check for trailing silence. If the interval between the last prev_end and the segment's end time exceeds the threshold, this final gap is also appended to the results. The function returns a list of (silence_start, silence_end) tuples representing all valid silent intervals.

Code Implementation: Using find_silences in Video-Use

To implement silence gap detection in your own workflow, import the find_silences and words_in_range functions from helpers/timeline_view.py:

from pathlib import Path
from helpers.timeline_view import words_in_range, find_silences

# Load a transcript generated by the video-use pipeline

transcript_path = Path("edit/transcripts/example_video.json")

# Define the interval you want to analyse (seconds)

segment_start = 30.0
segment_end   = 90.0

# Pull all words that intersect the interval

words = words_in_range(transcript_path, segment_start, segment_end)

# Detect silences – default threshold is 0.4 s

silences = find_silences(words, segment_start, segment_end)

print("Silence gaps ≥ 400 ms:", silences)

The words_in_range helper function extracts only the transcript entries that intersect your specified time window, ensuring the silence detection runs on relevant content only.

Configuring the Silence Threshold

The silence gap detection sensitivity is controlled via the threshold parameter. The default value of 0.4 seconds (400 ms) identifies natural pauses between sentences, but you can tune this for different content types:


# Aggressive cutting for fast-paced content (200 ms gaps)

short_silences = find_silences(words, segment_start, segment_end, threshold=0.2)
print("Silence gaps ≥ 200 ms:", short_silences)

# Conservative cutting for dramatic pauses (600 ms gaps)

long_silences = find_silences(words, segment_start, segment_end, threshold=0.6)

Changing the threshold requires no modifications to the underlying algorithm, making it straightforward to adapt the tool for podcasts, tutorials, or high-energy vlogs.

Converting Gaps to Cut Timestamps

Once you have detected silence gaps, convert them into precise cut points by selecting the midpoint of the longest silence:

if silences:
    longest = max(silences, key=lambda s: s[1] - s[0])
    clean_cut = (longest[0] + longest[1]) / 2
    print(f"Suggested clean-cut timestamp: {clean_cut:.2f}s")

This approach ensures your cuts land in the middle of silent periods, avoiding mid-word splits.

Integration with Video-Use's Timeline Visualization

The silence gap detection algorithm powers additional features throughout the Video-Use codebase. In helpers/pack_transcripts.py (lines 40-50), silence thresholds drive phrase-grouping logic that clusters words into semantic segments. Meanwhile, helpers/render.py leverages the same detection data to generate timeline visualizations that mark silent intervals on PNG exports, providing visual confirmation of where cuts will occur.

Because the algorithm relies solely on timestamp arithmetic without audio waveform analysis, it executes deterministically and completes in linear time relative to the word count.

Summary

  • Silence gap detection in Video-Use scans transcript timestamps to find pauses exceeding a configurable threshold (default 0.4 s).
  • The find_silences function in helpers/timeline_view.py (lines 35-48) implements the core algorithm using sequential iteration and gap measurement.
  • Non-speech tokens with type: "spacing" are automatically filtered to prevent false positives.
  • The algorithm handles both inter-word gaps and trailing silence at segment boundaries.
  • Threshold tuning requires only changing the threshold parameter, making the system adaptable to different content styles without code modification.

Frequently Asked Questions

What is the default silence threshold in Video-Use?

The default threshold is 0.4 seconds (400 milliseconds). This value is defined in the find_silences function signature in helpers/timeline_view.py and represents a natural pause between sentences in most spoken content.

How does Video-Use handle non-speech tokens during silence detection?

The algorithm explicitly ignores entries where the type field equals "spacing". These entries represent non-verbal separators in the transcript, and skipping them ensures that only actual speech gaps are measured for cutting purposes.

Can I use the silence detection without the full Video-Use pipeline?

Yes. The find_silences function in helpers/timeline_view.py operates independently on any list of word-level dictionaries containing start, end, and type keys. You can import and use it with custom transcript data as long as the timestamp format matches the expected structure.

How do I choose the best silence gap for a clean cut?

Select the longest silence gap using max(silences, key=lambda s: s[1] - s[0]), then calculate the midpoint between silence_start and silence_end. This places your cut at the center of the quietest moment, maximizing the distance from spoken words.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →