# How pack_transcripts.py Segments Phrases by Silence or Speaker Change

> Discover how pack_transcripts.py segments phrases by identifying silence, speaker changes, and temporal gaps. Optimize your transcript processing with our detailed guide.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: internals
- Published: 2026-07-05

---

**The `group_into_phrases` function in [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) detects phrase boundaries by monitoring three conditions: explicit silence entries (spacing objects) exceeding a configurable threshold, changes in `speaker_id` between consecutive tokens, and temporal gaps between words that exceed the silence threshold even without explicit spacing annotations.**

The `browser-use/video-use` repository provides a specialized transcription pipeline that converts raw Scribe output into compact, phrase-level markdown files. The [`pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/pack_transcripts.py) script performs the critical segmentation step, analyzing word-level JSON to identify natural breaks in speech. This article examines the exact implementation details of how the script determines when one phrase ends and another begins.

## The Core Segmentation Logic

The segmentation engine resides in the **`group_into_phrases`** function, defined at lines 38-44 of [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py). This function iterates through the list of word-level entries from the Scribe transcript and maintains a running buffer of current words, tracking the active speaker and start time.

When specific conditions are met during iteration, the function triggers a `flush()` operation that finalizes the current phrase and resets the buffers for the next segment. The decision logic implements three distinct detection rules to handle various transcript edge cases.

## Three Rules for Phrase Boundary Detection

The script employs a multi-layered approach to detect phrase boundaries, ensuring robust segmentation even with noisy or incomplete transcript data.

### Silence Threshold Detection via Spacing Entries

Scribe transcripts include explicit *spacing* entries representing silent intervals between words. The script detects these entries by checking `if t == "spacing"` (lines 90-99) and calculates the gap duration as `gap = end - start`.

When `gap >= silence_threshold` (default 0.5 seconds), the script immediately calls `flush()` to end the current phrase. This handles natural pauses in speech where the speaker stops for half a second or longer.

### Speaker Change Detection via speaker_id

Each word or audio event in the transcript carries a `speaker_id` field. The script maintains a `current_speaker` variable tracking the active speaker for the phrase under construction.

At lines 108-110, the code checks: `if current_speaker is not None and speaker is not None and speaker != current_speaker: flush()`. When the incoming token's speaker differs from the active phrase's speaker, the current phrase is finalized and a new phrase begins with the new speaker ID.

### Safety Gap Detection Between Tokens

As a defensive measure against transcripts missing explicit spacing entries, the script implements a fallback gap detection mechanism. It maintains a `prev_end` variable tracking the end timestamp of the previous token.

At lines 112-114, the code calculates `start - prev_end` and compares it against the silence threshold. If this gap meets or exceeds the threshold, `flush()` triggers even without an explicit spacing entry, ensuring long pauses are never swallowed into adjacent phrases.

## How the flush() Function Constructs Phrases

When any boundary condition triggers, the nested **`flush()`** helper (lines 54-82) assembles the accumulated words into a structured phrase dictionary. The construction process follows four steps:

1. **Text Collection**: Concatenates words from `current_words`, strips whitespace, and wraps audio events in parentheses.
2. **Punctuation Cleaning**: Fixes stray spaces before commas and periods using regex cleanup (line 74).
3. **Timestamp Determination**: Sets `start` from the first word's timestamp and `end` from the last word's `end` field (or its `start` if `end` is missing) (line 75).
4. **Phrase Assembly**: Appends a dictionary containing `start`, `end`, `text`, and `speaker_id` to the results list (lines 76-80).

After flushing, the script resets `current_words`, `current_start`, and `current_speaker` to `None`, preparing the buffers for the next phrase accumulation.

## Complete Processing Pipeline

The script executes a three-stage pipeline to transform raw transcripts:

1. **JSON Parsing**: Loads the transcript file and extracts the word list via `words = data.get("words", [])` (line 28).
2. **Phrase Grouping**: Calls `group_into_phrases(words, silence_threshold=0.5)` to apply the segmentation rules (line 30).
3. **Markdown Rendering**: The `render_markdown` function (lines 59-61) formats each phrase as `[start-end]` with optional speaker tags, producing the final [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md) output.

## Practical Usage Examples

### Command-Line Segmentation

Run the script from the repository root to process all transcripts in an edit directory:

```bash
python helpers/pack_transcripts.py --edit-dir /path/to/edit --silence-threshold 0.6

```

This reads all `*.json` files under `<edit>/transcripts/`, applies the segmentation rules with a 0.6-second silence threshold, and writes [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md) in the edit directory.

### Programmatic Phrase Extraction

Import and use the segmentation logic directly in Python:

```python
from pathlib import Path
import json
from helpers.pack_transcripts import group_into_phrases

# Load a transcript JSON file

transcript_path = Path("my_edit/transcripts/scene1.json")
data = json.loads(transcript_path.read_text())
words = data["words"]

# Segment into phrases (silence threshold = 0.5s)

phrases = group_into_phrases(words, silence_threshold=0.5)

for p in phrases:
    print(f"[{p['start']:.2f}-{p['end']:.2f}] {p.get('speaker_id','?')}: {p['text']}")

```

The function returns a list of dictionaries: `[{ "start": 1.23, "end": 3.45, "text": "...", "speaker_id": "speaker_0" }, ...]`.

### Manual Markdown Generation

To render phrases to markdown format programmatically:

```python
from helpers.pack_transcripts import render_markdown

# entries is a list of (name, duration, phrases) tuples

markdown = render_markdown(entries, silence_threshold=0.5)
print(markdown)

```

This produces the same formatted output as the CLI, with timestamps and speaker annotations.

## Summary

- **Three detection rules** drive phrase segmentation: explicit spacing entries (≥0.5s default), `speaker_id` changes, and inferred gaps between consecutive tokens.
- **The `flush()` function** (lines 54-82) handles phrase finalization by cleaning text, fixing punctuation spacing, and calculating accurate start/end timestamps.
- **The `group_into_phrases` function** serves as the primary API, accepting a word list and silence threshold parameter to return structured phrase dictionaries.
- **Source location**: All logic resides in [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) within the `browser-use/video-use` repository.

## Frequently Asked Questions

### What is the default silence threshold in pack_transcripts.py?

The default silence threshold is **0.5 seconds** (500 milliseconds). This value is passed to `group_into_phrases` as the `silence_threshold` parameter and applies to both explicit spacing entries and inferred gaps between consecutive words.

### How does pack_transcripts.py handle transcripts without speaker IDs?

If `speaker_id` fields are missing or `None`, the speaker change detection at lines 108-110 evaluates to `False` due to the `is not None` checks. The script continues processing but omits speaker labels from the output markdown, focusing solely on temporal gaps for segmentation.

### Can I adjust phrase segmentation without modifying the source code?

Yes. Use the `--silence-threshold` CLI argument to specify a custom threshold in seconds (e.g., `--silence-threshold 1.0` for 1-second pauses). When calling `group_into_phrases` programmatically, pass your desired threshold as the second argument: `group_into_phrases(words, silence_threshold=0.3)`.

### Where does the raw word-level data processed by pack_transcripts.py originate?

The input data comes from Scribe transcription JSON files, typically generated by [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) or batch-processed via [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) in the same repository. These files contain word-level entries with `start`, `end`, `text`, and optional `speaker_id` fields, plus explicit `spacing` entries representing silent intervals.