# Understanding the Packed Transcript Format and Phrase-Level Grouping in video-use

> Learn how video-use simplifies ElevenLabs Scribe JSON into packed markdown with phrase-level grouping and timestamps. Extract precise video data easily.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: deep-dive
- Published: 2026-07-03

---

**The `video-use` toolkit converts raw ElevenLabs Scribe JSON into a compact markdown file called [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md), where words are grouped into timestamped phrases based on silence thresholds and speaker changes.**

This packed transcript format serves as the primary text representation that LLMs consume when reasoning about video edits. The transformation from raw audio transcript to structured markdown happens in [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py), which implements a deterministic grouping algorithm to balance readability with precise timing data.

## What Is the Packed Transcript Format?

The packed transcript format is a space-efficient markdown representation generated from ElevenLabs Scribe output. While the raw JSON contains individual word-level entries—including `word`, `audio_event`, and `spacing` tokens—the packed format collapses these into **phrases** that respect natural speech boundaries.

Each phrase in the resulting [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md) file contains:
- Precise start and end timestamps (formatted as `[NNN.NN-MMM.MM]`)
- An optional speaker identifier (e.g., `S0`)
- The normalized text content

This format typically compresses an hour of transcribed content into approximately 12 KB, giving large language models essential word-boundary data without overwhelming context windows.

## The Phrase-Level Grouping Algorithm

The core logic resides in the `group_into_phrases` function within [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py). This function traverses the raw `words` array from Scribe's JSON output and builds logical speech segments through stateful iteration.

### Processing Raw Scribe Output

The input JSON structure contains an array of word objects under the `words` key. Each entry includes timing metadata (`start`, `end`) and a `type` field distinguishing between actual words, audio events, and spacing tokens. The algorithm must handle these heterogeneous entries while maintaining temporal accuracy.

### The group_into_phrases Function

**Signature:** `group_into_phrases(words, silence_threshold=0.5)`

The function accepts the raw word list and a configurable silence threshold (defaulting to **0.5 seconds**). It returns a list of dictionaries, each containing `start`, `end`, `text`, and `speaker_id` keys.

Internally, the function maintains a running buffer of words. When specific boundary conditions are met, it invokes a `flush()` helper that:
- Concatenates buffered words into normalized text
- Captures the phrase's start time (first word) and end time (last word)
- Resets the buffer for the next phrase

### Flush Conditions: Silence and Speaker Changes

The algorithm triggers a phrase flush on two primary conditions:

1. **Silence Detection:** When encountering a `spacing` entry where `gap = end - start` exceeds the threshold, or when the gap between the previous token's `end` and the current token's `start` exceeds the limit.

2. **Speaker Change Detection:** When the current token's `speaker_id` differs from the previous token's speaker, indicating a turn in conversation.

After the final iteration, `flush()` is called once more to ensure the last phrase is emitted.

## Rendering Phrases as Markdown

Once phrases are grouped, the `render_markdown` function formats them for LLM consumption. The output follows a strict structure:

- A top-level heading describing the transcript source
- Level-2 headings for each source file showing duration and phrase count
- Fixed-width timestamp lines using `format_time` to ensure six-character alignment

Each line follows this pattern:

```markdown
[002.52-005.36] S0 This is the spoken text content.

```

This consistent formatting allows regex-based parsing while remaining human-readable for manual editing workflows.

## End-to-End Workflow

The packed transcript generation integrates into the standard `video-use` pipeline:

1. **Transcription:** [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) calls ElevenLabs Scribe and stores raw JSON in `<edit_dir>/transcripts/<video>.json`
2. **Packing:** [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) processes all JSON files and writes `<edit_dir>/takes_packed.md`

Run the packing script via CLI:

```bash

# Pack all transcripts with default 0.5s threshold

python helpers/pack_transcripts.py --edit-dir ./my_edit

# Use custom silence threshold (0.8 seconds)

python helpers/pack_transcripts.py \
    --edit-dir ./my_edit \
    --silence-threshold 0.8 \
    -o ./my_edit/custom_packed.md

```

The script outputs a completion summary:

```

packed 3 transcripts → ./my_edit/takes_packed.md
  215 phrases, 12m 34.5s total runtime
  14.2 KB

```

## Programmatic Access

You can invoke the grouping logic directly from Python for custom processing:

```python
from pathlib import Path
import json
from helpers.pack_transcripts import group_into_phrases

# Load raw Scribe output

data = json.loads(Path("edit/transcripts/example.json").read_text())
words = data["words"]

# Group with default 0.5s silence threshold

phrases = group_into_phrases(words)

for p in phrases:
    speaker = f"S{p['speaker_id']}" if p.get('speaker_id') is not None else ""
    print(f"[{p['start']:.2f}-{p['end']:.2f}] {speaker} {p['text']}")

```

## Summary

- **[`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py)** implements the phrase-level grouping logic via the `group_into_phrases` function.
- The algorithm uses a **default 0.5-second silence threshold** and **speaker change detection** to determine phrase boundaries.
- Output is written to **[`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md)** as compact markdown with fixed-width timestamps and speaker tags.
- The format reduces raw transcript data by orders of magnitude while preserving word-level timing precision for LLM editing workflows.

## Frequently Asked Questions

### What is the default silence threshold and how do I change it?

The default silence threshold is **0.5 seconds**. You can override this via the `--silence-threshold` CLI flag when running [`pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/pack_transcripts.py), or by passing a different value to the `silence_threshold` parameter in the `group_into_phrases` function signature.

### How does the packed transcript format handle speaker changes?

The grouping algorithm detects speaker changes by comparing the `speaker_id` field of consecutive tokens in the raw JSON. When a change is detected, the current phrase is immediately flushed and a new phrase begins with the new speaker identifier, ensuring that each phrase contains text from only one speaker.

### What is the difference between the raw JSON and the packed markdown format?

The raw JSON contains granular word-level data including spacing tokens and audio events, while the packed markdown format ([`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md)) condenses this into human-readable phrases with speaker attribution and precise timing brackets. The packed format is optimized for LLM context windows and manual review, whereas the JSON preserves raw Scribe output for archival purposes.

### Why does video-use use a packed format instead of raw JSON for LLM consumption?

The packed format reduces token count by grouping words into phrases while maintaining essential timing metadata. An hour of content typically fits into ~12 KB of markdown, compared to significantly larger JSON representations. This compression allows the LLM to process longer video segments within context limits while still accessing precise cut points through the timestamped phrase boundaries.