Understanding the Packed Transcript Format and Phrase-Level Grouping in video-use

The video-use toolkit converts raw ElevenLabs Scribe JSON into a compact markdown file called takes_packed.md, where words are grouped into timestamped phrases based on silence thresholds and speaker changes.

This packed transcript format serves as the primary text representation that LLMs consume when reasoning about video edits. The transformation from raw audio transcript to structured markdown happens in helpers/pack_transcripts.py, which implements a deterministic grouping algorithm to balance readability with precise timing data.

What Is the Packed Transcript Format?

The packed transcript format is a space-efficient markdown representation generated from ElevenLabs Scribe output. While the raw JSON contains individual word-level entries—including word, audio_event, and spacing tokens—the packed format collapses these into phrases that respect natural speech boundaries.

Each phrase in the resulting takes_packed.md file contains:

  • Precise start and end timestamps (formatted as [NNN.NN-MMM.MM])
  • An optional speaker identifier (e.g., S0)
  • The normalized text content

This format typically compresses an hour of transcribed content into approximately 12 KB, giving large language models essential word-boundary data without overwhelming context windows.

The Phrase-Level Grouping Algorithm

The core logic resides in the group_into_phrases function within helpers/pack_transcripts.py. This function traverses the raw words array from Scribe's JSON output and builds logical speech segments through stateful iteration.

Processing Raw Scribe Output

The input JSON structure contains an array of word objects under the words key. Each entry includes timing metadata (start, end) and a type field distinguishing between actual words, audio events, and spacing tokens. The algorithm must handle these heterogeneous entries while maintaining temporal accuracy.

The group_into_phrases Function

Signature: group_into_phrases(words, silence_threshold=0.5)

The function accepts the raw word list and a configurable silence threshold (defaulting to 0.5 seconds). It returns a list of dictionaries, each containing start, end, text, and speaker_id keys.

Internally, the function maintains a running buffer of words. When specific boundary conditions are met, it invokes a flush() helper that:

  • Concatenates buffered words into normalized text
  • Captures the phrase's start time (first word) and end time (last word)
  • Resets the buffer for the next phrase

Flush Conditions: Silence and Speaker Changes

The algorithm triggers a phrase flush on two primary conditions:

  1. Silence Detection: When encountering a spacing entry where gap = end - start exceeds the threshold, or when the gap between the previous token's end and the current token's start exceeds the limit.

  2. Speaker Change Detection: When the current token's speaker_id differs from the previous token's speaker, indicating a turn in conversation.

After the final iteration, flush() is called once more to ensure the last phrase is emitted.

Rendering Phrases as Markdown

Once phrases are grouped, the render_markdown function formats them for LLM consumption. The output follows a strict structure:

  • A top-level heading describing the transcript source
  • Level-2 headings for each source file showing duration and phrase count
  • Fixed-width timestamp lines using format_time to ensure six-character alignment

Each line follows this pattern:

[002.52-005.36] S0 This is the spoken text content.

This consistent formatting allows regex-based parsing while remaining human-readable for manual editing workflows.

End-to-End Workflow

The packed transcript generation integrates into the standard video-use pipeline:

  1. Transcription: helpers/transcribe.py calls ElevenLabs Scribe and stores raw JSON in <edit_dir>/transcripts/<video>.json
  2. Packing: helpers/pack_transcripts.py processes all JSON files and writes <edit_dir>/takes_packed.md

Run the packing script via CLI:


# Pack all transcripts with default 0.5s threshold

python helpers/pack_transcripts.py --edit-dir ./my_edit

# Use custom silence threshold (0.8 seconds)

python helpers/pack_transcripts.py \
    --edit-dir ./my_edit \
    --silence-threshold 0.8 \
    -o ./my_edit/custom_packed.md

The script outputs a completion summary:


packed 3 transcripts → ./my_edit/takes_packed.md
  215 phrases, 12m 34.5s total runtime
  14.2 KB

Programmatic Access

You can invoke the grouping logic directly from Python for custom processing:

from pathlib import Path
import json
from helpers.pack_transcripts import group_into_phrases

# Load raw Scribe output

data = json.loads(Path("edit/transcripts/example.json").read_text())
words = data["words"]

# Group with default 0.5s silence threshold

phrases = group_into_phrases(words)

for p in phrases:
    speaker = f"S{p['speaker_id']}" if p.get('speaker_id') is not None else ""
    print(f"[{p['start']:.2f}-{p['end']:.2f}] {speaker} {p['text']}")

Summary

  • helpers/pack_transcripts.py implements the phrase-level grouping logic via the group_into_phrases function.
  • The algorithm uses a default 0.5-second silence threshold and speaker change detection to determine phrase boundaries.
  • Output is written to takes_packed.md as compact markdown with fixed-width timestamps and speaker tags.
  • The format reduces raw transcript data by orders of magnitude while preserving word-level timing precision for LLM editing workflows.

Frequently Asked Questions

What is the default silence threshold and how do I change it?

The default silence threshold is 0.5 seconds. You can override this via the --silence-threshold CLI flag when running pack_transcripts.py, or by passing a different value to the silence_threshold parameter in the group_into_phrases function signature.

How does the packed transcript format handle speaker changes?

The grouping algorithm detects speaker changes by comparing the speaker_id field of consecutive tokens in the raw JSON. When a change is detected, the current phrase is immediately flushed and a new phrase begins with the new speaker identifier, ensuring that each phrase contains text from only one speaker.

What is the difference between the raw JSON and the packed markdown format?

The raw JSON contains granular word-level data including spacing tokens and audio events, while the packed markdown format (takes_packed.md) condenses this into human-readable phrases with speaker attribution and precise timing brackets. The packed format is optimized for LLM context windows and manual review, whereas the JSON preserves raw Scribe output for archival purposes.

Why does video-use use a packed format instead of raw JSON for LLM consumption?

The packed format reduces token count by grouping words into phrases while maintaining essential timing metadata. An hour of content typically fits into ~12 KB of markdown, compared to significantly larger JSON representations. This compression allows the LLM to process longer video segments within context limits while still accessing precise cut points through the timestamped phrase boundaries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →