ElevenLabs Scribe Transcript: Word-Level Timestamps and Speaker Diarization Explained

The ElevenLabs Scribe transcript provides word-level timestamps, speaker diarization labels, and audio event tagging when processing video content in the browser-use/video-use repository.

When building automated video editing pipelines, having precise textual metadata eliminates the need to process raw audio or video frames. The video-use project leverages the ElevenLabs Scribe API to generate rich transcripts that enable frame-accurate editing decisions by Large Language Models (LLMs). This transcript data includes timing information down to individual words, speaker identification, and non-speech audio detection.

Core Data Provided by ElevenLabs Scribe

The Scribe integration in video-use requests three specific data types by configuring the API call parameters in helpers/transcribe.py.

Word-Level Timestamps

To obtain precise timing data, the call_scribe function (lines 65-69) sends timestamps_granularity: "word" to the API. This returns start and end times (in seconds) for every spoken word, allowing the editing pipeline to snap cuts to exact word boundaries rather than arbitrary time intervals.

Speaker Diarization

The transcript includes speaker identification when diarize: "true" is passed in the same API request. Each word entry contains a speaker field (e.g., S0, S1) indicating who spoke at that moment. This enables the LLM to understand conversation flow and attribute dialogue correctly when generating editing instructions.

Audio Event Tagging

The Scribe API also identifies non-speech sounds when tag_audio_events: "true" is enabled. These entries appear with type: "audio_event" and include descriptors like (laughter), (applause), or (sigh), providing context about ambient audio that might influence editing decisions.

JSON Schema and Response Structure

Scribe writes the raw response to <edit_dir>/transcripts/<video_stem>.json with the following structure:

{
  "words": [
    {
      "type": "word",
      "text": "the",
      "start": 2.52,
      "end": 2.80,
      "speaker": "S0"
    },
    {
      "type": "audio_event",
      "text": "(laughter)",
      "start": 5.10,
      "end": 5.45
    }
  ]
}

The type field distinguishes between regular words, spacing markers, and audio events. The start and end values provide sub-second precision, while the speaker field enables speaker-aware editing as enforced by Rule 6 in SKILL.md: "never cut inside a word."

Working with Transcript Data in Python

You can extract specific data types from the generated JSON using the following patterns:

import json
from pathlib import Path

# Load transcript generated by helpers/transcribe.py

transcript_path = Path("edit/transcripts/example.mp4.json")
data = json.loads(transcript_path.read_text())

# Extract word-level timestamps

words = [
    (w["text"], w["start"], w["end"])
    for w in data["words"]
    if w["type"] == "word"
]

# Map speaker segments

speaker_segments = [
    (w["speaker"], w["start"], w["end"])
    for w in data["words"]
    if w["type"] == "word" and "speaker" in w
]

# Identify audio events

audio_events = [
    (w["text"], w["start"], w["end"])
    for w in data["words"]
    if w["type"] == "audio_event"
]

This extraction logic mirrors the implementation in helpers/pack_transcripts.py, which groups words into phrases for LLM consumption, and helpers/timeline_view.py, which renders word-level labels over waveforms for visual inspection.

Integration Points in the video-use Pipeline

The transcript data flows through several key components:

  • helpers/transcribe.py: Builds the Scribe request with timestamps_granularity, diarize, and tag_audio_events parameters, then writes the raw JSON to disk.
  • helpers/pack_transcripts.py: Processes the words array to generate takes_packed.md, a compact representation consumed by the LLM.
  • helpers/timeline_view.py: Uses timestamp data to render word boundaries and silence shading for manual review.
  • SKILL.md: References word-level timestamps in Rules 6 and 7 to enforce cuts that respect natural speech boundaries.

Summary

  • Word-level timestamps (start/end in seconds) enable frame-accurate editing by identifying exact word boundaries.
  • Speaker diarization (speaker field) identifies who is speaking at any moment, allowing dialogue-aware cuts.
  • Audio event tagging captures non-speech sounds like laughter or applause for contextual editing.
  • The transcript JSON is consumed by pack_transcripts.py and timeline_view.py to create LLM-readable editing instructions and visual timelines.

Frequently Asked Questions

What is speaker diarization in ElevenLabs Scribe?

Speaker diarization is the process of identifying "who spoke when" in an audio recording. In the video-use implementation, Scribe returns a speaker field (such as S0 or S1) for each word, enabling the system to distinguish between multiple speakers without requiring separate audio tracks.

How accurate are the word-level timestamps?

The timestamps provided by Scribe are accurate to sub-second precision, typically aligning within milliseconds of actual speech. According to the implementation in helpers/transcribe.py, the timestamps_granularity: "word" parameter ensures each word entry contains precise start and end times, which the pipeline uses to enforce the "never cut inside a word" rule.

Can Scribe detect non-speech audio events?

Yes, when tag_audio_events: "true" is enabled in the API request, Scribe identifies ambient sounds like laughter, applause, sighs, and background noise. These appear as entries with type: "audio_event" in the words array, providing context that helps the LLM make informed decisions about transition points and content pacing.

How does video-use use the transcript data for editing?

The video-use pipeline uses the transcript as a text-only representation of the video. helpers/pack_transcripts.py converts the word-level data into phrases for the LLM, while timeline_view.py uses the timestamps to visualize word boundaries. The LLM then references these timestamps when generating edit commands, ensuring cuts align with natural speech patterns rather than arbitrary timecodes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →