# ElevenLabs Scribe Transcript: Word-Level Timestamps and Speaker Diarization Explained

> Explore ElevenLabs Scribe transcript features: word-level timestamps, speaker diarization, and audio event tagging for browser-use/video-use. Understand your video content better.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: deep-dive
- Published: 2026-07-01

---

**The ElevenLabs Scribe transcript provides word-level timestamps, speaker diarization labels, and audio event tagging when processing video content in the browser-use/video-use repository.**

When building automated video editing pipelines, having precise textual metadata eliminates the need to process raw audio or video frames. The `video-use` project leverages the ElevenLabs Scribe API to generate rich transcripts that enable frame-accurate editing decisions by Large Language Models (LLMs). This transcript data includes timing information down to individual words, speaker identification, and non-speech audio detection.

## Core Data Provided by ElevenLabs Scribe

The Scribe integration in `video-use` requests three specific data types by configuring the API call parameters in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py).

### Word-Level Timestamps

To obtain precise timing data, the `call_scribe` function (lines 65-69) sends `timestamps_granularity: "word"` to the API. This returns **start and end times** (in seconds) for every spoken word, allowing the editing pipeline to snap cuts to exact word boundaries rather than arbitrary time intervals.

### Speaker Diarization

The transcript includes **speaker identification** when `diarize: "true"` is passed in the same API request. Each word entry contains a `speaker` field (e.g., `S0`, `S1`) indicating who spoke at that moment. This enables the LLM to understand conversation flow and attribute dialogue correctly when generating editing instructions.

### Audio Event Tagging

The Scribe API also identifies non-speech sounds when `tag_audio_events: "true"` is enabled. These entries appear with `type: "audio_event"` and include descriptors like `(laughter)`, `(applause)`, or `(sigh)`, providing context about ambient audio that might influence editing decisions.

## JSON Schema and Response Structure

Scribe writes the raw response to `<edit_dir>/transcripts/<video_stem>.json` with the following structure:

```json
{
  "words": [
    {
      "type": "word",
      "text": "the",
      "start": 2.52,
      "end": 2.80,
      "speaker": "S0"
    },
    {
      "type": "audio_event",
      "text": "(laughter)",
      "start": 5.10,
      "end": 5.45
    }
  ]
}

```

The `type` field distinguishes between regular words, spacing markers, and audio events. The `start` and `end` values provide sub-second precision, while the `speaker` field enables speaker-aware editing as enforced by **Rule 6** in [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md): "never cut inside a word."

## Working with Transcript Data in Python

You can extract specific data types from the generated JSON using the following patterns:

```python
import json
from pathlib import Path

# Load transcript generated by helpers/transcribe.py

transcript_path = Path("edit/transcripts/example.mp4.json")
data = json.loads(transcript_path.read_text())

# Extract word-level timestamps

words = [
    (w["text"], w["start"], w["end"])
    for w in data["words"]
    if w["type"] == "word"
]

# Map speaker segments

speaker_segments = [
    (w["speaker"], w["start"], w["end"])
    for w in data["words"]
    if w["type"] == "word" and "speaker" in w
]

# Identify audio events

audio_events = [
    (w["text"], w["start"], w["end"])
    for w in data["words"]
    if w["type"] == "audio_event"
]

```

This extraction logic mirrors the implementation in [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py), which groups words into phrases for LLM consumption, and [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py), which renders word-level labels over waveforms for visual inspection.

## Integration Points in the video-use Pipeline

The transcript data flows through several key components:

- **[`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py)**: Builds the Scribe request with `timestamps_granularity`, `diarize`, and `tag_audio_events` parameters, then writes the raw JSON to disk.
- **[`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py)**: Processes the `words` array to generate [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md), a compact representation consumed by the LLM.
- **[`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py)**: Uses timestamp data to render word boundaries and silence shading for manual review.
- **[`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md)**: References word-level timestamps in **Rules 6 and 7** to enforce cuts that respect natural speech boundaries.

## Summary

- **Word-level timestamps** (`start`/`end` in seconds) enable frame-accurate editing by identifying exact word boundaries.
- **Speaker diarization** (`speaker` field) identifies who is speaking at any moment, allowing dialogue-aware cuts.
- **Audio event tagging** captures non-speech sounds like laughter or applause for contextual editing.
- The transcript JSON is consumed by [`pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/pack_transcripts.py) and [`timeline_view.py`](https://github.com/browser-use/video-use/blob/main/timeline_view.py) to create LLM-readable editing instructions and visual timelines.

## Frequently Asked Questions

### What is speaker diarization in ElevenLabs Scribe?

Speaker diarization is the process of identifying "who spoke when" in an audio recording. In the `video-use` implementation, Scribe returns a `speaker` field (such as `S0` or `S1`) for each word, enabling the system to distinguish between multiple speakers without requiring separate audio tracks.

### How accurate are the word-level timestamps?

The timestamps provided by Scribe are accurate to sub-second precision, typically aligning within milliseconds of actual speech. According to the implementation in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py), the `timestamps_granularity: "word"` parameter ensures each word entry contains precise `start` and `end` times, which the pipeline uses to enforce the "never cut inside a word" rule.

### Can Scribe detect non-speech audio events?

Yes, when `tag_audio_events: "true"` is enabled in the API request, Scribe identifies ambient sounds like laughter, applause, sighs, and background noise. These appear as entries with `type: "audio_event"` in the `words` array, providing context that helps the LLM make informed decisions about transition points and content pacing.

### How does video-use use the transcript data for editing?

The `video-use` pipeline uses the transcript as a text-only representation of the video. [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) converts the word-level data into phrases for the LLM, while [`timeline_view.py`](https://github.com/browser-use/video-use/blob/main/timeline_view.py) uses the timestamps to visualize word boundaries. The LLM then references these timestamps when generating edit commands, ensuring cuts align with natural speech patterns rather than arbitrary timecodes.