# How Speaker Diarization and Audio Event Tagging Work in video-use

> Learn how video-use leverages ElevenLabs to perform speaker diarization and audio event tagging, transforming raw JSON into easy-to-use markdown transcripts for LLM reasoning.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: deep-dive
- Published: 2026-06-29

---

**video-use obtains speaker diarization and audio-event tags from ElevenLabs Scribe and converts the raw JSON into compact, timestamped markdown transcripts that the LLM can reason over.**

The `browser-use/video-use` repository implements a two-layer audio processing pipeline that transforms raw video into structured text. By leveraging the ElevenLabs Scribe API, the system identifies individual speakers and detects non-speech sounds like laughter or applause. This guide explains the mechanical implementation of **speaker diarization** and **audio event tagging** across the transcription and packaging modules.

## Audio Extraction and Scribe Processing

The pipeline begins in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py), where the system prepares audio data and dispatches it to ElevenLabs' specialized transcription service.

### Extracting Audio with FFmpeg

First, the `extract_audio` function converts the source video into a mono 16 kHz WAV file using FFmpeg. This standardization ensures consistent quality for the downstream Scribe API regardless of the input format.

### Configuring the Scribe API Request

The script sends the audio to the `SCRIBE_URL` endpoint with a precisely structured payload. As shown in lines 64-70 of [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py), the request body always includes:

```python
{
    "model_id": "scribe_v1",
    "diarize": "true",
    "tag_audio_events": "true",
    "timestamps_granularity": "word"
}

```

Optional CLI arguments can append `language_code` and `num_speakers` parameters to improve diarization accuracy for specific content.

## The Scribe Response Format

Scribe returns a JSON object containing a `words` array that supports three distinct entry types. This granular format enables precise timeline reconstruction.

### Word-Level Timestamps and Speaker IDs

Entries with `type: "word"` contain `start` and `end` timestamps, the spoken `text`, and a `speaker_id` field (e.g., `"speaker_0"`). The presence of `speaker_id` implements the **speaker diarization** feature, tagging every token with its originating voice.

### Audio Events and Silence Markers

The response also includes `audio_event` entries for non-speech sounds and `spacing` entries for silence gaps. Audio events carry the raw label (such as `laughter` or `applause`) in the `text` field, while spacing entries indicate the duration of pauses between utterances.

## Packaging Diarization and Audio Event Data for LLM Consumption

Raw JSON remains unwieldy for LLM reasoning. The [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) module condenses this data into [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md), a markdown file optimized for AI consumption.

### Grouping Logic and Speaker Changes

The `group_into_phrases` function (lines 38-84) iterates through the `words` array and splits the transcript into phrases whenever it encounters:

- A silence duration exceeding the `silence_threshold` (default **0.5 seconds**)
- A change in `speaker_id`

This logic ensures that each line in the final output represents a continuous utterance from a single speaker.

### Rendering Audio Events and Speaker Tags

While building phrases, the script handles entry type differentiation. Word tokens pass through verbatim, while `audio_event` entries undergo transformation at lines 64-67:

```python
if t == "audio_event":
    # Convert to parenthesized format

    text = f"({text})"

```

The `render_markdown` function (lines 50-61) formats each phrase as:

```

[timestamp_start-timestamp_end] S{speaker_id} {text}

```

Speaker IDs map to sequential tags like **S0**, **S1**, etc., providing clear attribution without exposing raw API identifiers.

### Output Format

The resulting [`edit/takes_packed.md`](https://github.com/browser-use/video-use/blob/main/edit/takes_packed.md) contains lines such as:

```

[000.00-003.21] S0 Welcome to the demo.
[003.21-005.78] S1 (laughter) Thanks for watching.

```

This format embeds **audio event** tags directly within the speaker text while maintaining strict temporal boundaries.

## Implementation Examples

To process a video with full diarization and event tagging, run the transcription helper:

```bash
python helpers/transcribe.py path/to/video.mp4 \
    --edit-dir ./edit \
    --num-speakers 2

```

This command writes [`edit/transcripts/video.json`](https://github.com/browser-use/video-use/blob/main/edit/transcripts/video.json) containing the complete Scribe payload with per-word speaker assignments.

Next, package the transcripts for LLM analysis:

```bash
python helpers/pack_transcripts.py --edit-dir ./edit

```

The script generates [`edit/takes_packed.md`](https://github.com/browser-use/video-use/blob/main/edit/takes_packed.md) with speaker-tagged phrases and parenthesized audio events.

To inspect the raw diarization data programmatically:

```python
import json
import pathlib

data = json.loads(pathlib.Path("edit/transcripts/video.json").read_text())
for w in data["words"]:
    if w["type"] == "audio_event":
        print(f"Event {w['text']} at {w['start']:.2f}s")
    elif w["type"] == "word":
        print(f"Speaker {w['speaker_id']} said: {w['text']}")

```

## Summary

- **Audio extraction** in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) prepares mono 16 kHz WAV files using FFmpeg.
- The ElevenLabs Scribe API performs **speaker diarization** via the `diarize: true` flag and **audio event tagging** via `tag_audio_events: true`.
- Scribe returns word-level JSON where `speaker_id` identifies voices and `audio_event` entries capture non-speech sounds.
- [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) condenses raw JSON into [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md) using `group_into_phrases`, splitting on speaker changes or 0.5-second silences.
- Final output uses `S0`, `S1` speaker tags and parenthesized event labels like `(laughter)` for immediate LLM comprehension.

## Frequently Asked Questions

### How does video-use handle speaker identification accuracy?

The system passes an optional `num_speakers` parameter to the ElevenLabs Scribe API, which improves diarization accuracy by constraining the model's speaker count expectations. According to [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py), this value comes from CLI arguments and integrates directly into the Scribe request payload alongside the mandatory `diarize: true` setting.

### What types of audio events does the system detect?

The Scribe model identifies non-speech sounds such as laughter, applause, sighs, and background noise. These events appear as `audio_event` entries in the JSON response with descriptive labels (e.g., `laughter`). The [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) script converts these into parenthesized tags embedded directly within the speaker's text block.

### Can I adjust the silence threshold for phrase segmentation?

Yes. The `group_into_phrases` function in [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) uses a configurable `silence_threshold` parameter that defaults to 0.5 seconds. Increasing this value creates longer phrases by allowing longer pauses within a single utterance, while decreasing it splits speech more aggressively at brief hesitations.

### Where does the system store the raw diarization data?

Raw JSON output from ElevenLabs Scribe persists in [`edit/transcripts/video.json`](https://github.com/browser-use/video-use/blob/main/edit/transcripts/video.json) (or equivalent filenames based on input). This file retains per-word timestamps, speaker IDs, and audio event markers. The packed markdown version at [`edit/takes_packed.md`](https://github.com/browser-use/video-use/blob/main/edit/takes_packed.md) provides a human-readable abstraction, while the JSON serves visualization tools like [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py).