# How Speaker Diarization Works with the ElevenLabs Scribe API in video-use

> Learn how speaker diarization works with the ElevenLabs Scribe API in video-use. Extract audio, send to ElevenLabs with diarize true, and get speaker tagged transcripts.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: deep-dive
- Published: 2026-07-03

---

**Speaker diarization in the video-use library works by extracting mono 16kHz audio from video files and sending it to the ElevenLabs Scribe API with the `diarize: "true"` parameter, which returns a JSON transcript where each word is tagged with a speaker identifier like `spk_0`.**

The browser-use/video-use repository delegates speaker identification entirely to ElevenLabs' cloud-based Scribe service rather than running diarization models locally. By configuring specific flags in the API request payload, the library receives rich metadata including word-level timestamps and speaker labels that map each utterance to a specific person.

## Audio Extraction for Diarization

Before sending data to the API, `extract_audio()` in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) (lines 49-55) uses **ffmpeg** to convert the source video into a mono 16kHz WAV file. This specific format ensures compatibility with the ElevenLabs Scribe model and maintains consistent quality for speaker separation algorithms.

```bash

# ffmpeg command extracts single-channel audio at 16kHz

ffmpeg -i input.mp4 -ar 16000 -ac 1 output.wav

```

The function returns a file path to this processed audio, which serves as the multipart payload for the subsequent API call.

## Enabling Speaker Diarization in the API Request

The `call_scribe()` function constructs the HTTP POST payload that activates diarization. In [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) (lines 64-69), the code builds a dictionary with specific flags required by the ElevenLabs endpoint:

```python
data = {
    "model_id": "scribe_v1",
    "diarize": "true",               # Enables speaker diarization

    "tag_audio_events": "true",      # Includes non-speech audio timestamps

    "timestamps_granularity": "word",# Provides word-level precision

}

```

When the number of speakers is known beforehand, you can optionally include `"num_speakers"` in this payload to improve diarization accuracy. The function sends this data along with the audio file to `https://api.elevenlabs.io/v1/speech-to-text` using the `XI-API-Key` header for authentication.

## Processing the Diarized Response

The Scribe API returns a JSON payload containing a top-level `words` array. As implemented in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) (lines 75-84), each element in this array includes the spoken text, start/end timestamps, and a `speaker` field containing identifiers like `"spk_0"`, `"spk_1"`, etc.

```json
{
  "words": [
    {
      "text": "Hello",
      "start": 0.0,
      "end": 0.5,
      "speaker": "spk_0"
    },
    {
      "text": "world",
      "start": 0.6,
      "end": 1.0,
      "speaker": "spk_1"
    }
  ]
}

```

This structure allows downstream applications to attribute specific phrases to individual speakers and generate per-speaker transcripts or visual cues.

## Saving and Storing Diarization Data

After receiving the API response, `transcribe_one()` persists the complete JSON output to disk. According to [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) (lines 22-33 and 119-131), the file is written to `<edit_dir>/transcripts/<video_stem>.json`, preserving the original `words` array with speaker labels intact.

This storage approach keeps the rich diarization metadata available for subsequent processing steps, such as generating speaker-specific subtitles or analyzing conversation patterns.

## Batch Processing with Diarization

For processing multiple videos, [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) automates the workflow across an entire directory. The script invokes the same transcription pipeline for each video file, automatically applying the diarization flags and saving individual JSON transcripts for every source video.

## Implementation Examples

### Command-Line Usage

Transcribe a single video with default diarization settings:

```bash

# Ensure ELEVENLABS_API_KEY is set in your environment or .env file

python helpers/transcribe.py path/to/video.mp4

```

Specify language and expected speaker count for improved accuracy:

```bash
python helpers/transcribe.py path/to/video.mp4 \
    --language en \
    --num-speakers 2

```

Process an entire directory:

```bash
python helpers/transcribe_batch.py videos/

```

### Python API Integration

Call the transcription helper directly from your Python code:

```python
from pathlib import Path
from helpers.transcribe import transcribe_one, load_api_key

video_path = Path("example.mp4")
edit_dir = Path("example_edit")
api_key = load_api_key()  # Reads from .env or environment variable

# Perform transcription with speaker diarization

transcript_path = transcribe_one(
    video=video_path,
    edit_dir=edit_dir,
    api_key=api_key,
    language="en",
    num_speakers=2,  # Optional: improves diarization accuracy

)

print(f"Diarized transcript saved to {transcript_path}")

```

## Summary

- **Audio preprocessing**: `extract_audio()` in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) converts video to mono 16kHz WAV using ffmpeg to meet ElevenLabs API requirements.
- **Diarization activation**: The `diarize: "true"` flag in the request payload enables speaker identification in the Scribe API response.
- **Word-level metadata**: Each entry in the returned `words` array contains a `speaker` field (e.g., `spk_0`) alongside precise timestamps.
- **Persistent storage**: Complete JSON responses including speaker labels are saved to `<edit_dir>/transcripts/<video_stem>.json`.
- **Batch support**: [`transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/transcribe_batch.py) applies the same diarization workflow across multiple video files automatically.

## Frequently Asked Questions

### Does video-use perform speaker diarization locally?

No. According to the source code in `browser-use/video-use`, speaker diarization is delegated entirely to the ElevenLabs Scribe cloud API. The local code only handles audio extraction, request formatting, and result storage.

### What audio format is required for the ElevenLabs Scribe API?

The API requires a mono (single-channel) WAV file with a 16kHz sample rate. The `extract_audio()` function in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) (lines 49-55) automatically converts source videos to this format using ffmpeg before sending the data to ElevenLabs.

### How are speakers identified in the output JSON?

Speakers are labeled with identifiers like `spk_0`, `spk_1`, etc. Each word object in the `words` array contains a `speaker` field indicating which person uttered that specific word. The API determines these labels automatically, though accuracy improves when you specify the expected `num_speakers` in the request.

### Can I improve diarization accuracy by specifying the number of speakers?

Yes. When calling `transcribe_one()` or using the CLI with `--num-speakers`, the value is passed to the ElevenLabs API to help the diarization model better distinguish between voices. While optional, providing this parameter often produces more accurate speaker separation, especially in recordings with overlapping speech.