How Speaker Diarization Works with the ElevenLabs Scribe API in video-use
Speaker diarization in the video-use library works by extracting mono 16kHz audio from video files and sending it to the ElevenLabs Scribe API with the diarize: "true" parameter, which returns a JSON transcript where each word is tagged with a speaker identifier like spk_0.
The browser-use/video-use repository delegates speaker identification entirely to ElevenLabs' cloud-based Scribe service rather than running diarization models locally. By configuring specific flags in the API request payload, the library receives rich metadata including word-level timestamps and speaker labels that map each utterance to a specific person.
Audio Extraction for Diarization
Before sending data to the API, extract_audio() in helpers/transcribe.py (lines 49-55) uses ffmpeg to convert the source video into a mono 16kHz WAV file. This specific format ensures compatibility with the ElevenLabs Scribe model and maintains consistent quality for speaker separation algorithms.
# ffmpeg command extracts single-channel audio at 16kHz
ffmpeg -i input.mp4 -ar 16000 -ac 1 output.wav
The function returns a file path to this processed audio, which serves as the multipart payload for the subsequent API call.
Enabling Speaker Diarization in the API Request
The call_scribe() function constructs the HTTP POST payload that activates diarization. In helpers/transcribe.py (lines 64-69), the code builds a dictionary with specific flags required by the ElevenLabs endpoint:
data = {
"model_id": "scribe_v1",
"diarize": "true", # Enables speaker diarization
"tag_audio_events": "true", # Includes non-speech audio timestamps
"timestamps_granularity": "word",# Provides word-level precision
}
When the number of speakers is known beforehand, you can optionally include "num_speakers" in this payload to improve diarization accuracy. The function sends this data along with the audio file to https://api.elevenlabs.io/v1/speech-to-text using the XI-API-Key header for authentication.
Processing the Diarized Response
The Scribe API returns a JSON payload containing a top-level words array. As implemented in helpers/transcribe.py (lines 75-84), each element in this array includes the spoken text, start/end timestamps, and a speaker field containing identifiers like "spk_0", "spk_1", etc.
{
"words": [
{
"text": "Hello",
"start": 0.0,
"end": 0.5,
"speaker": "spk_0"
},
{
"text": "world",
"start": 0.6,
"end": 1.0,
"speaker": "spk_1"
}
]
}
This structure allows downstream applications to attribute specific phrases to individual speakers and generate per-speaker transcripts or visual cues.
Saving and Storing Diarization Data
After receiving the API response, transcribe_one() persists the complete JSON output to disk. According to helpers/transcribe.py (lines 22-33 and 119-131), the file is written to <edit_dir>/transcripts/<video_stem>.json, preserving the original words array with speaker labels intact.
This storage approach keeps the rich diarization metadata available for subsequent processing steps, such as generating speaker-specific subtitles or analyzing conversation patterns.
Batch Processing with Diarization
For processing multiple videos, helpers/transcribe_batch.py automates the workflow across an entire directory. The script invokes the same transcription pipeline for each video file, automatically applying the diarization flags and saving individual JSON transcripts for every source video.
Implementation Examples
Command-Line Usage
Transcribe a single video with default diarization settings:
# Ensure ELEVENLABS_API_KEY is set in your environment or .env file
python helpers/transcribe.py path/to/video.mp4
Specify language and expected speaker count for improved accuracy:
python helpers/transcribe.py path/to/video.mp4 \
--language en \
--num-speakers 2
Process an entire directory:
python helpers/transcribe_batch.py videos/
Python API Integration
Call the transcription helper directly from your Python code:
from pathlib import Path
from helpers.transcribe import transcribe_one, load_api_key
video_path = Path("example.mp4")
edit_dir = Path("example_edit")
api_key = load_api_key() # Reads from .env or environment variable
# Perform transcription with speaker diarization
transcript_path = transcribe_one(
video=video_path,
edit_dir=edit_dir,
api_key=api_key,
language="en",
num_speakers=2, # Optional: improves diarization accuracy
)
print(f"Diarized transcript saved to {transcript_path}")
Summary
- Audio preprocessing:
extract_audio()inhelpers/transcribe.pyconverts video to mono 16kHz WAV using ffmpeg to meet ElevenLabs API requirements. - Diarization activation: The
diarize: "true"flag in the request payload enables speaker identification in the Scribe API response. - Word-level metadata: Each entry in the returned
wordsarray contains aspeakerfield (e.g.,spk_0) alongside precise timestamps. - Persistent storage: Complete JSON responses including speaker labels are saved to
<edit_dir>/transcripts/<video_stem>.json. - Batch support:
transcribe_batch.pyapplies the same diarization workflow across multiple video files automatically.
Frequently Asked Questions
Does video-use perform speaker diarization locally?
No. According to the source code in browser-use/video-use, speaker diarization is delegated entirely to the ElevenLabs Scribe cloud API. The local code only handles audio extraction, request formatting, and result storage.
What audio format is required for the ElevenLabs Scribe API?
The API requires a mono (single-channel) WAV file with a 16kHz sample rate. The extract_audio() function in helpers/transcribe.py (lines 49-55) automatically converts source videos to this format using ffmpeg before sending the data to ElevenLabs.
How are speakers identified in the output JSON?
Speakers are labeled with identifiers like spk_0, spk_1, etc. Each word object in the words array contains a speaker field indicating which person uttered that specific word. The API determines these labels automatically, though accuracy improves when you specify the expected num_speakers in the request.
Can I improve diarization accuracy by specifying the number of speakers?
Yes. When calling transcribe_one() or using the CLI with --num-speakers, the value is passed to the ElevenLabs API to help the diarization model better distinguish between voices. While optional, providing this parameter often produces more accurate speaker separation, especially in recordings with overlapping speech.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →