What Audio Transcription Service Does video-use Use? ElevenLabs Scribe Integration Guide
TLDR: The video-use repository uses ElevenLabs Scribe as its audio transcription service, sending extracted audio to https://api.elevenlabs.io/v1/speech-to-text via the helpers/transcribe.py module with the scribe_v1 model.
The browser-use/video-use project automates video processing workflows requiring accurate speech-to-text conversion. For this core functionality, the codebase integrates exclusively with ElevenLabs Scribe, implementing the API calls in dedicated helper modules that handle authentication, audio extraction, and speaker diarization.
How video-use Implements ElevenLabs Scribe
Core Transcription Module
In helpers/transcribe.py, the transcription service implementation defines SCRIBE_URL pointing to https://api.elevenlabs.io/v1/speech-to-text. The transcribe_one() function extracts audio from video files into WAV format and transmits it to this endpoint. The request payload specifies "scribe_v1" as the model identifier and includes diarization parameters to support multi-speaker recognition.
API Authentication
The module requires an ELEVENLABS_API_KEY environment variable, documented in .env.example. The load_api_key() function retrieves this credential from environment variables or .env files, ensuring secure authentication with the ElevenLabs Scribe API without exposing keys in the source code.
Transcribing Videos with ElevenLabs Scribe
Single Video Transcription
Process individual video files using the command-line interface or Python API.
From the terminal:
python helpers/transcribe.py path/to/video.mp4 \
--edit-dir path/to/edit \
--language en \
--num-speakers 2
Programmatically:
from pathlib import Path
from helpers.transcribe import load_api_key, transcribe_one
video_path = Path("video.mp4")
edit_dir = Path("edit")
api_key = load_api_key()
transcript_path = transcribe_one(
video=video_path,
edit_dir=edit_dir,
api_key=api_key,
language="en",
num_speakers=2,
verbose=True
)
print(f"Transcript saved to: {transcript_path}")
Batch Processing
For high-volume workflows, helpers/transcribe_batch.py provides parallel processing across multiple workers.
python helpers/transcribe_batch.py /path/to/videos \
--workers 4 \
--language en \
--num-speakers 2
This script distributes videos across the specified number of workers, each invoking the ElevenLabs Scribe API independently to maximize throughput.
Configuration Options
The ElevenLabs Scribe integration supports several parameters:
- language: ISO-639-1 code (e.g.,
"en","es") specifying the spoken language - num_speakers: Integer enabling speaker diarization by indicating the expected number of distinct speakers in the audio
- verbose: Boolean flag controlling debug output during processing
These options map directly to the JSON payload sent to the ElevenLabs API endpoint.
Summary
- video-use employs ElevenLabs Scribe as its exclusive audio transcription service via the
https://api.elevenlabs.io/v1/speech-to-textendpoint. - The implementation resides in
helpers/transcribe.py, using thescribe_v1model with support for speaker diarization. - Authentication requires the
ELEVENLABS_API_KEYenvironment variable, loaded through theload_api_key()helper. - Both single-file (
transcribe_one) and batch processing (transcribe_batch.py) workflows are supported. - Configuration options include language specification, speaker count, and verbose logging.
Frequently Asked Questions
What specific ElevenLabs Scribe model does video-use use?
According to the source code in helpers/transcribe.py, video-use explicitly sends "scribe_v1" as the model identifier in the API request payload. This targets the specific Scribe model version optimized for accurate transcription and speaker diarization.
How does video-use handle ElevenLabs API authentication?
The repository uses the load_api_key() function to retrieve the ELEVENLABS_API_KEY from environment variables or .env files. This approach keeps credentials out of version control while making them available to the transcribe_one() function for authenticated API calls.
Can video-use distinguish between multiple speakers in a video?
Yes. By specifying the num_speakers parameter (either via command line or the Python API), the code enables speaker diarization in the ElevenLabs Scribe request. This returns transcripts segmented by individual speakers rather than a single continuous text block.
Is there a way to process multiple videos simultaneously?
Yes. The helpers/transcribe_batch.py script supports parallel processing using the --workers flag. This distributes transcription jobs across multiple processes, each making independent calls to the ElevenLabs Scribe API to maximize throughput.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →