What Audio Transcription Service Does video-use Use? ElevenLabs Scribe Integration Guide

TLDR: The video-use repository uses ElevenLabs Scribe as its audio transcription service, sending extracted audio to https://api.elevenlabs.io/v1/speech-to-text via the helpers/transcribe.py module with the scribe_v1 model.

The browser-use/video-use project automates video processing workflows requiring accurate speech-to-text conversion. For this core functionality, the codebase integrates exclusively with ElevenLabs Scribe, implementing the API calls in dedicated helper modules that handle authentication, audio extraction, and speaker diarization.

How video-use Implements ElevenLabs Scribe

Core Transcription Module

In helpers/transcribe.py, the transcription service implementation defines SCRIBE_URL pointing to https://api.elevenlabs.io/v1/speech-to-text. The transcribe_one() function extracts audio from video files into WAV format and transmits it to this endpoint. The request payload specifies "scribe_v1" as the model identifier and includes diarization parameters to support multi-speaker recognition.

API Authentication

The module requires an ELEVENLABS_API_KEY environment variable, documented in .env.example. The load_api_key() function retrieves this credential from environment variables or .env files, ensuring secure authentication with the ElevenLabs Scribe API without exposing keys in the source code.

Transcribing Videos with ElevenLabs Scribe

Single Video Transcription

Process individual video files using the command-line interface or Python API.

From the terminal:

python helpers/transcribe.py path/to/video.mp4 \
    --edit-dir path/to/edit \
    --language en \
    --num-speakers 2

Programmatically:

from pathlib import Path
from helpers.transcribe import load_api_key, transcribe_one

video_path = Path("video.mp4")
edit_dir = Path("edit")
api_key = load_api_key()

transcript_path = transcribe_one(
    video=video_path,
    edit_dir=edit_dir,
    api_key=api_key,
    language="en",
    num_speakers=2,
    verbose=True
)

print(f"Transcript saved to: {transcript_path}")

Batch Processing

For high-volume workflows, helpers/transcribe_batch.py provides parallel processing across multiple workers.

python helpers/transcribe_batch.py /path/to/videos \
    --workers 4 \
    --language en \
    --num-speakers 2

This script distributes videos across the specified number of workers, each invoking the ElevenLabs Scribe API independently to maximize throughput.

Configuration Options

The ElevenLabs Scribe integration supports several parameters:

  • language: ISO-639-1 code (e.g., "en", "es") specifying the spoken language
  • num_speakers: Integer enabling speaker diarization by indicating the expected number of distinct speakers in the audio
  • verbose: Boolean flag controlling debug output during processing

These options map directly to the JSON payload sent to the ElevenLabs API endpoint.

Summary

  • video-use employs ElevenLabs Scribe as its exclusive audio transcription service via the https://api.elevenlabs.io/v1/speech-to-text endpoint.
  • The implementation resides in helpers/transcribe.py, using the scribe_v1 model with support for speaker diarization.
  • Authentication requires the ELEVENLABS_API_KEY environment variable, loaded through the load_api_key() helper.
  • Both single-file (transcribe_one) and batch processing (transcribe_batch.py) workflows are supported.
  • Configuration options include language specification, speaker count, and verbose logging.

Frequently Asked Questions

What specific ElevenLabs Scribe model does video-use use?

According to the source code in helpers/transcribe.py, video-use explicitly sends "scribe_v1" as the model identifier in the API request payload. This targets the specific Scribe model version optimized for accurate transcription and speaker diarization.

How does video-use handle ElevenLabs API authentication?

The repository uses the load_api_key() function to retrieve the ELEVENLABS_API_KEY from environment variables or .env files. This approach keeps credentials out of version control while making them available to the transcribe_one() function for authenticated API calls.

Can video-use distinguish between multiple speakers in a video?

Yes. By specifying the num_speakers parameter (either via command line or the Python API), the code enables speaker diarization in the ElevenLabs Scribe request. This returns transcripts segmented by individual speakers rather than a single continuous text block.

Is there a way to process multiple videos simultaneously?

Yes. The helpers/transcribe_batch.py script supports parallel processing using the --workers flag. This distributes transcription jobs across multiple processes, each making independent calls to the ElevenLabs Scribe API to maximize throughput.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →