How Speaker Diarization and Audio Event Tagging Work in video-use

video-use obtains speaker diarization and audio-event tags from ElevenLabs Scribe and converts the raw JSON into compact, timestamped markdown transcripts that the LLM can reason over.

The browser-use/video-use repository implements a two-layer audio processing pipeline that transforms raw video into structured text. By leveraging the ElevenLabs Scribe API, the system identifies individual speakers and detects non-speech sounds like laughter or applause. This guide explains the mechanical implementation of speaker diarization and audio event tagging across the transcription and packaging modules.

Audio Extraction and Scribe Processing

The pipeline begins in helpers/transcribe.py, where the system prepares audio data and dispatches it to ElevenLabs' specialized transcription service.

Extracting Audio with FFmpeg

First, the extract_audio function converts the source video into a mono 16 kHz WAV file using FFmpeg. This standardization ensures consistent quality for the downstream Scribe API regardless of the input format.

Configuring the Scribe API Request

The script sends the audio to the SCRIBE_URL endpoint with a precisely structured payload. As shown in lines 64-70 of helpers/transcribe.py, the request body always includes:

{
    "model_id": "scribe_v1",
    "diarize": "true",
    "tag_audio_events": "true",
    "timestamps_granularity": "word"
}

Optional CLI arguments can append language_code and num_speakers parameters to improve diarization accuracy for specific content.

The Scribe Response Format

Scribe returns a JSON object containing a words array that supports three distinct entry types. This granular format enables precise timeline reconstruction.

Word-Level Timestamps and Speaker IDs

Entries with type: "word" contain start and end timestamps, the spoken text, and a speaker_id field (e.g., "speaker_0"). The presence of speaker_id implements the speaker diarization feature, tagging every token with its originating voice.

Audio Events and Silence Markers

The response also includes audio_event entries for non-speech sounds and spacing entries for silence gaps. Audio events carry the raw label (such as laughter or applause) in the text field, while spacing entries indicate the duration of pauses between utterances.

Packaging Diarization and Audio Event Data for LLM Consumption

Raw JSON remains unwieldy for LLM reasoning. The helpers/pack_transcripts.py module condenses this data into takes_packed.md, a markdown file optimized for AI consumption.

Grouping Logic and Speaker Changes

The group_into_phrases function (lines 38-84) iterates through the words array and splits the transcript into phrases whenever it encounters:

  • A silence duration exceeding the silence_threshold (default 0.5 seconds)
  • A change in speaker_id

This logic ensures that each line in the final output represents a continuous utterance from a single speaker.

Rendering Audio Events and Speaker Tags

While building phrases, the script handles entry type differentiation. Word tokens pass through verbatim, while audio_event entries undergo transformation at lines 64-67:

if t == "audio_event":
    # Convert to parenthesized format

    text = f"({text})"

The render_markdown function (lines 50-61) formats each phrase as:


[timestamp_start-timestamp_end] S{speaker_id} {text}

Speaker IDs map to sequential tags like S0, S1, etc., providing clear attribution without exposing raw API identifiers.

Output Format

The resulting edit/takes_packed.md contains lines such as:


[000.00-003.21] S0 Welcome to the demo.
[003.21-005.78] S1 (laughter) Thanks for watching.

This format embeds audio event tags directly within the speaker text while maintaining strict temporal boundaries.

Implementation Examples

To process a video with full diarization and event tagging, run the transcription helper:

python helpers/transcribe.py path/to/video.mp4 \
    --edit-dir ./edit \
    --num-speakers 2

This command writes edit/transcripts/video.json containing the complete Scribe payload with per-word speaker assignments.

Next, package the transcripts for LLM analysis:

python helpers/pack_transcripts.py --edit-dir ./edit

The script generates edit/takes_packed.md with speaker-tagged phrases and parenthesized audio events.

To inspect the raw diarization data programmatically:

import json
import pathlib

data = json.loads(pathlib.Path("edit/transcripts/video.json").read_text())
for w in data["words"]:
    if w["type"] == "audio_event":
        print(f"Event {w['text']} at {w['start']:.2f}s")
    elif w["type"] == "word":
        print(f"Speaker {w['speaker_id']} said: {w['text']}")

Summary

  • Audio extraction in helpers/transcribe.py prepares mono 16 kHz WAV files using FFmpeg.
  • The ElevenLabs Scribe API performs speaker diarization via the diarize: true flag and audio event tagging via tag_audio_events: true.
  • Scribe returns word-level JSON where speaker_id identifies voices and audio_event entries capture non-speech sounds.
  • helpers/pack_transcripts.py condenses raw JSON into takes_packed.md using group_into_phrases, splitting on speaker changes or 0.5-second silences.
  • Final output uses S0, S1 speaker tags and parenthesized event labels like (laughter) for immediate LLM comprehension.

Frequently Asked Questions

How does video-use handle speaker identification accuracy?

The system passes an optional num_speakers parameter to the ElevenLabs Scribe API, which improves diarization accuracy by constraining the model's speaker count expectations. According to helpers/transcribe.py, this value comes from CLI arguments and integrates directly into the Scribe request payload alongside the mandatory diarize: true setting.

What types of audio events does the system detect?

The Scribe model identifies non-speech sounds such as laughter, applause, sighs, and background noise. These events appear as audio_event entries in the JSON response with descriptive labels (e.g., laughter). The helpers/pack_transcripts.py script converts these into parenthesized tags embedded directly within the speaker's text block.

Can I adjust the silence threshold for phrase segmentation?

Yes. The group_into_phrases function in helpers/pack_transcripts.py uses a configurable silence_threshold parameter that defaults to 0.5 seconds. Increasing this value creates longer phrases by allowing longer pauses within a single utterance, while decreasing it splits speech more aggressively at brief hesitations.

Where does the system store the raw diarization data?

Raw JSON output from ElevenLabs Scribe persists in edit/transcripts/video.json (or equivalent filenames based on input). This file retains per-word timestamps, speaker IDs, and audio event markers. The packed markdown version at edit/takes_packed.md provides a human-readable abstraction, while the JSON serves visualization tools like helpers/timeline_view.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →