How to Batch Process Multiple Video Files Using the video-use Pipeline

Run helpers/transcribe_batch.py to parallelize transcription across multiple worker threads, then execute helpers/pack_transcripts.py to consolidate outputs into takes_packed.md, enabling the LLM-driven editing pipeline to process entire directories automatically.

The video-use repository by browser-use implements an AI-powered video editing workflow that converts raw footage into polished edits through transcription, packing, reasoning, and rendering stages. When you need to batch process multiple video files, the architecture allows you to parallelize only the transcription step while the remaining pipeline stages operate sequentially on the consolidated transcript. This design leverages ThreadPoolExecutor for concurrent ElevenLabs Scribe uploads and automatically handles the downstream edit generation.

Architecture of the Batch Pipeline

The batch workflow follows a seven-stage pipeline where only the initial discovery and transcription steps require explicit parallelization. Once transcripts exist, the system treats the packed output as a single source for editing.

Step 1: Video Discovery

The find_videos function in helpers/transcribe_batch.py recursively scans the supplied directory for common video extensions including .mp4, .mov, and .mkv. This function filters the directory contents and returns a list of video paths ready for processing.

Step 2: Parallel Transcription

Each discovered file is handed to transcribe_one (defined in helpers/transcribe.py) inside a ThreadPoolExecutor block within the main() function of helpers/transcribe_batch.py. The script checks for existing transcripts to avoid duplicate API calls, then extracts audio and uploads to ElevenLabs Scribe. JSON responses are saved to <videos_dir>/edit/transcripts/<stem>.json, preserving the original filename stem for traceability.

Step 3: Packing Transcripts

After all videos finish transcribing, helpers/pack_transcripts.py reads every *.json file in edit/transcripts/ and concatenates them into takes_packed.md. This markdown file contains word-level timestamps and speaker diarization data formatted for LLM consumption, placed alongside the source files in the edit directory.

Step 4: Automated Editing and Rendering

The LLM reads takes_packed.md from the pack step and proposes cuts, generating an EDL stored in edit/edl.md. The helpers/render.py script consumes this EDL to produce edit/final.mp4. Finally, helpers/timeline_view.py generates PNG previews of cut boundaries for self-evaluation, allowing up to three re-render attempts if the LLM detects issues with the output.

Running the Batch Transcription Script

Execute the batch processor with default settings (4 workers) by pointing it at your video directory:

python helpers/transcribe_batch.py /path/to/my/videos

This command scans the directory, transcribes any missing files using ElevenLabs Scribe, and writes individual JSON transcripts to edit/transcripts/.

After transcription completes, pack the results for the LLM:

python helpers/pack_transcripts.py /path/to/my/videos

With takes_packed.md generated, invoke your agent (Claude Code, Codex, or similar) with an instruction like > edit these into a launch video. The skill automatically handles EDL creation, rendering through helpers/render.py, and quality validation via helpers/timeline_view.py.

Configuring Parallel Processing Options

Control concurrency, language detection, and output locations using command-line flags:

python helpers/transcribe_batch.py /path/to/my/videos \
    --workers 8 \
    --language en \
    --num-speakers 2 \
    --edit-dir /custom/output/edit
  • --workers: Sets the ThreadPoolExecutor worker count (default: 4). Increase for faster processing of large batches, respecting ElevenLabs API rate limits.
  • --language: Forces a specific language code (e.g., en, es) instead of auto-detection.
  • --num-speakers: Improves speaker diarization accuracy when you know the exact speaker count.
  • --edit-dir: Redirects transcript output to a non-standard path, useful for segregating different project versions.

The pack_transcripts.py script is idempotent—you can re-run it after adding new videos to the directory without affecting existing packed data.

Summary

  • Parallelize transcription only: The helpers/transcribe_batch.py script handles concurrent uploads via ThreadPoolExecutor, while downstream stages process the consolidated takes_packed.md automatically.
  • Store transcripts correctly: Individual JSON files land in edit/transcripts/ with stems matching the source videos, as implemented in helpers/transcribe.py.
  • Pack before editing: Always run helpers/pack_transcripts.py after the batch finishes to create the LLM-ready markdown file.
  • Configure workers: Use the --workers flag to scale concurrency based on your API limits and local CPU resources.
  • Automated quality control: The pipeline includes self-evaluation via helpers/timeline_view.py with a maximum of three render attempts.

Frequently Asked Questions

How does the batch script avoid re-transcribing existing files?

The transcribe_one function in helpers/transcribe.py checks for the existence of a JSON file at transcripts_dir / f"{video.stem}.json" before initiating the ElevenLabs Scribe API call. If the transcript exists, the function skips that video, making the batch process idempotent and safe to re-run on directories containing mixed processed and unprocessed footage.

Can I process videos with different languages in the same batch?

Yes. While you can force a specific language using the --language flag (e.g., --language en), omitting this parameter allows ElevenLabs Scribe to auto-detect the language for each file individually. The transcribe_batch.py script passes each video to transcribe_one independently, so mixed-language batches process correctly without manual segmentation.

What happens if one video fails during parallel transcription?

The ThreadPoolExecutor block in main() of helpers/transcribe_batch.py processes each video in an isolated thread. If transcribe_one encounters an error (network failure, corrupt file, or API error), that specific task fails independently without stopping the entire batch. Successful transcripts are written to edit/transcripts/, while failed files can be re-processed by running the script again after resolving the underlying issue.

How do I adjust the number of concurrent transcriptions?

Modify the --workers parameter when invoking helpers/transcribe_batch.py. The default value of 4 balances throughput with ElevenLabs API rate limits, but you can increase this to 8 or higher for large local batches or decrease it to 1 or 2 if encountering rate limit errors. The worker count directly controls the ThreadPoolExecutor max_workers parameter in the source code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →