How Video-Use Implements Transcript Caching to Avoid Redundant ElevenLabs API Calls
Video-Use stores the complete JSON result returned by ElevenLabs Scribe on disk and reuses it on subsequent runs to eliminate unnecessary API calls.
The browser-use/video-use repository provides a transcription pipeline that integrates with ElevenLabs Scribe. To prevent wasting API credits and reduce processing time, the implementers designed a file-based transcript caching strategy that persists full API responses and verifies their existence before initiating any network requests or audio extraction.
Single-File Transcript Caching in helpers/transcribe.py
For individual video processing, the caching logic lives in helpers/transcribe.py. The system writes the complete JSON response from ElevenLabs to disk and verifies its existence before extracting audio or uploading files.
Cache Location and Naming Convention
Transcripts are stored at <edit_dir>/transcripts/<video_stem>.json, where video_stem represents the source video's base filename without extension. This predictable path structure allows the system to locate existing transcripts instantly without querying an external database.
The Cache Guard Implementation
The caching guard operates at lines 100-108 of helpers/transcribe.py. Before processing begins, the script checks out_path.exists(). If the file is present, the function prints cached: <name>.json and returns the existing path immediately, bypassing the entire upload and transcription workflow.
# Conceptual implementation based on the source
if out_path.exists():
print(f"cached: {out_path.name}")
return out_path
This guard executes before any audio extraction or network activity, ensuring zero redundant API calls for previously processed videos.
Batch Processing with Pre-Filtering
The batch transcription runner in helpers/transcribe_batch.py extends the caching strategy to parallel processing workflows. Rather than submitting all videos to the worker pool indiscriminately, it pre-filters the workload based on cache state.
Filtering Already Cached Videos
Around lines 72-74 of helpers/transcribe_batch.py, the script constructs a list named already_cached containing videos that have existing transcript files under <edit_dir>/transcripts/<stem>.json. Only videos absent from this list proceed to the transcription workers.
Worker Pool Optimization
By excluding cached files before creating the worker pool, the system ensures that parallel workers process only new content. This prevents duplicate uploads and maintains efficient resource utilization across the batch.
# First run processes all videos
python helpers/transcribe_batch.py videos/
# found 10 videos (0 cached, 10 to transcribe)
# transcribing 10 files with 4 parallel workers
# Subsequent run skips cached files
python helpers/transcribe_batch.py videos/
# found 10 videos (4 cached, 6 to transcribe)
# transcribing 6 files with 4 parallel workers
Practical Usage Examples
Processing a Single Video
When running helpers/transcribe.py against the same file twice, the second execution returns immediately from cache:
# First run - performs extraction and API call
python helpers/transcribe.py path/to/video.mp4
# → extracting audio … uploading … saved: video.json
# Second run - uses cached transcript
python helpers/transcribe.py path/to/video.mp4
# → cached: video.json
Programmatic Access
You can integrate the caching logic directly into Python scripts:
from helpers.transcribe import transcribe_one, load_api_key
from pathlib import Path
video = Path("videos/intro.mp4")
edit_dir = Path("videos/edit")
api_key = load_api_key()
# First call performs upload; second call returns cached path
transcript_path = transcribe_one(video, edit_dir, api_key)
transcript_path = transcribe_one(video, edit_dir, api_key) # Cached
Environment Configuration
The caching system relies on the ELEVENLABS_API_KEY defined in your environment, as shown in .env.example. Without valid credentials, the system cannot generate new transcripts, but it will still serve cached results if they exist.
Summary
- Video-Use implements a file-based transcript caching strategy to avoid redundant ElevenLabs API calls by persisting complete JSON responses to disk.
- Single-file processing in
helpers/transcribe.pychecks<edit_dir>/transcripts/<video_stem>.jsonfor existence before extracting audio or uploading, returning the cached path immediately if found. - Batch processing in
helpers/transcribe_batch.pypre-filters videos using analready_cachedlist, ensuring only uncached videos enter the parallel worker pool. - Both implementations check cache status before any network activity, guaranteeing each video transmits to ElevenLabs at most once while saving API credits and processing time.
Frequently Asked Questions
Where does Video-Use store cached transcripts?
Video-Use stores cached transcripts in <edit_dir>/transcripts/<video_stem>.json, where edit_dir represents your configured editing directory and video_stem is the source video's filename without extension. This JSON file contains the complete ElevenLabs Scribe API response.
How does the batch processor avoid re-transcribing videos?
The batch processor in helpers/transcribe_batch.py builds an already_cached list by checking for existing transcript files before creating the worker pool. Only videos not present in this list proceed to transcription, effectively filtering duplicates before any API calls occur.
Does the cache check happen before or after audio extraction?
The cache check happens before audio extraction. In helpers/transcribe.py, the existence of the output JSON is verified at lines 100-108 prior to any audio processing or network upload, ensuring the system skips expensive operations entirely for cached content.
Is the caching mechanism configurable or automatic?
The caching mechanism operates automatically without configuration options. The system always checks for existing transcripts at the defined path and returns them if present. There is no flag to disable caching; to force re-transcription, you must manually delete the corresponding JSON file from the transcripts directory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →