# How Video-Use Implements Transcript Caching to Avoid Redundant ElevenLabs API Calls

> Discover how Video-Use implements transcript caching to avoid redundant ElevenLabs API calls. Learn how to store and reuse JSON results for efficient transcription.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: internals
- Published: 2026-07-03

---

**Video-Use stores the complete JSON result returned by ElevenLabs Scribe on disk and reuses it on subsequent runs to eliminate unnecessary API calls.**

The browser-use/video-use repository provides a transcription pipeline that integrates with ElevenLabs Scribe. To prevent wasting API credits and reduce processing time, the implementers designed a file-based transcript caching strategy that persists full API responses and verifies their existence before initiating any network requests or audio extraction.

## Single-File Transcript Caching in helpers/transcribe.py

For individual video processing, the caching logic lives in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py). The system writes the complete JSON response from ElevenLabs to disk and verifies its existence before extracting audio or uploading files.

### Cache Location and Naming Convention

Transcripts are stored at `<edit_dir>/transcripts/<video_stem>.json`, where `video_stem` represents the source video's base filename without extension. This predictable path structure allows the system to locate existing transcripts instantly without querying an external database.

### The Cache Guard Implementation

The caching guard operates at lines 100-108 of [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py). Before processing begins, the script checks `out_path.exists()`. If the file is present, the function prints `cached: <name>.json` and returns the existing path immediately, bypassing the entire upload and transcription workflow.

```python

# Conceptual implementation based on the source

if out_path.exists():
    print(f"cached: {out_path.name}")
    return out_path

```

This guard executes before any audio extraction or network activity, ensuring zero redundant API calls for previously processed videos.

## Batch Processing with Pre-Filtering

The batch transcription runner in [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) extends the caching strategy to parallel processing workflows. Rather than submitting all videos to the worker pool indiscriminately, it pre-filters the workload based on cache state.

### Filtering Already Cached Videos

Around lines 72-74 of [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py), the script constructs a list named `already_cached` containing videos that have existing transcript files under `<edit_dir>/transcripts/<stem>.json`. Only videos absent from this list proceed to the transcription workers.

### Worker Pool Optimization

By excluding cached files before creating the worker pool, the system ensures that parallel workers process only new content. This prevents duplicate uploads and maintains efficient resource utilization across the batch.

```bash

# First run processes all videos

python helpers/transcribe_batch.py videos/

# found 10 videos (0 cached, 10 to transcribe)

# transcribing 10 files with 4 parallel workers

# Subsequent run skips cached files

python helpers/transcribe_batch.py videos/

# found 10 videos (4 cached, 6 to transcribe)

# transcribing 6 files with 4 parallel workers

```

## Practical Usage Examples

### Processing a Single Video

When running [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) against the same file twice, the second execution returns immediately from cache:

```bash

# First run - performs extraction and API call

python helpers/transcribe.py path/to/video.mp4

# → extracting audio … uploading … saved: video.json

# Second run - uses cached transcript

python helpers/transcribe.py path/to/video.mp4

# → cached: video.json

```

### Programmatic Access

You can integrate the caching logic directly into Python scripts:

```python
from helpers.transcribe import transcribe_one, load_api_key
from pathlib import Path

video = Path("videos/intro.mp4")
edit_dir = Path("videos/edit")
api_key = load_api_key()

# First call performs upload; second call returns cached path

transcript_path = transcribe_one(video, edit_dir, api_key)
transcript_path = transcribe_one(video, edit_dir, api_key)  # Cached

```

### Environment Configuration

The caching system relies on the `ELEVENLABS_API_KEY` defined in your environment, as shown in `.env.example`. Without valid credentials, the system cannot generate new transcripts, but it will still serve cached results if they exist.

## Summary

- **Video-Use** implements a file-based transcript caching strategy to avoid redundant ElevenLabs API calls by persisting complete JSON responses to disk.
- **Single-file processing** in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) checks `<edit_dir>/transcripts/<video_stem>.json` for existence before extracting audio or uploading, returning the cached path immediately if found.
- **Batch processing** in [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) pre-filters videos using an `already_cached` list, ensuring only uncached videos enter the parallel worker pool.
- Both implementations check cache status before any network activity, guaranteeing each video transmits to ElevenLabs at most once while saving API credits and processing time.

## Frequently Asked Questions

### Where does Video-Use store cached transcripts?

Video-Use stores cached transcripts in `<edit_dir>/transcripts/<video_stem>.json`, where `edit_dir` represents your configured editing directory and `video_stem` is the source video's filename without extension. This JSON file contains the complete ElevenLabs Scribe API response.

### How does the batch processor avoid re-transcribing videos?

The batch processor in [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) builds an `already_cached` list by checking for existing transcript files before creating the worker pool. Only videos not present in this list proceed to transcription, effectively filtering duplicates before any API calls occur.

### Does the cache check happen before or after audio extraction?

The cache check happens **before** audio extraction. In [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py), the existence of the output JSON is verified at lines 100-108 prior to any audio processing or network upload, ensuring the system skips expensive operations entirely for cached content.

### Is the caching mechanism configurable or automatic?

The caching mechanism operates automatically without configuration options. The system always checks for existing transcripts at the defined path and returns them if present. There is no flag to disable caching; to force re-transcription, you must manually delete the corresponding JSON file from the transcripts directory.