Read Frog Technical Approach to Real-Time YouTube Subtitle Translation
Read Frog translates YouTube subtitles in real-time by chaining four tightly-coupled subsystems—subtitle acquisition, format normalization, a sliding-window translation coordinator, and a background queue with intelligent caching—to deliver low-latency translations synchronized with video playback.
Read Frog is an open-source browser extension that performs real-time translation of YouTube video subtitles without requiring users to provide their own API keys. The technical architecture processes caption tracks as time-series data, leveraging a TranslationCoordinator that schedules jobs based on the viewer's current playback position while a background queue handles provider-specific optimizations and result caching.
The Four-Layer Translation Architecture
The codebase divides the workflow into four distinct subsystems that operate sequentially:
- Subtitle Acquisition – Detects the current video, selects the optimal caption track, and fetches raw timed-text events from YouTube's API. Implemented in
src/utils/subtitles/fetchers/youtube/index.ts. - Subtitle Normalisation – Filters noise, detects subtitle formats (karaoke, scrolling ASR, or standard), and parses events into a unified
SubtitlesFragment[]structure. Located insrc/utils/subtitles/fetchers/youtube/parser. - Translation Coordination – Maintains a sliding window around the current playback time, batches upcoming fragments, and schedules them for translation. Core logic resides in
src/entrypoints/subtitles.content/translation-coordinator.ts. - Background Translation Queue – Receives translation requests, checks the cache, and delegates to either a request-queue (single-call providers) or a batch-queue (LLM providers). Implemented in
src/entrypoints/background/translation-queues.ts.
Subtitle Acquisition and Format Parsing
When a YouTube watch page loads, initYoutubeSubtitles() in src/entrypoints/subtitles.content/init-youtube-subtitles.ts mounts the UI and instantiates a YoutubeSubtitlesFetcher.
The fetcher executes the following chain:
- Video Detection –
getYoutubeVideoId()extracts the video identifier, whilewaitForPlayerStateensures the player is ready. - Track Selection –
selectTrackprioritizes captions in this order: user-selected → human-uploaded → auto-generated. - URL Construction –
buildSubtitleUrlgenerates the authenticated timed-text endpoint using player data including the POT token. - Error Handling – HTTP 403/404/429 errors fail fast; transient 5xx errors trigger retry logic.
Once fetched, raw events pass through the parser pipeline in src/utils/subtitles/fetchers/youtube/parser. The system detects format types via detectFormat and applies specialized parsers:
- Karaoke subtitles –
parseKaraokeSubtitleshandles word-level timing - Scrolling ASR –
parseScrollingAsrSubtitlesprocesses auto-generated captions - Standard captions –
parseStandardSubtitlesfor traditional subtitle tracks
If AI segmentation is enabled (config.videoSubtitles.aiSegmentation), the AI-based parser reorganizes fragments for semantic coherence. Finally, optimizeSubtitles merges tiny cues and removes duplicates before passing the normalized fragments to the coordinator.
Real-Time Translation Coordination
The TranslationCoordinator class manages the critical path between video playback and translation delivery. Upon calling start(), it attaches DOM listeners for timeupdate and seeked events.
On each playback tick, the coordinator executes translateNearby(currentTimeMs):
- A look-ahead window (
TRANSLATE_LOOK_AHEAD_MS, default 30 seconds) defines which future subtitles are "nearby" - Up to
TRANSLATION_BATCH_SIZE(default 10) fragments are batched to balance latency against API efficiency - Only untranslated fragments within the window are selected
The coordinator calls translateSubtitles() from src/utils/subtitles/processor/translator.ts, which builds a deterministic hash of the fragment text, provider configuration, source/target language, and (for AI-aware providers) video title context. It then sends an enqueueSubtitlesTranslateRequest message to the background script via sendMessage.
// Inside TranslationCoordinator.translateNearby()
const batch = fragments
.filter(f => /* within look-ahead window and not yet translated */)
.slice(0, TRANSLATION_BATCH_SIZE);
if (batch.length) {
// Sends each fragment to the background translation queue
const translated = await translateSubtitles(batch, this.videoContext);
this.onTranslated(translated); // UI refresh
}
Background Queue and Intelligent Caching
The background script in src/entrypoints/background/translation-queues.ts handles the enqueueSubtitlesTranslateRequest message:
- Cache Lookup – Checks
db.translationCachefor existing results using the hash components built earlier. Cache hits return instantly without API calls. - Context Preparation – For LLM providers (
isLLMProviderConfig), the system may callgetOrGenerateSummaryto create a short video summary that improves translation quality through contextual awareness. - Queue Selection – Simple providers use the request queue (individual API calls), while LLM providers use the batch queue to maximize token efficiency.
- Result Storage – Fresh translations are stored in the cache with their hash keys for future reuse.
The translated fragments return to the TranslationCoordinator via the onTranslated callback, triggering UI components (subtitles-view.tsx, subtitles-translate-button.tsx) to render the text overlay. The display mode—original only, translation only, or bilingual side-by-side—respects user settings stored in the configuration.
Summary
- Read Frog's technical approach combines client-side subtitle fetching with a background translation queue to achieve real-time performance.
- The system selects optimal caption tracks from YouTube's API using
YoutubeSubtitlesFetcherand normalizes disparate formats (karaoke, ASR, standard) into unified fragments. - TranslationCoordinator implements a sliding window algorithm with 30-second look-ahead and 10-fragment batching to minimize API latency while maintaining synchronization with playback.
- A two-tier caching layer—checking
db.translationCachebefore API calls—eliminates redundant translations and reduces costs. - AI-aware providers receive video summaries for contextually accurate translations, processed through specialized batch queues in the background script.
Frequently Asked Questions
How does Read Frog handle different YouTube subtitle formats?
The parser subsystem detects format types via detectFormat and routes events to specialized handlers: parseKaraokeSubtitles for word-timed lyrics, parseScrollingAsrSubtitles for auto-generated captions, and parseStandardSubtitles for traditional tracks. If AI segmentation is enabled, the system further reorganizes fragments for semantic coherence before translation.
What is the default translation batch size and look-ahead window?
The system uses a 30-second look-ahead window (TRANSLATE_LOOK_AHEAD_MS) and batches up to 10 fragments (TRANSLATION_BATCH_SIZE) per translation request. This balances real-time responsiveness with API efficiency, ensuring subtitles appear before the viewer reaches that timestamp.
How does the caching mechanism prevent duplicate translations?
Before making API calls, the background queue checks db.translationCache using a deterministic hash built from fragment text, provider config, and language pair. If the hash exists, the cached result returns immediately. This eliminates redundant API calls when users rewind or rewatch video segments.
Which translation providers does Read Frog support?
The architecture abstracts providers into two categories: single-call providers processed through the request queue, and LLM providers processed through the batch queue. For AI-aware providers, the system can generate video summaries via getOrGenerateSummary to provide contextual grounding for higher-quality translations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →