Read Frog Technical Approach to Real-Time YouTube Subtitle Translation

Read Frog translates YouTube subtitles in real-time by chaining four tightly-coupled subsystems—subtitle acquisition, format normalization, a sliding-window translation coordinator, and a background queue with intelligent caching—to deliver low-latency translations synchronized with video playback.

Read Frog is an open-source browser extension that performs real-time translation of YouTube video subtitles without requiring users to provide their own API keys. The technical architecture processes caption tracks as time-series data, leveraging a TranslationCoordinator that schedules jobs based on the viewer's current playback position while a background queue handles provider-specific optimizations and result caching.

The Four-Layer Translation Architecture

The codebase divides the workflow into four distinct subsystems that operate sequentially:

  • Subtitle Acquisition – Detects the current video, selects the optimal caption track, and fetches raw timed-text events from YouTube's API. Implemented in src/utils/subtitles/fetchers/youtube/index.ts.
  • Subtitle Normalisation – Filters noise, detects subtitle formats (karaoke, scrolling ASR, or standard), and parses events into a unified SubtitlesFragment[] structure. Located in src/utils/subtitles/fetchers/youtube/parser.
  • Translation Coordination – Maintains a sliding window around the current playback time, batches upcoming fragments, and schedules them for translation. Core logic resides in src/entrypoints/subtitles.content/translation-coordinator.ts.
  • Background Translation Queue – Receives translation requests, checks the cache, and delegates to either a request-queue (single-call providers) or a batch-queue (LLM providers). Implemented in src/entrypoints/background/translation-queues.ts.

Subtitle Acquisition and Format Parsing

When a YouTube watch page loads, initYoutubeSubtitles() in src/entrypoints/subtitles.content/init-youtube-subtitles.ts mounts the UI and instantiates a YoutubeSubtitlesFetcher.

The fetcher executes the following chain:

  1. Video Detection – getYoutubeVideoId() extracts the video identifier, while waitForPlayerState ensures the player is ready.
  2. Track Selection – selectTrack prioritizes captions in this order: user-selected → human-uploaded → auto-generated.
  3. URL Construction – buildSubtitleUrl generates the authenticated timed-text endpoint using player data including the POT token.
  4. Error Handling – HTTP 403/404/429 errors fail fast; transient 5xx errors trigger retry logic.

Once fetched, raw events pass through the parser pipeline in src/utils/subtitles/fetchers/youtube/parser. The system detects format types via detectFormat and applies specialized parsers:

  • Karaoke subtitles – parseKaraokeSubtitles handles word-level timing
  • Scrolling ASR – parseScrollingAsrSubtitles processes auto-generated captions
  • Standard captions – parseStandardSubtitles for traditional subtitle tracks

If AI segmentation is enabled (config.videoSubtitles.aiSegmentation), the AI-based parser reorganizes fragments for semantic coherence. Finally, optimizeSubtitles merges tiny cues and removes duplicates before passing the normalized fragments to the coordinator.

Real-Time Translation Coordination

The TranslationCoordinator class manages the critical path between video playback and translation delivery. Upon calling start(), it attaches DOM listeners for timeupdate and seeked events.

On each playback tick, the coordinator executes translateNearby(currentTimeMs):

  • A look-ahead window (TRANSLATE_LOOK_AHEAD_MS, default 30 seconds) defines which future subtitles are "nearby"
  • Up to TRANSLATION_BATCH_SIZE (default 10) fragments are batched to balance latency against API efficiency
  • Only untranslated fragments within the window are selected

The coordinator calls translateSubtitles() from src/utils/subtitles/processor/translator.ts, which builds a deterministic hash of the fragment text, provider configuration, source/target language, and (for AI-aware providers) video title context. It then sends an enqueueSubtitlesTranslateRequest message to the background script via sendMessage.

// Inside TranslationCoordinator.translateNearby()
const batch = fragments
  .filter(f => /* within look-ahead window and not yet translated */)
  .slice(0, TRANSLATION_BATCH_SIZE);

if (batch.length) {
  // Sends each fragment to the background translation queue
  const translated = await translateSubtitles(batch, this.videoContext);
  this.onTranslated(translated); // UI refresh
}

Background Queue and Intelligent Caching

The background script in src/entrypoints/background/translation-queues.ts handles the enqueueSubtitlesTranslateRequest message:

  1. Cache Lookup – Checks db.translationCache for existing results using the hash components built earlier. Cache hits return instantly without API calls.
  2. Context Preparation – For LLM providers (isLLMProviderConfig), the system may call getOrGenerateSummary to create a short video summary that improves translation quality through contextual awareness.
  3. Queue Selection – Simple providers use the request queue (individual API calls), while LLM providers use the batch queue to maximize token efficiency.
  4. Result Storage – Fresh translations are stored in the cache with their hash keys for future reuse.

The translated fragments return to the TranslationCoordinator via the onTranslated callback, triggering UI components (subtitles-view.tsx, subtitles-translate-button.tsx) to render the text overlay. The display mode—original only, translation only, or bilingual side-by-side—respects user settings stored in the configuration.

Summary

  • Read Frog's technical approach combines client-side subtitle fetching with a background translation queue to achieve real-time performance.
  • The system selects optimal caption tracks from YouTube's API using YoutubeSubtitlesFetcher and normalizes disparate formats (karaoke, ASR, standard) into unified fragments.
  • TranslationCoordinator implements a sliding window algorithm with 30-second look-ahead and 10-fragment batching to minimize API latency while maintaining synchronization with playback.
  • A two-tier caching layer—checking db.translationCache before API calls—eliminates redundant translations and reduces costs.
  • AI-aware providers receive video summaries for contextually accurate translations, processed through specialized batch queues in the background script.

Frequently Asked Questions

How does Read Frog handle different YouTube subtitle formats?

The parser subsystem detects format types via detectFormat and routes events to specialized handlers: parseKaraokeSubtitles for word-timed lyrics, parseScrollingAsrSubtitles for auto-generated captions, and parseStandardSubtitles for traditional tracks. If AI segmentation is enabled, the system further reorganizes fragments for semantic coherence before translation.

What is the default translation batch size and look-ahead window?

The system uses a 30-second look-ahead window (TRANSLATE_LOOK_AHEAD_MS) and batches up to 10 fragments (TRANSLATION_BATCH_SIZE) per translation request. This balances real-time responsiveness with API efficiency, ensuring subtitles appear before the viewer reaches that timestamp.

How does the caching mechanism prevent duplicate translations?

Before making API calls, the background queue checks db.translationCache using a deterministic hash built from fragment text, provider config, and language pair. If the hash exists, the cached result returns immediately. This eliminates redundant API calls when users rewind or rewatch video segments.

Which translation providers does Read Frog support?

The architecture abstracts providers into two categories: single-call providers processed through the request queue, and LLM providers processed through the batch queue. For AI-aware providers, the system can generate video summaries via getOrGenerateSummary to provide contextual grounding for higher-quality translations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →