# Read Frog Technical Approach to Real-Time YouTube Subtitle Translation

> Discover Read Frog's real-time YouTube subtitle translation approach. Learn how subsystems chain for low-latency, synchronized subtitles with video playback.

- Repository: [MengXi/read-frog](https://github.com/mengxi-ream/read-frog)
- Tags: technical-approach
- Published: 2026-03-07

---

**Read Frog translates YouTube subtitles in real-time by chaining four tightly-coupled subsystems—subtitle acquisition, format normalization, a sliding-window translation coordinator, and a background queue with intelligent caching—to deliver low-latency translations synchronized with video playback.**

Read Frog is an open-source browser extension that performs real-time translation of YouTube video subtitles without requiring users to provide their own API keys. The technical architecture processes caption tracks as time-series data, leveraging a **TranslationCoordinator** that schedules jobs based on the viewer's current playback position while a background queue handles provider-specific optimizations and result caching.

## The Four-Layer Translation Architecture

The codebase divides the workflow into four distinct subsystems that operate sequentially:

- **Subtitle Acquisition** – Detects the current video, selects the optimal caption track, and fetches raw timed-text events from YouTube's API. Implemented in [`src/utils/subtitles/fetchers/youtube/index.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/subtitles/fetchers/youtube/index.ts).
- **Subtitle Normalisation** – Filters noise, detects subtitle formats (karaoke, scrolling ASR, or standard), and parses events into a unified `SubtitlesFragment[]` structure. Located in `src/utils/subtitles/fetchers/youtube/parser`.
- **Translation Coordination** – Maintains a sliding window around the current playback time, batches upcoming fragments, and schedules them for translation. Core logic resides in [`src/entrypoints/subtitles.content/translation-coordinator.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/subtitles.content/translation-coordinator.ts).
- **Background Translation Queue** – Receives translation requests, checks the cache, and delegates to either a **request-queue** (single-call providers) or a **batch-queue** (LLM providers). Implemented in [`src/entrypoints/background/translation-queues.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/background/translation-queues.ts).

## Subtitle Acquisition and Format Parsing

When a YouTube watch page loads, `initYoutubeSubtitles()` in [`src/entrypoints/subtitles.content/init-youtube-subtitles.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/subtitles.content/init-youtube-subtitles.ts) mounts the UI and instantiates a `YoutubeSubtitlesFetcher`.

The fetcher executes the following chain:

1. **Video Detection** – `getYoutubeVideoId()` extracts the video identifier, while `waitForPlayerState` ensures the player is ready.
2. **Track Selection** – `selectTrack` prioritizes captions in this order: user-selected → human-uploaded → auto-generated.
3. **URL Construction** – `buildSubtitleUrl` generates the authenticated timed-text endpoint using player data including the POT token.
4. **Error Handling** – HTTP 403/404/429 errors fail fast; transient 5xx errors trigger retry logic.

Once fetched, raw events pass through the parser pipeline in `src/utils/subtitles/fetchers/youtube/parser`. The system detects format types via `detectFormat` and applies specialized parsers:

- **Karaoke subtitles** – `parseKaraokeSubtitles` handles word-level timing
- **Scrolling ASR** – `parseScrollingAsrSubtitles` processes auto-generated captions
- **Standard captions** – `parseStandardSubtitles` for traditional subtitle tracks

If **AI segmentation** is enabled (`config.videoSubtitles.aiSegmentation`), the AI-based parser reorganizes fragments for semantic coherence. Finally, `optimizeSubtitles` merges tiny cues and removes duplicates before passing the normalized fragments to the coordinator.

## Real-Time Translation Coordination

The `TranslationCoordinator` class manages the critical path between video playback and translation delivery. Upon calling `start()`, it attaches DOM listeners for `timeupdate` and `seeked` events.

On each playback tick, the coordinator executes `translateNearby(currentTimeMs)`:

- A **look-ahead window** (`TRANSLATE_LOOK_AHEAD_MS`, default 30 seconds) defines which future subtitles are "nearby"
- Up to `TRANSLATION_BATCH_SIZE` (default 10) fragments are batched to balance latency against API efficiency
- Only untranslated fragments within the window are selected

The coordinator calls `translateSubtitles()` from [`src/utils/subtitles/processor/translator.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/subtitles/processor/translator.ts), which builds a deterministic hash of the fragment text, provider configuration, source/target language, and (for AI-aware providers) video title context. It then sends an `enqueueSubtitlesTranslateRequest` message to the background script via `sendMessage`.

```typescript
// Inside TranslationCoordinator.translateNearby()
const batch = fragments
  .filter(f => /* within look-ahead window and not yet translated */)
  .slice(0, TRANSLATION_BATCH_SIZE);

if (batch.length) {
  // Sends each fragment to the background translation queue
  const translated = await translateSubtitles(batch, this.videoContext);
  this.onTranslated(translated); // UI refresh
}

```

## Background Queue and Intelligent Caching

The background script in [`src/entrypoints/background/translation-queues.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/background/translation-queues.ts) handles the `enqueueSubtitlesTranslateRequest` message:

1. **Cache Lookup** – Checks `db.translationCache` for existing results using the hash components built earlier. Cache hits return instantly without API calls.
2. **Context Preparation** – For LLM providers (`isLLMProviderConfig`), the system may call `getOrGenerateSummary` to create a short video summary that improves translation quality through contextual awareness.
3. **Queue Selection** – Simple providers use the **request queue** (individual API calls), while LLM providers use the **batch queue** to maximize token efficiency.
4. **Result Storage** – Fresh translations are stored in the cache with their hash keys for future reuse.

The translated fragments return to the `TranslationCoordinator` via the `onTranslated` callback, triggering UI components ([`subtitles-view.tsx`](https://github.com/mengxi-ream/read-frog/blob/main/subtitles-view.tsx), [`subtitles-translate-button.tsx`](https://github.com/mengxi-ream/read-frog/blob/main/subtitles-translate-button.tsx)) to render the text overlay. The display mode—original only, translation only, or bilingual side-by-side—respects user settings stored in the configuration.

## Summary

- Read Frog's technical approach combines client-side subtitle fetching with a background translation queue to achieve real-time performance.
- The system selects optimal caption tracks from YouTube's API using `YoutubeSubtitlesFetcher` and normalizes disparate formats (karaoke, ASR, standard) into unified fragments.
- **TranslationCoordinator** implements a sliding window algorithm with 30-second look-ahead and 10-fragment batching to minimize API latency while maintaining synchronization with playback.
- A two-tier caching layer—checking `db.translationCache` before API calls—eliminates redundant translations and reduces costs.
- AI-aware providers receive video summaries for contextually accurate translations, processed through specialized batch queues in the background script.

## Frequently Asked Questions

### How does Read Frog handle different YouTube subtitle formats?

The parser subsystem detects format types via `detectFormat` and routes events to specialized handlers: `parseKaraokeSubtitles` for word-timed lyrics, `parseScrollingAsrSubtitles` for auto-generated captions, and `parseStandardSubtitles` for traditional tracks. If AI segmentation is enabled, the system further reorganizes fragments for semantic coherence before translation.

### What is the default translation batch size and look-ahead window?

The system uses a **30-second look-ahead window** (`TRANSLATE_LOOK_AHEAD_MS`) and batches up to **10 fragments** (`TRANSLATION_BATCH_SIZE`) per translation request. This balances real-time responsiveness with API efficiency, ensuring subtitles appear before the viewer reaches that timestamp.

### How does the caching mechanism prevent duplicate translations?

Before making API calls, the background queue checks `db.translationCache` using a deterministic hash built from fragment text, provider config, and language pair. If the hash exists, the cached result returns immediately. This eliminates redundant API calls when users rewind or rewatch video segments.

### Which translation providers does Read Frog support?

The architecture abstracts providers into two categories: **single-call providers** processed through the request queue, and **LLM providers** processed through the batch queue. For AI-aware providers, the system can generate video summaries via `getOrGenerateSummary` to provide contextual grounding for higher-quality translations.