How Read Frog Leverages Mozilla Readability for Context-Aware Translation of Web Content

Read Frog uses Mozilla's Readability library to extract the main article content from web pages, stripping navigation menus and advertisements, then feeds that clean context to AI translation providers for more accurate and coherent translations.

Context-aware translation requires more than raw text extraction from the DOM. The open-source browser extension Read Frog (mengxi-ream/read-frog) solves this by integrating Mozilla's Readability parser to isolate meaningful article content before sending it to large language model (LLM) providers. This approach ensures that AI translators receive the full narrative context of a page rather than fragmented menu items, advertisements, or script noise.

Why Raw DOM Extraction Fails for AI Translation

Extracting document.body.textContent floods translation models with irrelevant data. Navigation bars, cookie banners, sidebar widgets, and footer links dilute the semantic context, causing AI providers to generate translations that lack coherence or focus. Read Frog avoids this by parsing the semantic structure of the document to identify the actual article body.

Cloning and Sanitizing the DOM for Readability

Before parsing, Read Frog creates an isolated environment to protect the live page. The process begins in src/utils/host/translate/translate-text.ts by cloning the entire document and removing extension-specific artifacts.

A deep clone of the current page is created to prevent mutations to the active DOM. The removeDummyNodes function (imported from src/utils/content/utils.ts) then strips injected extension UI elements that would otherwise confuse the parser. This sanitization occurs at lines 94–100 of the translation module.

// src/utils/host/translate/translate-text.ts
const documentClone = document.cloneNode(true) as Document
await removeDummyNodes(documentClone) // strip extension UI

Parsing Articles with Mozilla Readability

The sanitized clone is passed to Mozilla's Readability constructor configured with an identity serializer to preserve HTML structure. The parse() method returns an article object containing title, textContent, and optional HTML content.

This implementation resides in the main translation flow and provides the foundation for context-aware translation:

const article = new Readability(documentClone, {
  serializer: el => el,
}).parse()

let textContent = ""
if (article?.textContent) {
  textContent = article.textContent // clean article body
}

The same parsing logic is reused in src/utils/content/analyze.ts (lines 18–23) for language detection and other content analyses, ensuring consistency across the extension.

Handling Parse Failures with Graceful Fallbacks

When Readability encounters unsupported layouts or complex web applications, the parser may return null or empty content. Read Frog implements a defensive fallback strategy at lines 108–110 of translate-text.ts to ensure translation proceeds regardless of parsing success.

if (!textContent) {
  textContent = document.body?.textContent || "" // raw fallback
}

This guarantees that users receive translations even on pages where Readability fails to identify a main article, though without the benefits of context-aware extraction.

Caching Extracted Content for Performance

Re-parsing the DOM on every translation request would waste computational resources. Read Frog maintains a cachedArticleData map keyed by URL to store extracted article objects, implemented at lines 52–65 of translate-text.ts.

When getOrFetchArticleData() is invoked, it first checks the cache. If the article data exists for the current URL, the extension reuses the cached title and textContent rather than re-running the Readability parser.

Injecting Article Context into Translation Requests

The context-aware capability activates when AI-aware providers are configured. The buildHashComponents function incorporates article metadata into the request hash and payload, ensuring the LLM receives the full page context.

At lines 47–56, the code checks enableAIContentAware and appends the article title and a truncated snippet of content (limited to 1000 characters for hash size efficiency):

if (enableAIContentAware && articleContext) {
  if (articleContext.title) {
    hashComponents.push(`title:${articleContext.title}`)
  }
  if (articleContext.textContent) {
    // keep hash size reasonable
    hashComponents.push(`content:${articleContext.textContent.slice(0, 1000)}`)
  }
}

These components become part of the request sent via sendMessage("enqueueTranslateRequest", ...):

return await sendMessage("enqueueTranslateRequest", {
  text,
  langConfig,
  providerConfig,
  scheduleAt: Date.now(),
  hash: Sha256Hex(...hashComponents),
  articleTitle,
  articleTextContent,
})

Summary

  • DOM Isolation: Read Frog clones the document and removes dummy nodes before parsing to ensure Readability processes clean HTML.
  • Semantic Extraction: Mozilla's Readability library identifies the main article body, excluding navigation and advertisements.
  • Resilient Fallbacks: When parsing fails, the extension falls back to document.body.textContent to maintain functionality.
  • Intelligent Caching: Article data is cached per URL via cachedArticleData to avoid redundant parsing operations.
  • AI Context Enhancement: The buildHashComponents function injects article titles and content snippets into LLM requests, enabling true context-aware translation.

Frequently Asked Questions

What is context-aware translation in Read Frog?

Context-aware translation refers to the process of providing AI translation models with the full semantic context of a web article—including its title and main body content—rather than isolated text fragments. By using Mozilla Readability to extract the primary article content, Read Frog ensures that LLM providers understand the broader narrative context, resulting in more accurate and natural translations.

How does Readability improve translation quality compared to body.textContent?

Mozilla Readability parses the DOM structure to identify the main content area of a page, filtering out navigation menus, sidebars, footers, and advertisements. According to the read-frog source code, feeding this cleaned article.textContent to AI providers eliminates noise that would otherwise confuse translation models, whereas raw document.body.textContent contains fragmented UI text that degrades translation coherence.

What happens when Readability cannot parse a page?

When the Readability parser fails to identify an article—such as on complex web applications or unsupported layouts—Read Frog falls back to document.body.textContent as implemented in src/utils/host/translate/translate-text.ts lines 108–110. This ensures that translation requests still proceed, though without the semantic filtering benefits of successful Readability parsing.

How does Read Frog cache article data to improve performance?

The extension maintains a cachedArticleData object that stores parsed article objects keyed by URL, preventing redundant DOM parsing when users trigger multiple translations on the same page. This caching mechanism, located at lines 52–65 of translate-text.ts, allows getOrFetchArticleData() to return immediately with existing context rather than re-cloning and re-parsing the document.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →