# How Read Frog Leverages Mozilla Readability for Context-Aware Translation of Web Content

> Read Frog uses Mozilla Readability to extract article content for AI translation. Get accurate, context-aware web translations by cleaning web pages before AI processing.

- Repository: [MengXi/read-frog](https://github.com/mengxi-ream/read-frog)
- Tags: how-to-guide
- Published: 2026-03-07

---

**Read Frog uses Mozilla's Readability library to extract the main article content from web pages, stripping navigation menus and advertisements, then feeds that clean context to AI translation providers for more accurate and coherent translations.**

Context-aware translation requires more than raw text extraction from the DOM. The open-source browser extension Read Frog (mengxi-ream/read-frog) solves this by integrating Mozilla's Readability parser to isolate meaningful article content before sending it to large language model (LLM) providers. This approach ensures that AI translators receive the full narrative context of a page rather than fragmented menu items, advertisements, or script noise.

## Why Raw DOM Extraction Fails for AI Translation

Extracting `document.body.textContent` floods translation models with irrelevant data. Navigation bars, cookie banners, sidebar widgets, and footer links dilute the semantic context, causing AI providers to generate translations that lack coherence or focus. Read Frog avoids this by parsing the semantic structure of the document to identify the actual article body.

## Cloning and Sanitizing the DOM for Readability

Before parsing, Read Frog creates an isolated environment to protect the live page. The process begins in [`src/utils/host/translate/translate-text.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/host/translate/translate-text.ts) by cloning the entire document and removing extension-specific artifacts.

A deep clone of the current page is created to prevent mutations to the active DOM. The `removeDummyNodes` function (imported from [`src/utils/content/utils.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/content/utils.ts)) then strips injected extension UI elements that would otherwise confuse the parser. This sanitization occurs at lines 94–100 of the translation module.

```typescript
// src/utils/host/translate/translate-text.ts
const documentClone = document.cloneNode(true) as Document
await removeDummyNodes(documentClone) // strip extension UI

```

## Parsing Articles with Mozilla Readability

The sanitized clone is passed to Mozilla's Readability constructor configured with an identity serializer to preserve HTML structure. The `parse()` method returns an `article` object containing `title`, `textContent`, and optional HTML `content`.

This implementation resides in the main translation flow and provides the foundation for context-aware translation:

```typescript
const article = new Readability(documentClone, {
  serializer: el => el,
}).parse()

let textContent = ""
if (article?.textContent) {
  textContent = article.textContent // clean article body
}

```

The same parsing logic is reused in [`src/utils/content/analyze.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/content/analyze.ts) (lines 18–23) for language detection and other content analyses, ensuring consistency across the extension.

## Handling Parse Failures with Graceful Fallbacks

When Readability encounters unsupported layouts or complex web applications, the parser may return null or empty content. Read Frog implements a defensive fallback strategy at lines 108–110 of [`translate-text.ts`](https://github.com/mengxi-ream/read-frog/blob/main/translate-text.ts) to ensure translation proceeds regardless of parsing success.

```typescript
if (!textContent) {
  textContent = document.body?.textContent || "" // raw fallback
}

```

This guarantees that users receive translations even on pages where Readability fails to identify a main article, though without the benefits of context-aware extraction.

## Caching Extracted Content for Performance

Re-parsing the DOM on every translation request would waste computational resources. Read Frog maintains a `cachedArticleData` map keyed by URL to store extracted article objects, implemented at lines 52–65 of [`translate-text.ts`](https://github.com/mengxi-ream/read-frog/blob/main/translate-text.ts).

When `getOrFetchArticleData()` is invoked, it first checks the cache. If the article data exists for the current URL, the extension reuses the cached `title` and `textContent` rather than re-running the Readability parser.

## Injecting Article Context into Translation Requests

The context-aware capability activates when AI-aware providers are configured. The `buildHashComponents` function incorporates article metadata into the request hash and payload, ensuring the LLM receives the full page context.

At lines 47–56, the code checks `enableAIContentAware` and appends the article title and a truncated snippet of content (limited to 1000 characters for hash size efficiency):

```typescript
if (enableAIContentAware && articleContext) {
  if (articleContext.title) {
    hashComponents.push(`title:${articleContext.title}`)
  }
  if (articleContext.textContent) {
    // keep hash size reasonable
    hashComponents.push(`content:${articleContext.textContent.slice(0, 1000)}`)
  }
}

```

These components become part of the request sent via `sendMessage("enqueueTranslateRequest", ...)`:

```typescript
return await sendMessage("enqueueTranslateRequest", {
  text,
  langConfig,
  providerConfig,
  scheduleAt: Date.now(),
  hash: Sha256Hex(...hashComponents),
  articleTitle,
  articleTextContent,
})

```

## Summary

- **DOM Isolation**: Read Frog clones the document and removes dummy nodes before parsing to ensure Readability processes clean HTML.
- **Semantic Extraction**: Mozilla's Readability library identifies the main article body, excluding navigation and advertisements.
- **Resilient Fallbacks**: When parsing fails, the extension falls back to `document.body.textContent` to maintain functionality.
- **Intelligent Caching**: Article data is cached per URL via `cachedArticleData` to avoid redundant parsing operations.
- **AI Context Enhancement**: The `buildHashComponents` function injects article titles and content snippets into LLM requests, enabling true context-aware translation.

## Frequently Asked Questions

### What is context-aware translation in Read Frog?

Context-aware translation refers to the process of providing AI translation models with the full semantic context of a web article—including its title and main body content—rather than isolated text fragments. By using Mozilla Readability to extract the primary article content, Read Frog ensures that LLM providers understand the broader narrative context, resulting in more accurate and natural translations.

### How does Readability improve translation quality compared to body.textContent?

Mozilla Readability parses the DOM structure to identify the main content area of a page, filtering out navigation menus, sidebars, footers, and advertisements. According to the read-frog source code, feeding this cleaned `article.textContent` to AI providers eliminates noise that would otherwise confuse translation models, whereas raw `document.body.textContent` contains fragmented UI text that degrades translation coherence.

### What happens when Readability cannot parse a page?

When the Readability parser fails to identify an article—such as on complex web applications or unsupported layouts—Read Frog falls back to `document.body.textContent` as implemented in [`src/utils/host/translate/translate-text.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/host/translate/translate-text.ts) lines 108–110. This ensures that translation requests still proceed, though without the semantic filtering benefits of successful Readability parsing.

### How does Read Frog cache article data to improve performance?

The extension maintains a `cachedArticleData` object that stores parsed article objects keyed by URL, preventing redundant DOM parsing when users trigger multiple translations on the same page. This caching mechanism, located at lines 52–65 of [`translate-text.ts`](https://github.com/mengxi-ream/read-frog/blob/main/translate-text.ts), allows `getOrFetchArticleData()` to return immediately with existing context rather than re-cloning and re-parsing the document.