# Read Article Side Panel Content Script in Read Frog: Architecture and Implementation

> Explore the Read Frog side panel content script architecture. Discover how it extracts article text and SEO metadata using DOM traversal for a seamless reading experience.

- Repository: [MengXi/read-frog](https://github.com/mengxi-ream/read-frog)
- Tags: architecture
- Published: 2026-03-07

---

**The Read Article side panel content script in Read Frog creates an isolated shadow-DOM overlay that extracts article text and SEO metadata from web pages using DOM traversal algorithms, then renders a resizable React side panel for reading.**

The Read Frog browser extension (available at `mengxi-ream/read-frog`) implements its Read Article feature through a sophisticated content script architecture. This script injects a fully isolated UI overlay onto any web page, extracts readable content using custom DOM traversal logic, and presents it in a resizable side panel without interfering with the host page's styles or scripts.

## Entry Point and Shadow DOM Initialization

The content script registers itself via [`src/entrypoints/side.content/index.tsx`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/side.content/index.tsx), using the WXT framework's `defineContentScript` API to match all URLs (`*://*/*` and `file:///*`).

### Content Script Registration

The entry point defines the script with CSS injection mode set to `"ui"`, allowing styles to be scoped to the extension's shadow DOM rather than the page:

```tsx
// src/entrypoints/side.content/index.tsx
export default defineContentScript({
  matches: ["*://*/*", "file:///*"],
  cssInjectionMode: "ui",
  async main(ctx) {
    const config = await getLocalConfig() ?? DEFAULT_CONFIG
    if (!isSiteEnabled(window.location.href, config)) return

    const ui = await createShadowRootUi(ctx, {
      name: kebabCase(APP_NAME),
      position: "overlay",
      anchor: "body",
      append: "last",
      // ...
    })
    ui.mount()
  },
})

```

The `isSiteEnabled` utility (from [`src/utils/site-control.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/site-control.ts)) checks the current URL against user preferences before initialization, preventing the script from running on disabled sites.

### Shadow DOM Setup and Style Injection

Inside the `onMount` callback, the script constructs an isolated rendering environment using `createShadowRootUi` from [`src/utils/shadow-root.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/shadow-root.ts). This creates a shadow root attached to the document body with `position: "overlay"`, ensuring complete style isolation:

```tsx
onMount: (container, shadow, shadowHost) => {
  const wrapper = insertShadowRootUIWrapperInto(container)
  shadowWrapper = wrapper

  addStyleToShadow(shadow)
  mirrorDynamicStyles("#_goober", shadow)
  protectInternalStyles()
  protectSelectAllShadowRoot(shadowHost, wrapper)
  // ...
}

```

The **style protection system** serves two critical functions: `addStyleToShadow` copies extension-wide CSS into the shadow boundary, while `mirrorDynamicStyles` synchronizes dynamically generated styles (like those from CSS-in-JS libraries) into the isolated context. `protectInternalStyles` prevents the host page from accidentally removing extension styles via aggressive CSS resets.

### State Hydration with Jotai

Before rendering React components, the script hydrates the Jotai atom store with persisted configuration data. The `useHydrateAtoms` hook injects values retrieved from Chrome extension storage via `getLocalConfig` ([`src/utils/config/storage.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/config/storage.ts)):

```tsx
const HydrateAtoms = ({
  initialValues,
  children,
}: { initialValues: [[typeof configAtom, Config]]; children: React.ReactNode }) => {
  useHydrateAtoms(initialValues)
  return children
}

const root = ReactDOM.createRoot(wrapper)
root.render(
  <QueryClientProvider client={queryClient}>
    <JotaiProvider store={store}>
      <HydrateAtoms initialValues={[[configAtom, config]]}>
        <ThemeProvider container={wrapper}>
          <TooltipProvider>
            <App />
          </TooltipProvider>
        </ThemeProvider>
      </HydrateAtoms>
    </JotaiProvider>
  </QueryClientProvider>,
)

```

## Article Extraction Engine

The core content processing logic resides in [`src/entrypoints/side.content/utils/article.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/side.content/utils/article.ts), providing two primary export functions: `flattenToParagraphs` for text extraction and `extractSeoInfo` for metadata collection.

### DOM Traversal with flattenToParagraphs

The **flattening algorithm** performs a depth-first traversal of the DOM tree to extract readable paragraphs while preserving semantic structure:

```ts
// src/entrypoints/side.content/utils/article.ts
export function flattenToParagraphs(root: Node) {
  const semanticBlocks = new Set(['P', 'ARTICLE', 'SECTION', 'DIV', 'MAIN', 'BLOCKQUOTE'])
  
  // Helper functions:
  // isBlockLevel - checks tag name or computed display style
  // hasBlockDescendant - verifies if node contains nested blocks
  // getTextWithSpaces - recursively concatenates text nodes with spacing
  
  // Returns array of cleaned paragraph strings (≥ 20 chars)
}

```

The algorithm identifies **leaf block elements**—block-level containers that do not contain other block elements—and extracts their text content. It handles inline elements by inserting spaces between them to prevent word concatenation, then filters out paragraphs shorter than 20 characters to remove navigation cruft and decorative elements.

### SEO Metadata Extraction

The `extractSeoInfo` function gathers comprehensive metadata for downstream processing such as summarization or translation:

```ts
export function extractSeoInfo(doc: Document) {
  return {
    title: doc.title,
    metaDescription: doc.querySelector('meta[name="description"]')?.getAttribute('content') ?? '',
    metaKeywords: doc.querySelector('meta[name="keywords"]')?.getAttribute('content') ?? '',
    canonicalUrl: doc.querySelector('link[rel="canonical"]')?.getAttribute('href') ?? '',
    ogTitle: doc.querySelector('meta[property="og:title"]')?.getAttribute('content') ?? '',
    ogDescription: doc.querySelector('meta[property="og:description"]')?.getAttribute('content') ?? '',
    h1Tags: Array.from(doc.querySelectorAll('h1')).map(h => h.textContent?.trim() ?? ''),
    structuredData: Array.from(
      doc.querySelectorAll('script[type="application/ld+json"]')
    ).map(script => {
      try { return JSON.parse(script.textContent || '{}') }
      catch { return {} }
    })
  }
}

```

This extracts standard meta tags, Open Graph properties, Twitter card data, H1 headings, and JSON-LD structured data schemas.

## Side Panel UI Components

The visual interface is implemented in [`src/entrypoints/side.content/components/side-content/index.tsx`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/side.content/components/side-content/index.tsx), featuring a resizable panel that slides in from the right edge of the viewport.

### Resizable Panel Implementation

The component uses mouse event listeners to calculate new widths based on cursor position, storing the value in a Jotai atom (`sideContent` via `configFieldsAtomMap.sideContent`):

```tsx
useEffect(() => {
  if (!isResizing) return
  const handleMouseMove = (e: MouseEvent) => {
    const newWidth = Math.max(MIN_SIDE_CONTENT_WIDTH, window.innerWidth - e.clientX)
    void setSideContent({ width: newWidth })
  }
  const handleMouseUp = () => setIsResizing(false)
  
  document.addEventListener('mousemove', handleMouseMove)
  document.addEventListener('mouseup', handleMouseUp)
  document.body.style.userSelect = 'none'
  
  return () => {
    document.removeEventListener('mousemove', handleMouseMove)
    document.removeEventListener('mouseup', handleMouseUp)
    document.body.style.userSelect = ''
  }
}, [isResizing, setSideContent])

```

The resize handle sits at the left edge of the panel with `cursor-ew-resize`, while a transparent full-screen overlay (`z-[2147483647]`) captures mouse events during dragging to prevent text selection on the underlying page.

### Page Width Adjustment

To prevent content from scrolling underneath the panel, the component dynamically injects a style tag into the document head when `isSideOpenAtom` is true:

```tsx
useEffect(() => {
  const styleId = `shrink-origin-for-${kebabCase(APP_NAME)}-side-content`
  let styleTag = document.getElementById(styleId)
  
  if (isSideOpen) {
    if (!styleTag) {
      styleTag = document.createElement('style')
      styleTag.id = styleId
      document.head.appendChild(styleTag)
    }
    styleTag.textContent = `
      html { width: calc(100% - ${sideContent.width}px) !important; position: relative !important; min-height: 100vh !important; }
    `
  } else if (styleTag) {
    document.head.removeChild(styleTag)
  }
}, [isSideOpen, sideContent.width])

```

This reduces the effective viewport width by the panel width, forcing the host page to reflow naturally alongside the extension UI.

## Application Orchestration

The `App` component in [`src/entrypoints/side.content/app.tsx`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/side.content/app.tsx) serves as the orchestration layer, triggering article extraction when the user opens the panel:

```tsx
export default function App() {
  const [article, setArticle] = useState<string[]>([])
  const isSideOpen = useAtomValue(isSideOpenAtom)

  useEffect(() => {
    if (!isSideOpen) return
    const paragraphs = flattenToParagraphs(document.body)
    const seo = extractSeoInfo(document)
    setArticle(paragraphs)
    // Data flows to downstream UI components
  }, [isSideOpen])

  return (
    <>
      <SideContent />
    </>
  )
}

```

This lazy-loads content only when the panel is first activated, improving initial page load performance.

## Configuration and State Management

The content script relies on several utility modules for state persistence:

- **[`src/utils/atoms/config.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/atoms/config.ts)** – Defines `configAtom` and related atoms for reactive state management
- **[`src/utils/config/storage.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/config/storage.ts)** – Implements `getLocalConfig()` for Chrome storage API integration
- **[`src/utils/site-control.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/site-control.ts)** – Provides `isSiteEnabled()` for URL-based feature gating
- **[`src/utils/shadow-root.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/shadow-root.ts)** – Contains `createShadowRootUi` and `insertShadowRootUIWrapperInto` for DOM isolation
- **[`src/utils/styles/utils.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/styles/utils.ts)** – Exports `addStyleToShadow`, `mirrorDynamicStyles`, and the `cn()` class name utility

These utilities ensure the content script remains thin and reusable across other entry points like the subtitle overlay or selection toolbar.

## Summary

- The Read Article feature runs as a content script in [`src/entrypoints/side.content/index.tsx`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/side.content/index.tsx) that creates an isolated shadow-DOM overlay on every webpage.
- **Style isolation** is achieved through `createShadowRootUi` combined with `addStyleToShadow` and `mirrorDynamicStyles`, preventing CSS leakage between the extension and host page.
- Article extraction uses `flattenToParagraphs` in [`src/entrypoints/side.content/utils/article.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/side.content/utils/article.ts) to perform depth-first DOM traversal, identifying leaf block elements and cleaning text content.
- SEO metadata collection via `extractSeoInfo` gathers standard meta tags, Open Graph data, H1 elements, and JSON-LD structured data.
- The side panel UI in [`components/side-content/index.tsx`](https://github.com/mengxi-ream/read-frog/blob/main/components/side-content/index.tsx) implements mouse-based resizing with dynamic page-width adjustment to prevent layout conflicts.
- State management uses Jotai atoms hydrated from Chrome extension storage, allowing persistent user preferences across sessions.

## Frequently Asked Questions

### How does the content script isolate its UI from the host page?

The script uses **shadow DOM encapsulation** via `createShadowRootUi` from [`src/utils/shadow-root.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/shadow-root.ts). This attaches a closed shadow root to a host element injected into the page body, creating a style boundary that prevents the host page's CSS from affecting the extension UI and vice versa. Additional protections like `protectInternalStyles` guard against aggressive page styles that might otherwise pierce the shadow boundary.

### What algorithm extracts article text from web pages?

The **`flattenToParagraphs`** function in [`src/entrypoints/side.content/utils/article.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/entrypoints/side.content/utils/article.ts) implements a recursive depth-first traversal that identifies block-level elements without nested block children (leaf blocks). It extracts text content from these nodes while preserving spacing between inline elements, then filters results to include only paragraphs with 20 or more characters, effectively removing navigation menus and decorative text.

### How does the side panel handle resizing without breaking page layout?

The `SideContent` component tracks mouse movements to calculate new widths stored in Jotai atoms. Simultaneously, it injects a dynamic style tag into the document head that sets `html { width: calc(100% - ${width}px) }` when open. This forces the host page to reflow within the remaining viewport space, preventing horizontal scrollbars and content occlusion.

### Where is the extension configuration stored and how is it accessed?

Configuration persists in **Chrome extension storage** via utilities in [`src/utils/config/storage.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/config/storage.ts). The `getLocalConfig()` function retrieves settings asynchronously during script initialization in [`index.tsx`](https://github.com/mengxi-ream/read-frog/blob/main/index.tsx), while `useHydrateAtoms` injects these values into the Jotai state store. The `isSiteEnabled` utility then checks the current URL against these preferences to determine whether the side panel should activate.