Read Article Side Panel Content Script in Read Frog: Architecture and Implementation

The Read Article side panel content script in Read Frog creates an isolated shadow-DOM overlay that extracts article text and SEO metadata from web pages using DOM traversal algorithms, then renders a resizable React side panel for reading.

The Read Frog browser extension (available at mengxi-ream/read-frog) implements its Read Article feature through a sophisticated content script architecture. This script injects a fully isolated UI overlay onto any web page, extracts readable content using custom DOM traversal logic, and presents it in a resizable side panel without interfering with the host page's styles or scripts.

Entry Point and Shadow DOM Initialization

The content script registers itself via src/entrypoints/side.content/index.tsx, using the WXT framework's defineContentScript API to match all URLs (*://*/* and file:///*).

Content Script Registration

The entry point defines the script with CSS injection mode set to "ui", allowing styles to be scoped to the extension's shadow DOM rather than the page:

// src/entrypoints/side.content/index.tsx
export default defineContentScript({
  matches: ["*://*/*", "file:///*"],
  cssInjectionMode: "ui",
  async main(ctx) {
    const config = await getLocalConfig() ?? DEFAULT_CONFIG
    if (!isSiteEnabled(window.location.href, config)) return

    const ui = await createShadowRootUi(ctx, {
      name: kebabCase(APP_NAME),
      position: "overlay",
      anchor: "body",
      append: "last",
      // ...
    })
    ui.mount()
  },
})

The isSiteEnabled utility (from src/utils/site-control.ts) checks the current URL against user preferences before initialization, preventing the script from running on disabled sites.

Shadow DOM Setup and Style Injection

Inside the onMount callback, the script constructs an isolated rendering environment using createShadowRootUi from src/utils/shadow-root.ts. This creates a shadow root attached to the document body with position: "overlay", ensuring complete style isolation:

onMount: (container, shadow, shadowHost) => {
  const wrapper = insertShadowRootUIWrapperInto(container)
  shadowWrapper = wrapper

  addStyleToShadow(shadow)
  mirrorDynamicStyles("#_goober", shadow)
  protectInternalStyles()
  protectSelectAllShadowRoot(shadowHost, wrapper)
  // ...
}

The style protection system serves two critical functions: addStyleToShadow copies extension-wide CSS into the shadow boundary, while mirrorDynamicStyles synchronizes dynamically generated styles (like those from CSS-in-JS libraries) into the isolated context. protectInternalStyles prevents the host page from accidentally removing extension styles via aggressive CSS resets.

State Hydration with Jotai

Before rendering React components, the script hydrates the Jotai atom store with persisted configuration data. The useHydrateAtoms hook injects values retrieved from Chrome extension storage via getLocalConfig (src/utils/config/storage.ts):

const HydrateAtoms = ({
  initialValues,
  children,
}: { initialValues: [[typeof configAtom, Config]]; children: React.ReactNode }) => {
  useHydrateAtoms(initialValues)
  return children
}

const root = ReactDOM.createRoot(wrapper)
root.render(
  <QueryClientProvider client={queryClient}>
    <JotaiProvider store={store}>
      <HydrateAtoms initialValues={[[configAtom, config]]}>
        <ThemeProvider container={wrapper}>
          <TooltipProvider>
            <App />
          </TooltipProvider>
        </ThemeProvider>
      </HydrateAtoms>
    </JotaiProvider>
  </QueryClientProvider>,
)

Article Extraction Engine

The core content processing logic resides in src/entrypoints/side.content/utils/article.ts, providing two primary export functions: flattenToParagraphs for text extraction and extractSeoInfo for metadata collection.

DOM Traversal with flattenToParagraphs

The flattening algorithm performs a depth-first traversal of the DOM tree to extract readable paragraphs while preserving semantic structure:

// src/entrypoints/side.content/utils/article.ts
export function flattenToParagraphs(root: Node) {
  const semanticBlocks = new Set(['P', 'ARTICLE', 'SECTION', 'DIV', 'MAIN', 'BLOCKQUOTE'])
  
  // Helper functions:
  // isBlockLevel - checks tag name or computed display style
  // hasBlockDescendant - verifies if node contains nested blocks
  // getTextWithSpaces - recursively concatenates text nodes with spacing
  
  // Returns array of cleaned paragraph strings (≥ 20 chars)
}

The algorithm identifies leaf block elements—block-level containers that do not contain other block elements—and extracts their text content. It handles inline elements by inserting spaces between them to prevent word concatenation, then filters out paragraphs shorter than 20 characters to remove navigation cruft and decorative elements.

SEO Metadata Extraction

The extractSeoInfo function gathers comprehensive metadata for downstream processing such as summarization or translation:

export function extractSeoInfo(doc: Document) {
  return {
    title: doc.title,
    metaDescription: doc.querySelector('meta[name="description"]')?.getAttribute('content') ?? '',
    metaKeywords: doc.querySelector('meta[name="keywords"]')?.getAttribute('content') ?? '',
    canonicalUrl: doc.querySelector('link[rel="canonical"]')?.getAttribute('href') ?? '',
    ogTitle: doc.querySelector('meta[property="og:title"]')?.getAttribute('content') ?? '',
    ogDescription: doc.querySelector('meta[property="og:description"]')?.getAttribute('content') ?? '',
    h1Tags: Array.from(doc.querySelectorAll('h1')).map(h => h.textContent?.trim() ?? ''),
    structuredData: Array.from(
      doc.querySelectorAll('script[type="application/ld+json"]')
    ).map(script => {
      try { return JSON.parse(script.textContent || '{}') }
      catch { return {} }
    })
  }
}

This extracts standard meta tags, Open Graph properties, Twitter card data, H1 headings, and JSON-LD structured data schemas.

Side Panel UI Components

The visual interface is implemented in src/entrypoints/side.content/components/side-content/index.tsx, featuring a resizable panel that slides in from the right edge of the viewport.

Resizable Panel Implementation

The component uses mouse event listeners to calculate new widths based on cursor position, storing the value in a Jotai atom (sideContent via configFieldsAtomMap.sideContent):

useEffect(() => {
  if (!isResizing) return
  const handleMouseMove = (e: MouseEvent) => {
    const newWidth = Math.max(MIN_SIDE_CONTENT_WIDTH, window.innerWidth - e.clientX)
    void setSideContent({ width: newWidth })
  }
  const handleMouseUp = () => setIsResizing(false)
  
  document.addEventListener('mousemove', handleMouseMove)
  document.addEventListener('mouseup', handleMouseUp)
  document.body.style.userSelect = 'none'
  
  return () => {
    document.removeEventListener('mousemove', handleMouseMove)
    document.removeEventListener('mouseup', handleMouseUp)
    document.body.style.userSelect = ''
  }
}, [isResizing, setSideContent])

The resize handle sits at the left edge of the panel with cursor-ew-resize, while a transparent full-screen overlay (z-[2147483647]) captures mouse events during dragging to prevent text selection on the underlying page.

Page Width Adjustment

To prevent content from scrolling underneath the panel, the component dynamically injects a style tag into the document head when isSideOpenAtom is true:

useEffect(() => {
  const styleId = `shrink-origin-for-${kebabCase(APP_NAME)}-side-content`
  let styleTag = document.getElementById(styleId)
  
  if (isSideOpen) {
    if (!styleTag) {
      styleTag = document.createElement('style')
      styleTag.id = styleId
      document.head.appendChild(styleTag)
    }
    styleTag.textContent = `
      html { width: calc(100% - ${sideContent.width}px) !important; position: relative !important; min-height: 100vh !important; }
    `
  } else if (styleTag) {
    document.head.removeChild(styleTag)
  }
}, [isSideOpen, sideContent.width])

This reduces the effective viewport width by the panel width, forcing the host page to reflow naturally alongside the extension UI.

Application Orchestration

The App component in src/entrypoints/side.content/app.tsx serves as the orchestration layer, triggering article extraction when the user opens the panel:

export default function App() {
  const [article, setArticle] = useState<string[]>([])
  const isSideOpen = useAtomValue(isSideOpenAtom)

  useEffect(() => {
    if (!isSideOpen) return
    const paragraphs = flattenToParagraphs(document.body)
    const seo = extractSeoInfo(document)
    setArticle(paragraphs)
    // Data flows to downstream UI components
  }, [isSideOpen])

  return (
    <>
      <SideContent />
    </>
  )
}

This lazy-loads content only when the panel is first activated, improving initial page load performance.

Configuration and State Management

The content script relies on several utility modules for state persistence:

These utilities ensure the content script remains thin and reusable across other entry points like the subtitle overlay or selection toolbar.

Summary

  • The Read Article feature runs as a content script in src/entrypoints/side.content/index.tsx that creates an isolated shadow-DOM overlay on every webpage.
  • Style isolation is achieved through createShadowRootUi combined with addStyleToShadow and mirrorDynamicStyles, preventing CSS leakage between the extension and host page.
  • Article extraction uses flattenToParagraphs in src/entrypoints/side.content/utils/article.ts to perform depth-first DOM traversal, identifying leaf block elements and cleaning text content.
  • SEO metadata collection via extractSeoInfo gathers standard meta tags, Open Graph data, H1 elements, and JSON-LD structured data.
  • The side panel UI in components/side-content/index.tsx implements mouse-based resizing with dynamic page-width adjustment to prevent layout conflicts.
  • State management uses Jotai atoms hydrated from Chrome extension storage, allowing persistent user preferences across sessions.

Frequently Asked Questions

How does the content script isolate its UI from the host page?

The script uses shadow DOM encapsulation via createShadowRootUi from src/utils/shadow-root.ts. This attaches a closed shadow root to a host element injected into the page body, creating a style boundary that prevents the host page's CSS from affecting the extension UI and vice versa. Additional protections like protectInternalStyles guard against aggressive page styles that might otherwise pierce the shadow boundary.

What algorithm extracts article text from web pages?

The flattenToParagraphs function in src/entrypoints/side.content/utils/article.ts implements a recursive depth-first traversal that identifies block-level elements without nested block children (leaf blocks). It extracts text content from these nodes while preserving spacing between inline elements, then filters results to include only paragraphs with 20 or more characters, effectively removing navigation menus and decorative text.

How does the side panel handle resizing without breaking page layout?

The SideContent component tracks mouse movements to calculate new widths stored in Jotai atoms. Simultaneously, it injects a dynamic style tag into the document head that sets html { width: calc(100% - ${width}px) } when open. This forces the host page to reflow within the remaining viewport space, preventing horizontal scrollbars and content occlusion.

Where is the extension configuration stored and how is it accessed?

Configuration persists in Chrome extension storage via utilities in src/utils/config/storage.ts. The getLocalConfig() function retrieves settings asynchronously during script initialization in index.tsx, while useHydrateAtoms injects these values into the Jotai state store. The isSiteEnabled utility then checks the current URL against these preferences to determine whether the side panel should activate.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →