How Does Hister’s Browser Extension Work: A Deep Dive Into the Source Code

Hister’s browser extension captures web page content via a content script, forwards it to a self-hosted server for indexing through a cookie-aware network layer, and manages per-tab UI state and skip rules in a background service worker.

The asciimoo/hister repository contains a self-hosted search engine with a companion browser extension that continuously indexes visited pages. Understanding how Hister’s browser extension works requires examining its four core components that handle everything from DOM extraction to authenticated API communication.

Architecture Overview

The extension is built for Manifest V3 compatibility across Chrome and Firefox. Its architecture separates concerns into distinct modules:

The Manifest File

The bootstrap configuration in webui/ext/src/manifest.json registers the extension’s capabilities. It declares manifest_version: 3 and requests permissions for tabs, storage, and cookies necessary for cross-origin authentication forwarding.

The manifest defines three keyboard commands mapped to specific actions:

  • Ctrl+I – Forces immediate indexing of the current page via indexCurrentTab()
  • Ctrl+B – Disables indexing for the current URL pattern
  • Ctrl+Y – Disables indexing for the current domain

It registers content.js to match <all_urls> and loads the background logic as a module from webui/ext/src/background/background.ts.

Content Script Data Extraction

The webui/ext/src/content/content.ts script executes in every browser tab. When a page finishes loading, it invokes the extract() function to harvest data and forward it to the background worker.

function extract(sendResponse, actionType, force) {
  // Guard against unsupported content types
  if (!isSupportedContentType(document.contentType)) return;

  // Register callback for search result interactions
  registerResultExtractor(window, (r) => {
    if (isContextValid()) chrome.runtime.sendMessage({ resultData: r });
  });

  // Extract HTML, title, URL, and cookies
  const d = extractPageData();

  // Forward to background script
  chrome.runtime.sendMessage(
    { pageData: d, action: actionType },
    (resp) => {
      // Schedule retry if server rejects the page
      scheduleUpdate();
    }
  );
}

The content script monitors Single Page Application (SPA) navigation through the window.navigation API or falls back to timer-based polling via scheduleUpdate(). It also listens for visibilitychange events to submit final snapshots before tabs are hidden.

When the server returns HTTP 406 (Not Acceptable), the script caches the URL in skippedUrl to prevent unnecessary resubmissions.

Background Script Orchestration

The webui/ext/src/background/background.ts service worker acts as the central hub. It receives messages from content scripts, maintains state, and updates the browser chrome UI.

UI State Management

The background script manipulates toolbar icons and badges to reflect indexing status:

  • setNormalIcon() – Indicates successful indexing
  • setGreyIcon() – Shows when a page is skipped due to rules
  • setErrorBadge() – Displays ! when server communication fails
  • clearBadge() – Removes status indicators

Skip-Rule Caching

The extension fetches regex-based skip rules from /api/rules and caches them locally with a 60-second TTL (SKIP_RULES_TTL). When settings change, the background script invalidates this cache and fetches fresh rules.

PDF Handling

For PDF URLs detected via isPDFUrl(), the background script downloads the file, base-64 encodes the content, and transmits it to /api/add_pdf through the sendPDFData() helper.

Command Handling

The chrome.commands.onCommand listener maps keyboard shortcuts to specific functions like indexCurrentTab() and disableIndexingForCurrentTab(), allowing users to override skip rules or exclude domains on demand.

Network Layer and Authentication

All server communication flows through webui/ext/src/modules/network.ts. The fetchAPI() wrapper automatically attaches stored authentication credentials and handles token refresh logic.

// Simplified authentication flow
async function fetchAPI(endpoint, options = {}) {
  const cookies = await chrome.storage.local.get(['histerCookies', 'histerToken']);
  
  const response = await fetch(serverURL + endpoint, {
    ...options,
    headers: {
      ...options.headers,
      'Cookie': cookies.histerCookies,
      'Authorization': `Bearer ${cookies.histerToken}`
    }
  });

  // Refresh cookies on 401/403 and retry
  if (response.status === 401 || response.status === 403) {
    await refreshServerCookies();
    return fetchAPI(endpoint, options); // Retry with fresh auth
  }
  
  return response;
}

The sendPageData() function enriches payloads with favicon data fetched via fetchFavicon() before posting to /api/add. For PDF documents, sendPDFData() constructs a payload containing the base-64 encoded file and custom headers.

End-to-End Indexing Flow

A typical user session follows this sequence:

  1. Page Load – The content script executes extract() and forwards data to the background worker
  2. Rule Check – The background script validates the URL against cached skip rules
  3. Transmission – sendPageData() POSTs the payload to <server>/api/add with authentication headers
  4. Success – On HTTP 201 (Created), the background calls setNormalIcon() and displays an indexed badge
  5. Skip – On HTTP 406, the background invokes setGreyIcon() and caches the skip rule
  6. Override – User presses Ctrl+I to trigger indexCurrentTab(), bypassing cached skip rules

Summary

  • Hister’s browser extension uses a content script (content.ts) to extract page data and monitor SPA navigation
  • The background service worker (background.ts) manages state, skip-rule caching, and UI feedback through icons and badges
  • Authentication is handled transparently in network.ts, which refreshes cookies automatically on 401/403 responses
  • Keyboard shortcuts (Ctrl+I, Ctrl+B, Ctrl+Y) allow users to force indexing or disable it for specific URLs and domains
  • All configuration persists in chrome.storage.local, enabling offline state management and automatic retry logic

Frequently Asked Questions

How does the extension handle Single Page Applications?

The content script in webui/ext/src/content/content.ts detects SPA navigation through the window.navigation API when available. For browsers without native support, it falls back to timer-based polling via scheduleUpdate() to capture URL changes and submit updated content to the server.

Can I programmatically trigger indexing from my own extension?

Yes. You can send a message to the active tab to request re-indexing:

chrome.tabs.query({ active: true, currentWindow: true }, (tabs) => {
  const tabId = tabs[0]?.id;
  if (tabId) {
    chrome.tabs.sendMessage(tabId, { action: 'reindex' });
  }
});

This invokes the extract() function in the content script, which forwards the request to the background worker and subsequently to the Hister server.

Where does the extension store authentication credentials?

The extension stores cookies and bearer tokens in chrome.storage.local under the keys histerCookies and histerToken. The fetchAPI() function in webui/ext/src/modules/network.ts retrieves these values for each request and automatically refreshes them when the server returns unauthorized status codes.

How do skip rules work and how long are they cached?

Skip rules are regex patterns fetched from /api/rules and cached in memory for 60 seconds (SKIP_RULES_TTL). The background script checks URLs against these patterns before sending data to the server. When a URL matches a skip rule, the extension displays a grey icon and avoids resubmitting the page until the cache expires or the user forces indexing with Ctrl+I.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →