How Does Hister’s Browser Extension Work: A Deep Dive Into the Source Code
Hister’s browser extension captures web page content via a content script, forwards it to a self-hosted server for indexing through a cookie-aware network layer, and manages per-tab UI state and skip rules in a background service worker.
The asciimoo/hister repository contains a self-hosted search engine with a companion browser extension that continuously indexes visited pages. Understanding how Hister’s browser extension works requires examining its four core components that handle everything from DOM extraction to authenticated API communication.
Architecture Overview
The extension is built for Manifest V3 compatibility across Chrome and Firefox. Its architecture separates concerns into distinct modules:
webui/ext/src/manifest.json– Declares permissions, entry points, and keyboard shortcutswebui/ext/src/content/content.ts– Runs in every tab to extract page data and monitor navigationwebui/ext/src/background/background.ts– Service worker that manages server communication and UI statewebui/ext/src/modules/network.ts– Handles cookie-aware HTTP requests and authentication refresh
The Manifest File
The bootstrap configuration in webui/ext/src/manifest.json registers the extension’s capabilities. It declares manifest_version: 3 and requests permissions for tabs, storage, and cookies necessary for cross-origin authentication forwarding.
The manifest defines three keyboard commands mapped to specific actions:
Ctrl+I– Forces immediate indexing of the current page viaindexCurrentTab()Ctrl+B– Disables indexing for the current URL patternCtrl+Y– Disables indexing for the current domain
It registers content.js to match <all_urls> and loads the background logic as a module from webui/ext/src/background/background.ts.
Content Script Data Extraction
The webui/ext/src/content/content.ts script executes in every browser tab. When a page finishes loading, it invokes the extract() function to harvest data and forward it to the background worker.
function extract(sendResponse, actionType, force) {
// Guard against unsupported content types
if (!isSupportedContentType(document.contentType)) return;
// Register callback for search result interactions
registerResultExtractor(window, (r) => {
if (isContextValid()) chrome.runtime.sendMessage({ resultData: r });
});
// Extract HTML, title, URL, and cookies
const d = extractPageData();
// Forward to background script
chrome.runtime.sendMessage(
{ pageData: d, action: actionType },
(resp) => {
// Schedule retry if server rejects the page
scheduleUpdate();
}
);
}
The content script monitors Single Page Application (SPA) navigation through the window.navigation API or falls back to timer-based polling via scheduleUpdate(). It also listens for visibilitychange events to submit final snapshots before tabs are hidden.
When the server returns HTTP 406 (Not Acceptable), the script caches the URL in skippedUrl to prevent unnecessary resubmissions.
Background Script Orchestration
The webui/ext/src/background/background.ts service worker acts as the central hub. It receives messages from content scripts, maintains state, and updates the browser chrome UI.
UI State Management
The background script manipulates toolbar icons and badges to reflect indexing status:
setNormalIcon()– Indicates successful indexingsetGreyIcon()– Shows when a page is skipped due to rulessetErrorBadge()– Displays!when server communication failsclearBadge()– Removes status indicators
Skip-Rule Caching
The extension fetches regex-based skip rules from /api/rules and caches them locally with a 60-second TTL (SKIP_RULES_TTL). When settings change, the background script invalidates this cache and fetches fresh rules.
PDF Handling
For PDF URLs detected via isPDFUrl(), the background script downloads the file, base-64 encodes the content, and transmits it to /api/add_pdf through the sendPDFData() helper.
Command Handling
The chrome.commands.onCommand listener maps keyboard shortcuts to specific functions like indexCurrentTab() and disableIndexingForCurrentTab(), allowing users to override skip rules or exclude domains on demand.
Network Layer and Authentication
All server communication flows through webui/ext/src/modules/network.ts. The fetchAPI() wrapper automatically attaches stored authentication credentials and handles token refresh logic.
// Simplified authentication flow
async function fetchAPI(endpoint, options = {}) {
const cookies = await chrome.storage.local.get(['histerCookies', 'histerToken']);
const response = await fetch(serverURL + endpoint, {
...options,
headers: {
...options.headers,
'Cookie': cookies.histerCookies,
'Authorization': `Bearer ${cookies.histerToken}`
}
});
// Refresh cookies on 401/403 and retry
if (response.status === 401 || response.status === 403) {
await refreshServerCookies();
return fetchAPI(endpoint, options); // Retry with fresh auth
}
return response;
}
The sendPageData() function enriches payloads with favicon data fetched via fetchFavicon() before posting to /api/add. For PDF documents, sendPDFData() constructs a payload containing the base-64 encoded file and custom headers.
End-to-End Indexing Flow
A typical user session follows this sequence:
- Page Load – The content script executes
extract()and forwards data to the background worker - Rule Check – The background script validates the URL against cached skip rules
- Transmission –
sendPageData()POSTs the payload to<server>/api/addwith authentication headers - Success – On HTTP 201 (Created), the background calls
setNormalIcon()and displays an indexed badge - Skip – On HTTP 406, the background invokes
setGreyIcon()and caches the skip rule - Override – User presses Ctrl+I to trigger
indexCurrentTab(), bypassing cached skip rules
Summary
- Hister’s browser extension uses a content script (
content.ts) to extract page data and monitor SPA navigation - The background service worker (
background.ts) manages state, skip-rule caching, and UI feedback through icons and badges - Authentication is handled transparently in
network.ts, which refreshes cookies automatically on 401/403 responses - Keyboard shortcuts (
Ctrl+I,Ctrl+B,Ctrl+Y) allow users to force indexing or disable it for specific URLs and domains - All configuration persists in
chrome.storage.local, enabling offline state management and automatic retry logic
Frequently Asked Questions
How does the extension handle Single Page Applications?
The content script in webui/ext/src/content/content.ts detects SPA navigation through the window.navigation API when available. For browsers without native support, it falls back to timer-based polling via scheduleUpdate() to capture URL changes and submit updated content to the server.
Can I programmatically trigger indexing from my own extension?
Yes. You can send a message to the active tab to request re-indexing:
chrome.tabs.query({ active: true, currentWindow: true }, (tabs) => {
const tabId = tabs[0]?.id;
if (tabId) {
chrome.tabs.sendMessage(tabId, { action: 'reindex' });
}
});
This invokes the extract() function in the content script, which forwards the request to the background worker and subsequently to the Hister server.
Where does the extension store authentication credentials?
The extension stores cookies and bearer tokens in chrome.storage.local under the keys histerCookies and histerToken. The fetchAPI() function in webui/ext/src/modules/network.ts retrieves these values for each request and automatically refreshes them when the server returns unauthorized status codes.
How do skip rules work and how long are they cached?
Skip rules are regex patterns fetched from /api/rules and cached in memory for 60 seconds (SKIP_RULES_TTL). The background script checks URLs against these patterns before sending data to the server. When a URL matches a skip rule, the extension displays a grey icon and avoids resubmitting the page until the cache expires or the user forces indexing with Ctrl+I.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →