Site-Specific Extractors in Hister: Complete Guide to Built-In Parsers

Hister provides 11 built-in site-specific extractors for platforms including Reddit, GitHub, Twitter, Mastodon, Bluesky, Wikipedia, Stack Exchange, Discourse, Lobsters, GoDoc, and Notion, each implementing the Extractor interface defined in server/extractor/sdk/sdk.go and registered in the ordered chain within server/extractor/registry.go.

Hister discovers and renders web content by executing a prioritized chain of extractors. These site-specific extractors transform messy HTML into clean, structured previews tailored to individual platforms, while the registry manages selection order and fallback behavior.

How Site-Specific Extractors Work

Each extractor implements the stable contract defined in server/extractor/sdk/sdk.go. The registry walks the extractor list in order, calling Match(*sdk.Document) bool on each until one returns true. The first matching extractor then performs the work via its Extract or Preview methods, returning sdk.ExtractResult or sdk.PreviewResponse respectively. If no site-specific extractor matches, Hister falls back to basicExtractor (strips HTML tags) or readabilityExtractor (algorithmic preview).

The DefaultExtractors slice in server/extractor/registry.go (lines 56-74) builds the ordered chain of fresh extractor instances. All extractors share common utilities including urlutil for hostname validation and sanitizer for safe HTML output.

Available Site-Specific Extractors in Hister

The following extractors target specific web services and are consulted before generic fallback extractors:

Each extractor implements both Extract and Preview capabilities, allowing Hister to either retrieve raw structured data or generate formatted HTML previews.

Requesting Specific Extractors via the MCP API

You can bypass automatic detection and force Hister to use a specific site extractor by providing the Extractor field in your request.

package main

import (
	"context"
	"fmt"
	"github.com/asciimoo/hister/client"
)

func main() {
	c := client.NewClient("http://localhost:8080")
	preview, err := c.GetPreview(context.TODO(), client.GetPreviewArgs{
		URL:       "https://reddit.com/r/golang/comments/xyz/example_post",
		Extractor: "Reddit", // force the Reddit extractor
	})
	if err != nil {
		panic(err)
	}
	fmt.Println("Title:", preview.Title)
	fmt.Println("HTML preview:", preview.HTML)
}

The Extractor string must match the extractor's registered name (e.g., "Reddit", "GitHub", "Wikipedia").

Listing All Registered Extractors Programmatically

To discover which extractors are available in your Hister instance, iterate over the default registry:

package main

import (
	"fmt"
	"github.com/asciimoo/hister/server/extractor"
)

func main() {
	for _, e := range extractor.DefaultRegistry().Extractors() {
		fmt.Printf("%s – %s (capabilities: %+v)\n",
			e.Name(), e.Description(), e.Capabilities())
	}
}

This prints every registered extractor, including all 11 site-specific implementations and generic fallback extractors like MarkdownExtractor and YtdlpExtractor.

Summary

  • Eleven site-specific extractors ship with Hister, covering major platforms from GitHub to Mastodon to Wikipedia.
  • Registration occurs in server/extractor/registry.go via DefaultExtractors, which establishes the evaluation order.
  • Selection relies on the Match(*sdk.Document) bool method, where each extractor implements URL pattern or hostname checks using urlutil.
  • Capabilities include both Extract (structured data retrieval) and Preview (HTML generation), defined in server/extractor/sdk/sdk.go.
  • Fallback behavior triggers when no site-specific extractor matches, using basicExtractor or readabilityExtractor to handle generic HTML.

Frequently Asked Questions

How do I force Hister to use a specific extractor instead of auto-detection?

Pass the Extractor field in your GetPreviewArgs or extraction request, setting it to the exact name of the extractor (e.g., "Twitter" or "GitHub"). This bypasses the registry's ordered Match chain and invokes the specified extractor directly.

What happens if no site-specific extractor matches the URL?

If all 11 site-specific extractors return false from their Match methods, Hister falls back to generic extractors. First, it attempts readabilityExtractor for algorithmic content extraction; if that fails or is unavailable, basicExtractor strips HTML tags and returns plain text.

Can I add custom site-specific extractors to Hister?

Yes. Implement the Extractor interface from server/extractor/sdk/sdk.go, providing Match, Extract, and Preview methods. Register your implementation in the DefaultExtractors slice within server/extractor/registry.go (around lines 56-74) to include it in the matching chain.

Which extractor handles Reddit versus Stack Exchange?

RedditExtractor (server/extractor/extractors/reddit/reddit.go) handles Reddit URLs, while StackExchangeExtractor (server/extractor/extractors/stackexchange/stackexchange.go) manages Stack Exchange network sites. Both support full thread extraction and comment parsing, but use distinct HTML parsing logic tailored to each platform's DOM structure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →