Site-Specific Extractors in Hister: Complete Guide to Built-In Parsers
Hister provides 11 built-in site-specific extractors for platforms including Reddit, GitHub, Twitter, Mastodon, Bluesky, Wikipedia, Stack Exchange, Discourse, Lobsters, GoDoc, and Notion, each implementing the Extractor interface defined in server/extractor/sdk/sdk.go and registered in the ordered chain within server/extractor/registry.go.
Hister discovers and renders web content by executing a prioritized chain of extractors. These site-specific extractors transform messy HTML into clean, structured previews tailored to individual platforms, while the registry manages selection order and fallback behavior.
How Site-Specific Extractors Work
Each extractor implements the stable contract defined in server/extractor/sdk/sdk.go. The registry walks the extractor list in order, calling Match(*sdk.Document) bool on each until one returns true. The first matching extractor then performs the work via its Extract or Preview methods, returning sdk.ExtractResult or sdk.PreviewResponse respectively. If no site-specific extractor matches, Hister falls back to basicExtractor (strips HTML tags) or readabilityExtractor (algorithmic preview).
The DefaultExtractors slice in server/extractor/registry.go (lines 56-74) builds the ordered chain of fresh extractor instances. All extractors share common utilities including urlutil for hostname validation and sanitizer for safe HTML output.
Available Site-Specific Extractors in Hister
The following extractors target specific web services and are consulted before generic fallback extractors:
- DiscourseExtractor (
server/extractor/extractors/discourse/discourse.go) — Parses posts and threads from Discourse-based forums. - RedditExtractor (
server/extractor/extractors/reddit/reddit.go) — Extracts Reddit posts, comments, and metadata. - StackExchangeExtractor (
server/extractor/extractors/stackexchange/stackexchange.go) — Handles Stack Exchange question pages and answers. - GoDocExtractor (
server/extractor/extractors/godoc/godoc.go) — Retrieves Go package documentation frompkg.go.dev. - GitHubExtractor (
server/extractor/extractors/github/github.go) — Scrapes repository pages for README content, descriptions, and stats. - LobstersExtractor (
server/extractor/extractors/lobsters/lobsters.go) — Extracts content from the Lobsters link-aggregation site. - WikipediaExtractor (
server/extractor/extractors/wikipedia/wikipedia.go) — Retrieves article excerpts and structured data from Wikipedia. - MastodonExtractor (
server/extractor/extractors/mastodon/mastodon.go) — Parses posts from Mastodon instances. - BlueskyExtractor (
server/extractor/extractors/bluesky/bluesky.go) — Extracts posts from Bluesky (bsky.app). - TwitterExtractor (
server/extractor/extractors/twitter/twitter.go) — Handles tweet rendering and thread extraction. - NotionExtractor (
server/extractor/extractors/notion/notion.go) — Pulls content from public Notion pages.
Each extractor implements both Extract and Preview capabilities, allowing Hister to either retrieve raw structured data or generate formatted HTML previews.
Requesting Specific Extractors via the MCP API
You can bypass automatic detection and force Hister to use a specific site extractor by providing the Extractor field in your request.
package main
import (
"context"
"fmt"
"github.com/asciimoo/hister/client"
)
func main() {
c := client.NewClient("http://localhost:8080")
preview, err := c.GetPreview(context.TODO(), client.GetPreviewArgs{
URL: "https://reddit.com/r/golang/comments/xyz/example_post",
Extractor: "Reddit", // force the Reddit extractor
})
if err != nil {
panic(err)
}
fmt.Println("Title:", preview.Title)
fmt.Println("HTML preview:", preview.HTML)
}
The Extractor string must match the extractor's registered name (e.g., "Reddit", "GitHub", "Wikipedia").
Listing All Registered Extractors Programmatically
To discover which extractors are available in your Hister instance, iterate over the default registry:
package main
import (
"fmt"
"github.com/asciimoo/hister/server/extractor"
)
func main() {
for _, e := range extractor.DefaultRegistry().Extractors() {
fmt.Printf("%s – %s (capabilities: %+v)\n",
e.Name(), e.Description(), e.Capabilities())
}
}
This prints every registered extractor, including all 11 site-specific implementations and generic fallback extractors like MarkdownExtractor and YtdlpExtractor.
Summary
- Eleven site-specific extractors ship with Hister, covering major platforms from GitHub to Mastodon to Wikipedia.
- Registration occurs in
server/extractor/registry.goviaDefaultExtractors, which establishes the evaluation order. - Selection relies on the
Match(*sdk.Document) boolmethod, where each extractor implements URL pattern or hostname checks usingurlutil. - Capabilities include both
Extract(structured data retrieval) andPreview(HTML generation), defined inserver/extractor/sdk/sdk.go. - Fallback behavior triggers when no site-specific extractor matches, using
basicExtractororreadabilityExtractorto handle generic HTML.
Frequently Asked Questions
How do I force Hister to use a specific extractor instead of auto-detection?
Pass the Extractor field in your GetPreviewArgs or extraction request, setting it to the exact name of the extractor (e.g., "Twitter" or "GitHub"). This bypasses the registry's ordered Match chain and invokes the specified extractor directly.
What happens if no site-specific extractor matches the URL?
If all 11 site-specific extractors return false from their Match methods, Hister falls back to generic extractors. First, it attempts readabilityExtractor for algorithmic content extraction; if that fails or is unavailable, basicExtractor strips HTML tags and returns plain text.
Can I add custom site-specific extractors to Hister?
Yes. Implement the Extractor interface from server/extractor/sdk/sdk.go, providing Match, Extract, and Preview methods. Register your implementation in the DefaultExtractors slice within server/extractor/registry.go (around lines 56-74) to include it in the matching chain.
Which extractor handles Reddit versus Stack Exchange?
RedditExtractor (server/extractor/extractors/reddit/reddit.go) handles Reddit URLs, while StackExchangeExtractor (server/extractor/extractors/stackexchange/stackexchange.go) manages Stack Exchange network sites. Both support full thread extraction and comment parsing, but use distinct HTML parsing logic tailored to each platform's DOM structure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →