# Site-Specific Extractors in Hister: Complete Guide to Built-In Parsers

> Explore Hister's 11 site-specific extractors for Reddit, GitHub, Twitter, and more. Integrate seamlessly with platforms using built-in parsers for efficient data extraction.

- Repository: [Adam Tauber/hister](https://github.com/asciimoo/hister)
- Tags: how-to-guide
- Published: 2026-08-27

---

**Hister provides 11 built-in site-specific extractors for platforms including Reddit, GitHub, Twitter, Mastodon, Bluesky, Wikipedia, Stack Exchange, Discourse, Lobsters, GoDoc, and Notion, each implementing the `Extractor` interface defined in [`server/extractor/sdk/sdk.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/sdk/sdk.go) and registered in the ordered chain within [`server/extractor/registry.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/registry.go).**

Hister discovers and renders web content by executing a prioritized chain of extractors. These site-specific extractors transform messy HTML into clean, structured previews tailored to individual platforms, while the registry manages selection order and fallback behavior.

## How Site-Specific Extractors Work

Each extractor implements the stable contract defined in [`server/extractor/sdk/sdk.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/sdk/sdk.go). The registry walks the extractor list in order, calling `Match(*sdk.Document) bool` on each until one returns `true`. The first matching extractor then performs the work via its `Extract` or `Preview` methods, returning `sdk.ExtractResult` or `sdk.PreviewResponse` respectively. If no site-specific extractor matches, Hister falls back to `basicExtractor` (strips HTML tags) or `readabilityExtractor` (algorithmic preview).

The `DefaultExtractors` slice in [`server/extractor/registry.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/registry.go) (lines 56-74) builds the ordered chain of fresh extractor instances. All extractors share common utilities including `urlutil` for hostname validation and `sanitizer` for safe HTML output.

## Available Site-Specific Extractors in Hister

The following extractors target specific web services and are consulted before generic fallback extractors:

- **DiscourseExtractor** ([`server/extractor/extractors/discourse/discourse.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/discourse/discourse.go)) — Parses posts and threads from Discourse-based forums.
- **RedditExtractor** ([`server/extractor/extractors/reddit/reddit.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/reddit/reddit.go)) — Extracts Reddit posts, comments, and metadata.
- **StackExchangeExtractor** ([`server/extractor/extractors/stackexchange/stackexchange.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/stackexchange/stackexchange.go)) — Handles Stack Exchange question pages and answers.
- **GoDocExtractor** ([`server/extractor/extractors/godoc/godoc.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/godoc/godoc.go)) — Retrieves Go package documentation from `pkg.go.dev`.
- **GitHubExtractor** ([`server/extractor/extractors/github/github.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/github/github.go)) — Scrapes repository pages for README content, descriptions, and stats.
- **LobstersExtractor** ([`server/extractor/extractors/lobsters/lobsters.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/lobsters/lobsters.go)) — Extracts content from the Lobsters link-aggregation site.
- **WikipediaExtractor** ([`server/extractor/extractors/wikipedia/wikipedia.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/wikipedia/wikipedia.go)) — Retrieves article excerpts and structured data from Wikipedia.
- **MastodonExtractor** ([`server/extractor/extractors/mastodon/mastodon.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/mastodon/mastodon.go)) — Parses posts from Mastodon instances.
- **BlueskyExtractor** ([`server/extractor/extractors/bluesky/bluesky.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/bluesky/bluesky.go)) — Extracts posts from Bluesky (`bsky.app`).
- **TwitterExtractor** ([`server/extractor/extractors/twitter/twitter.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/twitter/twitter.go)) — Handles tweet rendering and thread extraction.
- **NotionExtractor** ([`server/extractor/extractors/notion/notion.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/notion/notion.go)) — Pulls content from public Notion pages.

Each extractor implements both `Extract` and `Preview` capabilities, allowing Hister to either retrieve raw structured data or generate formatted HTML previews.

## Requesting Specific Extractors via the MCP API

You can bypass automatic detection and force Hister to use a specific site extractor by providing the `Extractor` field in your request.

```go
package main

import (
	"context"
	"fmt"
	"github.com/asciimoo/hister/client"
)

func main() {
	c := client.NewClient("http://localhost:8080")
	preview, err := c.GetPreview(context.TODO(), client.GetPreviewArgs{
		URL:       "https://reddit.com/r/golang/comments/xyz/example_post",
		Extractor: "Reddit", // force the Reddit extractor
	})
	if err != nil {
		panic(err)
	}
	fmt.Println("Title:", preview.Title)
	fmt.Println("HTML preview:", preview.HTML)
}

```

The `Extractor` string must match the extractor's registered name (e.g., `"Reddit"`, `"GitHub"`, `"Wikipedia"`).

## Listing All Registered Extractors Programmatically

To discover which extractors are available in your Hister instance, iterate over the default registry:

```go
package main

import (
	"fmt"
	"github.com/asciimoo/hister/server/extractor"
)

func main() {
	for _, e := range extractor.DefaultRegistry().Extractors() {
		fmt.Printf("%s – %s (capabilities: %+v)\n",
			e.Name(), e.Description(), e.Capabilities())
	}
}

```

This prints every registered extractor, including all 11 site-specific implementations and generic fallback extractors like `MarkdownExtractor` and `YtdlpExtractor`.

## Summary

- **Eleven site-specific extractors** ship with Hister, covering major platforms from GitHub to Mastodon to Wikipedia.
- **Registration occurs** in [`server/extractor/registry.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/registry.go) via `DefaultExtractors`, which establishes the evaluation order.
- **Selection relies on** the `Match(*sdk.Document) bool` method, where each extractor implements URL pattern or hostname checks using `urlutil`.
- **Capabilities** include both `Extract` (structured data retrieval) and `Preview` (HTML generation), defined in [`server/extractor/sdk/sdk.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/sdk/sdk.go).
- **Fallback behavior** triggers when no site-specific extractor matches, using `basicExtractor` or `readabilityExtractor` to handle generic HTML.

## Frequently Asked Questions

### How do I force Hister to use a specific extractor instead of auto-detection?

Pass the `Extractor` field in your `GetPreviewArgs` or extraction request, setting it to the exact name of the extractor (e.g., `"Twitter"` or `"GitHub"`). This bypasses the registry's ordered `Match` chain and invokes the specified extractor directly.

### What happens if no site-specific extractor matches the URL?

If all 11 site-specific extractors return `false` from their `Match` methods, Hister falls back to generic extractors. First, it attempts `readabilityExtractor` for algorithmic content extraction; if that fails or is unavailable, `basicExtractor` strips HTML tags and returns plain text.

### Can I add custom site-specific extractors to Hister?

Yes. Implement the `Extractor` interface from [`server/extractor/sdk/sdk.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/sdk/sdk.go), providing `Match`, `Extract`, and `Preview` methods. Register your implementation in the `DefaultExtractors` slice within [`server/extractor/registry.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/registry.go) (around lines 56-74) to include it in the matching chain.

### Which extractor handles Reddit versus Stack Exchange?

**RedditExtractor** ([`server/extractor/extractors/reddit/reddit.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/reddit/reddit.go)) handles Reddit URLs, while **StackExchangeExtractor** ([`server/extractor/extractors/stackexchange/stackexchange.go`](https://github.com/asciimoo/hister/blob/main/server/extractor/extractors/stackexchange/stackexchange.go)) manages Stack Exchange network sites. Both support full thread extraction and comment parsing, but use distinct HTML parsing logic tailored to each platform's DOM structure.