# How Hister Handles Content Extraction and Language Detection

> Discover how Hister extracts content using goquery and detects language with lingua-go for efficient, language-specific indexing. Learn more about its advanced capabilities.

- Repository: [Adam Tauber/hister](https://github.com/asciimoo/hister)
- Tags: how-to-guide
- Published: 2026-09-01

---

**Hister extracts canonical URLs and titles from HTML using `goquery` in [`server/document/fromhtml.go`](https://github.com/asciimoo/hister/blob/main/server/document/fromhtml.go), then detects document language via the `lingua-go` library in [`server/document/language.go`](https://github.com/asciimoo/hister/blob/main/server/document/language.go) to enable language-specific Bleve indexing.**

Hister is an open-source search engine that transforms raw HTML into searchable documents. Understanding how it handles **content extraction and language detection** is essential for configuring multilingual indexing pipelines. The implementation spans dedicated modules in the `server/document` and `server/indexer` packages, orchestrating metadata parsing and linguistic analysis.

## Content Extraction from HTML

The core extraction logic resides in **[`server/document/fromhtml.go`](https://github.com/asciimoo/hister/blob/main/server/document/fromhtml.go)**, where the `FromHTML` function parses HTML strings and populates a `*Document` struct with normalized metadata.

### Parsing HTML with goquery

Hister relies on the **`goquery`** library to navigate and query HTML DOM structures. The `FromHTML` function accepts an HTML string and systematically extracts critical metadata fields required for indexing.

### URL Resolution Strategy

Hister implements a hierarchical fallback strategy to determine the canonical document URL:

1.  **Canonical Link Tag** – Searches for `<link rel="canonical">` elements
2.  **Open Graph Meta Tag** – Falls back to `og:url` meta property
3.  **Twitter Card Meta Tag** – Checks `twitter:url` meta property
4.  **Caller Fallback** – Uses a fallback URL provided by the caller if no canonical source is found

If no URL is resolved through any of these methods, the function returns **`ErrNoURL`** and the document is rejected from indexing【fromhtml.go †L12‑L30】.

### Title Extraction Logic

Title extraction follows a similar priority pattern to URL resolution:

-   **`og:title`** or **`twitter:title`** meta tags are preferred when available
-   Falls back to the standard `<title>` element if social meta tags are absent【fromhtml.go †L62‑L77】

The result is a populated `Document` struct containing `URL`, `Title`, and raw `HTML` fields ready for further processing.

## Language Detection Pipeline

Hister's language detection is optional and controlled by the **`detect_languages`** configuration flag. When enabled, it leverages the **`lingua-go`** library to analyze text and assign ISO-639-1 language codes.

### The lingua-go Integration

The detection engine is initialized in **[`server/document/language.go`](https://github.com/asciimoo/hister/blob/main/server/document/language.go)**. The `NewLanguageDetector()` function constructs a detector preloaded with language models for all supported languages, including Arabic, Danish, English, French, German, Japanese, Korean, and Spanish【language.go†L42‑L74】.

To identify a document's language, the system calls:

```go
detector := document.NewLanguageDetector()
lang := detector.DetectLanguage(doc.Text)

```

The `DetectLanguage` method returns a two-letter ISO-639-1 code (e.g., `"en"`, `"de"`, `"no"`). Texts that fail classification return the string `"unknown"`【language.go†L98‑L107】.

### Supported Languages and Detection Logic

The detector supports a comprehensive set of languages defined in the `Languages` slice. During indexing in **[`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go)**, the detected language code is stored in the document fields:

```go
lang := detector.DetectLanguage(doc.Text)
if lang != document.UnknownLanguage {
    fields["language"] = lang
}

```

This metadata enables language-aware search faceting and filtering.

### Mapping Languages to Bleve Analyzers

Once a language is detected, Hister maps the code to a specific Bleve analyzer in **[`server/indexer/language.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/language.go)**. The mapping logic distinguishes CJK (Chinese, Japanese, Korean) text from other languages:

-   **CJK texts** – Uses the specialized `"cjk"` analyzer for proper tokenization of East Asian scripts
-   **Other languages** – Uses the language code directly as the analyzer name (e.g., `"en"` for English)【language.go†L36‑L42】

If language detection is disabled in the configuration, Hister logs a warning and falls back to the default analyzer【indexer.go†L456‑L458】.

## Configuration and Indexing Integration

### Enabling Detection via Configuration

Language detection is toggled via the **`DetectLanguages`** boolean field in **[`config/config.go`](https://github.com/asciimoo/hister/blob/main/config/config.go)**【config.go†L121】. When set to `false`, the indexer skips linguistic analysis entirely, reducing processing overhead for monolingual deployments.

### Storing Language Metadata

Hister maintains indexing consistency through metadata fingerprints. The **`AnalyzerFingerprint`** stored in [`server/indexer/fingerprint.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/fingerprint.go) records whether language detection was active during the original indexing operation【fingerprint.go†L11‑L20】. This ensures that re-indexing operations respect the original linguistic settings.

### API Exposure

The public API exposes the language field through **[`server/mcp.go`](https://github.com/asciimoo/hister/blob/main/server/mcp.go)**, where the schema defines the `language` field as the "detected language code"【mcp.go†L240‑L242】. This allows external consumers to filter search results by detected language.

## Practical Implementation Example

The following example demonstrates the complete extraction and detection workflow:

```go
package main

import (
    "fmt"
    "github.com/asciimoo/hister/server/document"
)

func main() {
    html := `<html><head>
        <title>Hola Mundo</title>
        <meta property="og:url" content="https://example.com/helloworld"/>
        <meta property="og:title" content="Hola Mundo"/>
    </head><body>¡Bienvenido!</body></html>`

    // Step 1: Extract URL and title from HTML
    doc, err := document.FromHTML(html)
    if err != nil {
        panic(err)
    }
    fmt.Printf("URL: %s\nTitle: %s\n", doc.URL, doc.Title)

    // Step 2: Detect language of the content
    detector := document.NewLanguageDetector()
    lang := detector.DetectLanguage("¡Bienvenido! Este es un ejemplo.")
    fmt.Printf("Detected language: %s\n", lang) // Output: "es"
}

```

## Summary

-   **Content extraction** occurs in [`server/document/fromhtml.go`](https://github.com/asciimoo/hister/blob/main/server/document/fromhtml.go) using `goquery`, with hierarchical fallback logic for URLs (canonical → OG → Twitter → fallback) and titles (OG/Twitter → HTML title)
-   **Language detection** relies on `lingua-go` via [`server/document/language.go`](https://github.com/asciimoo/hister/blob/main/server/document/language.go), supporting ISO-639-1 codes and automatically handling unknown texts
-   **Analyzer mapping** in [`server/indexer/language.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/language.go) routes CJK text to specialized analyzers while using standard language codes for others
-   **Configuration control** via `DetectLanguages` in [`config/config.go`](https://github.com/asciimoo/hister/blob/main/config/config.go) allows operators to disable detection for performance optimization
-   **Metadata persistence** through `AnalyzerFingerprint` ensures consistent re-indexing behavior across deployments

## Frequently Asked Questions

### How does Hister extract the canonical URL from HTML documents?

Hister checks `<link rel="canonical">` tags first, then falls back to Open Graph (`og:url`) and Twitter Card (`twitter:url`) meta tags. If none are present, it uses a caller-provided fallback URL. Without any of these sources, `FromHTML` returns `ErrNoURL` and the document is rejected.

### What languages does Hister support for automatic detection?

Hister supports all languages included in the `lingua-go` library's model set, including Arabic, Danish, English, French, German, Italian, Japanese, Korean, Norwegian, Portuguese, Russian, Spanish, and Swedish. The complete list is defined in the `Languages` slice in [`server/document/language.go`](https://github.com/asciimoo/hister/blob/main/server/document/language.go).

### How does Hister handle Chinese, Japanese, and Korean text analysis?

When the language detector identifies CJK text, [`server/indexer/language.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/language.go) maps these documents to Bleve's specialized `"cjk"` analyzer. This analyzer properly tokenizes East Asian scripts that lack whitespace delimiters, ensuring accurate search indexing for Chinese, Japanese, and Korean content.

### Can language detection be disabled in Hister?

Yes. Set `detect_languages: false` in the configuration file (controlled by the `DetectLanguages` boolean in [`config/config.go`](https://github.com/asciimoo/hister/blob/main/config/config.go)). When disabled, Hister logs a warning and indexes all documents using the default Bleve analyzer without linguistic classification.