How Hister Handles Content Extraction and Language Detection

Hister extracts canonical URLs and titles from HTML using goquery in server/document/fromhtml.go, then detects document language via the lingua-go library in server/document/language.go to enable language-specific Bleve indexing.

Hister is an open-source search engine that transforms raw HTML into searchable documents. Understanding how it handles content extraction and language detection is essential for configuring multilingual indexing pipelines. The implementation spans dedicated modules in the server/document and server/indexer packages, orchestrating metadata parsing and linguistic analysis.

Content Extraction from HTML

The core extraction logic resides in server/document/fromhtml.go, where the FromHTML function parses HTML strings and populates a *Document struct with normalized metadata.

Parsing HTML with goquery

Hister relies on the goquery library to navigate and query HTML DOM structures. The FromHTML function accepts an HTML string and systematically extracts critical metadata fields required for indexing.

URL Resolution Strategy

Hister implements a hierarchical fallback strategy to determine the canonical document URL:

  1. Canonical Link Tag – Searches for <link rel="canonical"> elements
  2. Open Graph Meta Tag – Falls back to og:url meta property
  3. Twitter Card Meta Tag – Checks twitter:url meta property
  4. Caller Fallback – Uses a fallback URL provided by the caller if no canonical source is found

If no URL is resolved through any of these methods, the function returns ErrNoURL and the document is rejected from indexing【fromhtml.go †L12‑L30】.

Title Extraction Logic

Title extraction follows a similar priority pattern to URL resolution:

  • og:title or twitter:title meta tags are preferred when available
  • Falls back to the standard <title> element if social meta tags are absent【fromhtml.go †L62‑L77】

The result is a populated Document struct containing URL, Title, and raw HTML fields ready for further processing.

Language Detection Pipeline

Hister's language detection is optional and controlled by the detect_languages configuration flag. When enabled, it leverages the lingua-go library to analyze text and assign ISO-639-1 language codes.

The lingua-go Integration

The detection engine is initialized in server/document/language.go. The NewLanguageDetector() function constructs a detector preloaded with language models for all supported languages, including Arabic, Danish, English, French, German, Japanese, Korean, and Spanish【language.go†L42‑L74】.

To identify a document's language, the system calls:

detector := document.NewLanguageDetector()
lang := detector.DetectLanguage(doc.Text)

The DetectLanguage method returns a two-letter ISO-639-1 code (e.g., "en", "de", "no"). Texts that fail classification return the string "unknown"【language.go†L98‑L107】.

Supported Languages and Detection Logic

The detector supports a comprehensive set of languages defined in the Languages slice. During indexing in server/indexer/indexer.go, the detected language code is stored in the document fields:

lang := detector.DetectLanguage(doc.Text)
if lang != document.UnknownLanguage {
    fields["language"] = lang
}

This metadata enables language-aware search faceting and filtering.

Mapping Languages to Bleve Analyzers

Once a language is detected, Hister maps the code to a specific Bleve analyzer in server/indexer/language.go. The mapping logic distinguishes CJK (Chinese, Japanese, Korean) text from other languages:

  • CJK texts – Uses the specialized "cjk" analyzer for proper tokenization of East Asian scripts
  • Other languages – Uses the language code directly as the analyzer name (e.g., "en" for English)【language.go†L36‑L42】

If language detection is disabled in the configuration, Hister logs a warning and falls back to the default analyzer【indexer.go†L456‑L458】.

Configuration and Indexing Integration

Enabling Detection via Configuration

Language detection is toggled via the DetectLanguages boolean field in config/config.go【config.go†L121】. When set to false, the indexer skips linguistic analysis entirely, reducing processing overhead for monolingual deployments.

Storing Language Metadata

Hister maintains indexing consistency through metadata fingerprints. The AnalyzerFingerprint stored in server/indexer/fingerprint.go records whether language detection was active during the original indexing operation【fingerprint.go†L11‑L20】. This ensures that re-indexing operations respect the original linguistic settings.

API Exposure

The public API exposes the language field through server/mcp.go, where the schema defines the language field as the "detected language code"【mcp.go†L240‑L242】. This allows external consumers to filter search results by detected language.

Practical Implementation Example

The following example demonstrates the complete extraction and detection workflow:

package main

import (
    "fmt"
    "github.com/asciimoo/hister/server/document"
)

func main() {
    html := `<html><head>
        <title>Hola Mundo</title>
        <meta property="og:url" content="https://example.com/helloworld"/>
        <meta property="og:title" content="Hola Mundo"/>
    </head><body>¡Bienvenido!</body></html>`

    // Step 1: Extract URL and title from HTML
    doc, err := document.FromHTML(html)
    if err != nil {
        panic(err)
    }
    fmt.Printf("URL: %s\nTitle: %s\n", doc.URL, doc.Title)

    // Step 2: Detect language of the content
    detector := document.NewLanguageDetector()
    lang := detector.DetectLanguage("¡Bienvenido! Este es un ejemplo.")
    fmt.Printf("Detected language: %s\n", lang) // Output: "es"
}

Summary

  • Content extraction occurs in server/document/fromhtml.go using goquery, with hierarchical fallback logic for URLs (canonical → OG → Twitter → fallback) and titles (OG/Twitter → HTML title)
  • Language detection relies on lingua-go via server/document/language.go, supporting ISO-639-1 codes and automatically handling unknown texts
  • Analyzer mapping in server/indexer/language.go routes CJK text to specialized analyzers while using standard language codes for others
  • Configuration control via DetectLanguages in config/config.go allows operators to disable detection for performance optimization
  • Metadata persistence through AnalyzerFingerprint ensures consistent re-indexing behavior across deployments

Frequently Asked Questions

How does Hister extract the canonical URL from HTML documents?

Hister checks <link rel="canonical"> tags first, then falls back to Open Graph (og:url) and Twitter Card (twitter:url) meta tags. If none are present, it uses a caller-provided fallback URL. Without any of these sources, FromHTML returns ErrNoURL and the document is rejected.

What languages does Hister support for automatic detection?

Hister supports all languages included in the lingua-go library's model set, including Arabic, Danish, English, French, German, Italian, Japanese, Korean, Norwegian, Portuguese, Russian, Spanish, and Swedish. The complete list is defined in the Languages slice in server/document/language.go.

How does Hister handle Chinese, Japanese, and Korean text analysis?

When the language detector identifies CJK text, server/indexer/language.go maps these documents to Bleve's specialized "cjk" analyzer. This analyzer properly tokenizes East Asian scripts that lack whitespace delimiters, ensuring accurate search indexing for Chinese, Japanese, and Korean content.

Can language detection be disabled in Hister?

Yes. Set detect_languages: false in the configuration file (controlled by the DetectLanguages boolean in config/config.go). When disabled, Hister logs a warning and indexes all documents using the default Bleve analyzer without linguistic classification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →