What Crawler Backends Does Hister Support? A Complete Guide to HTTP, Chromedp, and BiDi

Hister supports three distinct crawler backends: a lightweight HTTP fetcher for static content, a Chromedp backend for full JavaScript rendering via headless Chrome, and a BiDi backend using Chrome's Bidirectional Protocol for modern page interaction.

Hister is an open-source crawling engine designed to fetch and process web content efficiently. Understanding the different crawler backends available in asciimoo/hister is essential for selecting the right approach based on whether your target sites use server-side rendering or complex client-side JavaScript frameworks.

Overview of Hister's Crawler Architecture

The crawling engine abstracts all fetching logic behind the crawler.Crawler interface defined in server/crawler/crawler.go. Rather than hardcoding a specific fetching mechanism, Hister uses a factory pattern that instantiates different concrete implementations based on the CrawlerConfig.Backend configuration value.

This architectural choice allows operators to swap between lightweight static fetching and full browser automation without changing application code. The selection logic resides in the New factory function, which switches on the backend string and returns the appropriate fetcher implementation.

The Three Crawler Backends Explained

Hister provides three distinct backend implementations, each optimized for different crawling scenarios.

HTTP Backend (Default)

The HTTP backend is the default fetcher implementation found in server/crawler/http.go. This backend uses Go's standard net/http client to download raw HTML directly from target servers.

Key characteristics:

  • Zero JavaScript execution; downloads static HTML only
  • Fastest performance with minimal resource overhead
  • Ideal for traditional server-rendered websites, documentation, and blogs
  • Handles redirects and proxy configurations via standard HTTP client options

When no backend is specified in the configuration, or when an unrecognized value is provided, Hister automatically falls back to this implementation.

Chromedp Backend

The chromedp backend, implemented in server/crawler/chromedp.go, leverages the chromedp Go library to launch and control a headless Chrome instance. This enables full JavaScript rendering and DOM manipulation.

Key characteristics:

  • Executes JavaScript and renders the complete DOM as a user would see it
  • Supports configurable capture delays to wait for dynamic content loading
  • Can extract links and content from single-page applications (SPAs) like Notion or React-based sites
  • Requires significantly more memory and CPU resources than the HTTP backend

This backend is essential when crawling modern web applications that rely on client-side rendering or infinite scroll behaviors.

BiDi Backend

The BiDi backend, located in server/crawler/bidi.go, utilizes Chrome's Bidirectional Protocol (BiDi) for page interaction and rendering. This offers an alternative to chromedp with similar capabilities but through Chrome's newer DevTools Protocol interface.

Key characteristics:

  • Implements the Chrome DevTools BiDi specification for modern browser automation
  • Provides the same JavaScript rendering capabilities as chromedp
  • Offers an alternative protocol choice for environments where BiDi is preferred over the legacy Chrome DevTools Protocol used by chromedp
  • Supports the same configuration options for timeouts, proxies, and capture delays

How Backend Selection Works in the Code

The factory function New in server/crawler/crawler.go (lines 73-80) handles the instantiation logic based on the configuration:

switch cfg.Backend {
case "chromedp":
    f, err = newChromedpFetcher(cfg)
case "bidi":
    f, err = newBidiFetcher(cfg)
default: // "" or any other value
    f, err = newHTTPFetcher(cfg)
}

This switch statement demonstrates how Hister maps string configuration values to concrete fetcher implementations. The cfg.Backend field from config.CrawlerConfig drives the selection, with any empty string or unrecognized value defaulting to the HTTP fetcher for safety and performance.

Configuring the Crawler Backend

You can specify the backend through YAML configuration or programmatically via the Go API.

YAML Configuration

Define the backend in your configuration file as defined in config/config.go:

crawler:
  backend: chromedp        # options: http, chromedp, bidi

  capture_delay: "2s"      # only affects chromedp/bidi

  timeout: 30              # seconds

  proxy: http://localhost:3128

The backend field accepts the three string values discussed above. Note that capture_delay is only relevant for the chromedp and BiDi backends, as it specifies how long to wait after page load before capturing the DOM.

Programmatic Usage

To create a crawler instance with a specific backend in Go:

import (
    "context"
    "fmt"
    "github.com/asciimoo/hister/config"
    "github.com/asciimoo/hister/server/crawler"
)

func crawlExample() error {
    // Configure for JavaScript rendering
    cfg := &config.CrawlerConfig{
        Backend: "chromedp",
        Timeout: 30,
        CaptureDelay: "2s",
    }

    // Optional robots.txt cache
    robots, _ := crawler.NewRobotsCache(cfg.UserAgent, cfg.Proxy)

    // Factory creates the appropriate backend
    cr, err := crawler.New(cfg, robots)
    if err != nil {
        return err
    }
    defer cr.Close()

    // Execute crawl
    ctx := context.Background()
    docs, err := cr.Crawl(ctx, "https://example.com", crawler.NewValidator(&crawler.ValidatorRules{}))
    if err != nil {
        return err
    }

    for d := range docs {
        fmt.Println("Fetched:", d.URL)
        // d.HTML contains rendered content for chromedp/bidi
    }
    return nil
}

The crawler.New function returns a Crawler interface backed by the concrete implementation specified in the configuration, allowing the rest of your application to remain agnostic to the fetching mechanism.

Summary

  • Hister provides three crawler backends: HTTP (default), chromedp, and BiDi, each suited to different content types.
  • HTTP backend in server/crawler/http.go offers fast, resource-efficient static HTML fetching without JavaScript execution.
  • Chromedp backend in server/crawler/chromedp.go provides full browser automation for JavaScript-heavy sites using the chromedp library.
  • BiDi backend in server/crawler/bidi.go offers equivalent rendering capabilities through Chrome's Bidirectional Protocol.
  • Backend selection occurs in server/crawler/crawler.go via the New factory function based on the CrawlerConfig.Backend configuration value.

Frequently Asked Questions

What is the default crawler backend in Hister?

The HTTP backend is the default when no backend is specified or when an invalid value is provided in the configuration. This implementation uses Go's standard net/http client to fetch raw HTML without executing JavaScript, making it suitable for static websites and the most resource-efficient option. The fallback logic is implemented in the New function in server/crawler/crawler.go.

When should I use chromedp instead of the HTTP backend?

Use the chromedp backend when crawling modern web applications that rely on client-side JavaScript frameworks like React, Vue, or Angular. Sites like Notion, dashboard applications, or content loaded via XHR/fetch requests require JavaScript execution to render properly. The chromedp backend in server/crawler/chromedp.go launches a headless Chrome instance to execute this JavaScript and capture the fully rendered DOM.

How do I configure the capture delay for JavaScript rendering?

Set the capture_delay field in your CrawlerConfig to specify how long the crawler should wait after the initial page load before capturing content. This is configured in the YAML file or struct as a duration string (e.g., "2s"). This setting only affects the chromedp and BiDi backends, as implemented in config/config.go, allowing dynamic content to finish loading before extraction occurs.

What is the difference between chromedp and BiDi backends?

Both backends provide full Chrome browser automation and JavaScript execution, but they use different protocols to communicate with the browser. Chromedp uses the traditional Chrome DevTools Protocol (CDP) through the third-party chromedp Go library, while BiDi uses Chrome's newer Bidirectional Protocol specification. The BiDi backend in server/crawler/bidi.go offers the same rendering capabilities as chromedp but implements the emerging web standard for browser automation, potentially offering better future compatibility as BiDi becomes the dominant protocol.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →