How to Define Skip and Priority Rules for Indexing in Hister

Hister controls indexing behavior through a global Rules configuration structure that uses regex-based skip patterns to exclude URLs and priority patterns to control processing order.

The asciimoo/hister search engine provides granular control over which documents enter the index and how quickly they are processed. By configuring the Rules struct defined in the configuration layer, you can prevent irrelevant content from being crawled while ensuring high-priority URLs are indexed first.

Understanding the Rules Configuration Structure

At the core of Hister's indexing control is the Rules struct defined in config/config.go. This structure maintains two distinct rule sets that the indexer consults during every crawl operation.

The Rules Struct

The configuration defines Rules as a container for both exclusion and prioritization logic:

  • Skip (rules.Skip): A collection of regular expression strings (ReStrs) that match URLs which should be ignored entirely
  • Priority (rules.Priority): A collection of regex patterns paired with numeric values that determine processing precedence

According to the source code at config/config.go#L151-L165, these fields are initialized as slices that the indexer iterates during the document processing pipeline.

Configuring Skip Rules

Skip rules prevent the indexer from processing unwanted URLs. When the crawler encounters a URL matching any skip pattern, it immediately excludes the document from the indexing queue and emits a log entry indicating the skip action.

YAML Configuration Format

Define skip rules in your Hister configuration file using standard regular expressions:

rules:
  skip:
    - "^https?://(www\\.)?example\\.com/skip/.*$"
    - "^file://.*\\.tmp$"
    - "\\.pdf$"

Each entry in the skip array is compiled into a regex matcher during server initialization. According to the implementation in server/endpoints.go#L1248-L1281, these patterns are read from the configuration and applied to every URL discovered during crawling.

Programmatic Skip Rule Management

Use the Go client library to dynamically add skip rules:

// client/rules.go
func AddSkipRule(rules *config.Rules, pattern string) {
    rules.Skip.ReStrs = append(rules.Skip.ReStrs, pattern)
}

Reference: client/rules.go#L12-L18

Configuring Priority Rules

Priority rules assign numeric weights to URLs, allowing the indexer to process important documents before lower-priority ones. Higher integer values indicate higher priority, pushing matching URLs to the front of the indexing queue.

YAML Configuration Format

Priority rules require both a pattern and a priority value:

rules:
  priority:
    - pattern: "^https?://(www\\.)?important\\.com/.*$"
      priority: 10
    - pattern: "^https?://.*\\.pdf$"
      priority: 5

The priority field accepts a list of objects containing pattern (string) and priority (integer) keys. The indexer evaluates these rules after skip checks but before queue insertion.

Programmatic Priority Rule Management

Add priority rules dynamically using the client library:

func AddPriorityRule(rules *config.Rules, pattern string, prio int) {
    rules.Priority.ReStrs = append(rules.Priority.ReStrs,
        config.PriorityRule{Pattern: pattern, Priority: prio})
}

Reference: client/rules.go#L20-L26

Command-Line Interface Configuration

Hister provides CLI commands for rule management without requiring direct file editing. The rules set command parses arguments, updates the in-memory Rules object, and persists changes to the configuration file.

hister rules set --skip "^https?://.*\\.zip$" \
                 --priority "^https?://news\\.example\\.com/.*$=20"

This interface is implemented in cmd/rules.go, which serves as the entry point for administrative rule modifications.

How the Indexer Applies Rules

During the indexing pipeline, Hister evaluates rules in two distinct phases to determine document fate and queue position.

Skip Evaluation Logic

The core indexer checks skip rules before processing content:

// server/indexer/indexer.go (excerpt)
func (i *Indexer) shouldSkip(url string) bool {
    for _, re := range i.rules.Skip.ReStrs {
        if matched, _ := regexp.MatchString(re, url); matched {
            return true
        }
    }
    return false
}

Reference: server/indexer/indexer.go#L711-L720

When shouldSkip returns true, the indexer immediately discards the document and logs the exclusion event.

Priority Assignment

After passing skip filters, URLs are evaluated against priority patterns. The highest matching priority value is assigned to the document metadata, influencing the crawl queue ordering. The indexer consults rules.Priority.ReStrs to determine the final priority score before queue insertion.

Summary

  • Configuration Location: The Rules struct in config/config.go defines both skip and priority rule containers
  • Skip Rules: Regex patterns in rules.Skip.ReStrs exclude matching URLs from indexing entirely
  • Priority Rules: Pattern-value pairs in rules.Priority.ReStrs assign numeric precedence to control processing order
  • Implementation: The server/indexer/indexer.go file contains the shouldSkip function that encludes exclusion logic during crawling
  • Management: Rules can be configured via YAML files, CLI commands (cmd/rules.go), or the Go client library (client/rules.go)

Frequently Asked Questions

What regex syntax does Hister use for skip and priority patterns?

Hister uses Go's standard regular expression syntax (RE2) for pattern matching. Both rules.Skip.ReStrs and rules.Priority.ReStrs are passed directly to regexp.MatchString during evaluation, supporting standard anchors (^, $), character classes, and quantifiers without backreferences or advanced Perl features.

Can I combine multiple skip patterns for different URL types?

Yes, the Rules struct accepts multiple patterns in the Skip.ReStrs slice. The indexer iterates through every pattern in the array, and if any pattern matches, the URL is excluded. This allows you to maintain separate rules for file extensions, domains, and path patterns simultaneously.

How do priority values affect the crawling order?

Higher integer values in priority rules cause URLs to be processed earlier in the crawl queue. When the indexer discovers a URL matching a priority pattern, it assigns that numeric value to the document. The crawl scheduler uses these values to sort the queue, ensuring priority 10 documents are fetched before priority 1 documents.

Where does Hister persist skip and priority rules between restarts?

Rules are persisted in Hister's configuration file (typically config.yaml or the file specified via CLI flags). The config/config.go package handles serialization of the Rules struct, while server/endpoints.go manages runtime updates that trigger configuration reloads or restarts depending on your deployment mode.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →