How Hister Uses Bleve for Full-Text Search: Architecture and Implementation
Hister leverages Bleve v2 to power its full-text search through a multi-index architecture that supports language-aware sharding, custom text analyzers, and a domain-specific query language compiled into Bleve query objects.
Hister is an open-source search engine built in Go that utilizes the Bleve library for indexing and querying content. According to the asciimoo/hister source code, the implementation combines multiple language-specific indexes under a unified alias, enabling efficient full-text retrieval across multilingual document collections while keeping the search index lightweight and fast.
Index Architecture and Initialization
The foundation of Hister's search capability lies in how it initializes and structures its Bleve indexes.
Creating the Index Mapping
When the Indexer is instantiated via New in server/indexer/indexer.go, it calls initializeIndexer to open or create the underlying Bleve indexes. By default, Hister creates a main index at index.db. The mapping is constructed by createMapping(lang, keepStopwords), which registers a custom analyzer configured with a single-token tokenizer and a lowercase token filter. This design preserves exact-match capabilities while enabling tokenized search.
cfg := &config.Config{
App: config.AppConfig{DisablePreviews: false},
Indexer: config.IndexerConfig{
DetectLanguages: true,
KeepStopwords: false,
},
}
idx, err := indexer.New(cfg) // server/indexer/indexer.go#L25-L31
Language-Aware Index Sharding
When DetectLanguages is enabled, Hister creates separate indexes for each detected language (e.g., index_de.db for German, index_en.db for English). The getOrCreate(d.Language) function routes documents to their respective language indexes. These are managed under a single bleve.IndexAlias (i.idx), allowing search requests to query all language shards simultaneously.
doc.Language = "de" // German
idx.AddDocument(doc) // routed to index_de.db via getOrCreate
Document Ingestion Pipeline
Hister separates the concerns of indexing searchable text and storing large binary content, optimizing index performance.
The AddDocument Flow
The AddDocumentContext method in server/indexer/indexer.go handles document ingestion. It first validates the document, detects the language if enabled, and extracts content. Large blobs such as HTML content and favicons are written to a separate data store using SHA-256 keys, rather than being embedded in the Bleve index. The document metadata and text are then indexed via plan.target.Index(d.ID(), d).
doc := &document.Document{
URL: "https://example.com",
Title: "Example Page",
Text: "Bleve provides full-text search for Go programs.",
UserID: 1,
}
if err := idx.AddDocument(doc); err != nil {
log.Fatal(err)
} // server/indexer/indexer.go#L21-L30
Storage Optimization
By storing HTML and favicon data outside the index (in the data directory keyed by SHA-256), Hister keeps the Bleve indexes compact. This separation reduces index size and improves full-text query performance, as the search engine only processes lightweight document metadata and text fields.
Query Building and Execution
Hister implements a custom query DSL that compiles into native Bleve query objects, enabling complex search semantics.
Parsing the Query DSL
The querybuilder.ParseSearch function in server/indexer/querybuilder/builder.go parses user-provided search strings into a hierarchy of Bleve query.Query implementations. The builder supports match, term, phrase, regex, wildcard, and numeric range queries, translating the DSL into the appropriate Bleve query types.
Executing Searches
The Indexer.search method (around line 1616 in indexer.go) constructs a bleve.SearchRequest from the compiled query. If facets are enabled, addFacets appends facet definitions from the searchschema package. The request is executed against the IndexAlias (i.idx), which automatically broadcasts the query to all language-specific indexes and aggregates the results.
q := &indexer.Query{
Text: "full-text search",
Facets: true,
Limit: 20,
}
res, err := idx.Search(q) // server/indexer/indexer.go#L1616-L1645
Facets and Highlighting
Hister enhances search results with faceted navigation and text highlighting using Bleve's built-in capabilities.
Configuring Facets
The addFacets function registers facet requests for terms, numeric ranges, and date ranges based on definitions in searchschema. These facets allow users to filter results by categories, dates, or custom numeric fields after the initial full-text search.
Result Highlighting
Hister registers custom highlighters via registerHighlighters in indexer.go. It uses Bleve's simpleFragmenter and simpleHighlighter to generate highlighted snippets, with support for ANSI/TUI formatting in terminal interfaces.
Semantic Search Integration
While Bleve handles lexical search, Hister optionally augments results with semantic (vector) search capabilities.
When SemanticSearchEnabled is true, Hister initializes a vector store alongside the Bleve indexes. During query execution, the textual query is processed through an embedding pipeline. The results are then merged with the Bleve hits, combining lexical matching with vector similarity for enhanced relevance.
cfg.SemanticSearch.Enable = true
idx, _ = indexer.New(cfg) // embeds the vector store
q.SemanticEnabled = true
res, _ = idx.Search(q) // merges Bleve hits with vector similarity hits
Summary
- Hister builds on Bleve v2, using a main index and optional language-specific shards managed under an
IndexAlias. - Documents are ingested via
AddDocumentContext, which routes content to language-specific indexes while storing large blobs externally. - The
querybuilder.ParseSearchfunction translates Hister's DSL into Bleve query objects, supporting boolean logic, ranges, and wildcards. - Facets and highlighting are implemented through
addFacetsandregisterHighlighters, leveraging Bleve's fragmenter and highlighter interfaces. - Optional semantic search integrates vector embeddings with Bleve's lexical search for hybrid retrieval.
Frequently Asked Questions
What version of Bleve does Hister use?
Hister uses Bleve v2, the current major version of the Go full-text search library. The codebase imports the v2 API paths and utilizes modern Bleve features such as the IndexAlias interface for multi-index queries.
How does Hister handle multilingual content?
When language detection is enabled in the configuration, Hister creates separate Bleve indexes for each language (e.g., index_de.db, index_fr.db). The getOrCreate function routes documents to the appropriate index based on detected language, while the IndexAlias aggregates all indexes for unified search queries.
Why does Hister store HTML content outside the Bleve index?
Hister stores HTML and favicon data in a separate data directory using SHA-256 keys to keep the Bleve indexes small and fast. This architectural decision prevents text search operations from being slowed down by large binary payloads, as the index only contains lightweight metadata and searchable text fields.
Can Hister perform semantic search without Bleve?
No, the semantic search feature is designed to complement rather than replace Bleve. When enabled, Hister runs the query through both the vector store (for semantic similarity) and Bleve (for lexical matching), then merges the results. The Bleve index remains the primary search engine, with vector search as an optional augmentation layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →