How Hister Uses Bleve for Full-Text Search: Architecture and Implementation Guide
Hister implements full-text search using Bleve v2 by creating language-aware indexes with custom analyzers, routing documents through a single index alias, and translating a custom DSL into Bleve query objects for faceted, highlighted search results.
Hister is an open-source search engine that leverages Bleve v2 to provide fast, scalable full-text search capabilities. The implementation centers on a modular indexer architecture that supports multilingual content, custom tokenization, and optional semantic vector search augmentation. This guide examines how Hister uses Bleve for full-text search by breaking down the indexing pipeline, query execution, and result enhancement strategies.
Initializing the Bleve Index and Custom Mapping
When Hister instantiates an Indexer via New in server/indexer/indexer.go, it initializes the Bleve search backend through initializeIndexer (lines 91-106). This function opens or creates a default Bleve index named index.db in the configured data directory.
If language detection is enabled, Hister creates separate language-specific indexes (index_<lang>.db) for each detected language. The index mapping is constructed by createMapping (lines 2043-2070), which registers a custom analyzer specifically designed for Hister's use case:
- Single-token tokenizer: Preserves the original text structure while enabling tokenization
- Lowercase token filter: Ensures case-insensitive matching
This custom analyzer configuration keeps text intact for exact-match queries while supporting efficient tokenized search.
Indexing Documents with Language Detection
Document ingestion flows through AddDocument or AddDocumentContext (lines 25-31), which orchestrates the indexing pipeline:
- Language detection: Determines the document's language for routing
- Content extraction: Parses the document text and metadata
- External storage: Writes large blobs (HTML content, favicons) to separate data stores using SHA-256 keys, keeping the Bleve index lean
- Index routing: Selects the appropriate Bleve index via
getOrCreate(d.Language), routing to either the default index or a language-specific shard - Bleve indexing: Calls
plan.target.Index(d.ID(), d)to store the document
The prepareStorageWrite and applyDocumentWrite functions (lines 124-140) handle the actual Bleve Index operation, ensuring documents are persisted with their language-specific mappings.
doc := &document.Document{
URL: "https://example.com",
Title: "Example Page",
Text: "Bleve provides full‑text search for Go programs.",
UserID: 1,
}
if err := idx.AddDocument(doc); err != nil {
log.Fatal(err)
}
Building and Executing Search Queries
Hister translates user search strings into Bleve query objects through querybuilder.ParseSearch in server/indexer/querybuilder/builder.go. The parser converts Hister's custom DSL into a hierarchy of Bleve query.Query types:
- Match queries for full-text search
- Term queries for exact matches
- Phrase queries for sequential word matching
- Regex and wildcard queries for pattern matching
- Numeric range queries for date and number filtering
The search function (lines 1616-1645) assembles these into a bleve.SearchRequest and executes it against i.idx—a bleve.IndexAlias that aggregates all language-specific indexes. This alias capability allows a single search request to query across all language shards simultaneously.
q := &indexer.Query{
Text: "full‑text search",
Facets: true,
Limit: 20,
}
res, err := idx.Search(q)
if err != nil {
log.Fatal(err)
}
for _, d := range res.Documents {
fmt.Println(d.Title, d.URL)
}
Facets, Highlighting, and Result Enhancement
After executing the core Bleve search, Hister enriches results through two mechanisms defined in server/indexer/indexer.go:
Facet Aggregation (addFacets, lines 36-73)
- Translates schema definitions from
searchschema/*.gointo Bleve facet requests - Supports term facets, numeric ranges, and date ranges
- Enables faceted navigation in search results
Text Highlighting (registerHighlighters, lines 78-88)
- Registers Bleve's
simpleFragmenterandsimpleHighlighter - Returns highlighted snippets showing query term matches in context
- Supports ANSI/TUI formatting for terminal interfaces
Optional Semantic Search Integration
When vector search is enabled, Hister extends Bleve's capabilities by running queries through an embedding pipeline. The SemanticSearchEnabled flag (lines 77-78) gates this functionality. The system:
- Generates vector embeddings from the textual query
- Retrieves similar vectors from the vector store
- Merges vector similarity results with traditional Bleve hits
This architecture demonstrates how Hister uses Bleve as the primary inverted index while delegating semantic similarity to specialized vector stores.
Summary
- Hister uses Bleve v2 as its core full-text search engine, implemented primarily in
server/indexer/indexer.go - Language-aware indexing creates separate Bleve indexes per language when detection is enabled, with documents routed via
getOrCreate - Single index alias (
i.idx) aggregates all language shards, enabling cross-language search with one query - Custom analyzer combines single-token tokenization with lowercase filtering for flexible exact and partial matching
- External blob storage keeps the Bleve index small by storing HTML and favicons separately using SHA-256 keys
- Rich query DSL supports match, term, phrase, regex, wildcard, and range queries parsed by
querybuilder.ParseSearch - Faceted search and highlighting enhance results through
addFacetsandregisterHighlighters
Frequently Asked Questions
How does Hister handle multiple languages in Bleve?
Hister creates separate Bleve indexes for each detected language (e.g., index_de.db for German) when DetectLanguages is enabled. The getOrCreate function routes documents to the appropriate language-specific index based on detection results. An index alias (i.idx) aggregates all language indexes, allowing a single search to query across all languages simultaneously.
What is the purpose of the custom analyzer in Hister's Bleve implementation?
The custom analyzer, defined in createMapping, uses a single-token tokenizer paired with a lowercase token filter. This configuration preserves the original text structure for exact-match queries while enabling case-insensitive tokenized search, optimizing the index for Hister's specific retrieval patterns.
How does Hister keep the Bleve index size optimized?
Hister stores large binary blobs—including HTML content and favicons—outside the Bleve index in a separate data directory using SHA-256 keys. Only searchable metadata and text content are indexed by Bleve, significantly reducing index size and improving query performance.
Can Hister combine traditional Bleve search with semantic vector search?
Yes. When the vector store is enabled via SemanticSearchEnabled, Hister generates embeddings from the query text and retrieves similar vectors. These semantic results are merged with traditional Bleve hits, allowing the system to combine lexical matching from Bleve with semantic similarity from vector embeddings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →