How to Enable and Configure Semantic Search in Hister
Enable semantic search in Hister by setting semantic_search.enable: true in ~/.config/hister/config.yml and pointing embedding_endpoint to any OpenAI-compatible /v1/embeddings service.
Hister augments its traditional keyword search with vector-based semantic search to improve result relevance. This optional feature embeds queries and documents into high-dimensional vectors, allowing the engine to match conceptually similar content even without exact keyword overlap. The implementation resides in the asciimoo/hister repository and integrates via a pluggable embedder interface.
Configure Semantic Search in config.yml
The primary method to enable the feature is editing the Hister configuration YAML (default path ~/.config/hister/config.yml). Add the semantic_search block at the root level:
semantic_search:
enable: true
embedding_endpoint: "http://localhost:11434/v1/embeddings"
embedding_model: "qwen3-embedding:8b"
embedding_timeout: 300
api_key: ""
headers: {}
dimensions: 4096
max_context_length: 512
chunk_overlap: 64
max_embedding_batch_size: 8
query_prefix: "query: "
document_prefix: ""
similarity_threshold: 0.1
result_limit: 50
semantic_weight: 0.4
max_embedding_concurrency: 2
The SemanticSearch struct is defined in config/config.go (lines 68-86) and validates these fields at startup. After modifying the file, restart Hister or execute hister reload to apply changes. When enable is false, all other fields are ignored and the system falls back to pure keyword search.
Critical Configuration Parameters
embedding_endpoint: Must implement the OpenAI/v1/embeddingsAPI. The default points to a local Ollama server, but any compatible service (OpenAI, Azure, or custom) works.max_context_length: Specifies the token ceiling for document chunking. TheChunkAndEmbedfunction inserver/vectorstore/embedder.gouses this to split large documents before vectorization.similarity_threshold: Minimum cosine similarity (0.0-1.0) required for a vector match to be considered a hit. Adjust this to filter noise in dense retrieval.semantic_weight: Controls the blending ratio between keyword (BM25) and semantic (cosine) scores in the final ranking algorithm implemented inserver/indexer/indexer.go.
Toggle Semantic Search at Runtime
Once enabled in the configuration, you can control semantic search activation per query through three interfaces.
TUI Hot-Key
The terminal UI binds Ctrl+E to ActionToggleSemantic (defined in config/config.go, lines 162-168). Pressing this key flips the semantic flag for the current session, and the status bar displays "Semantic search: enabled" or "disabled".
HTTP API Parameter
When calling the /search endpoint, append semantic=1 (or semantic=true) as a query parameter. The server parses this in server/endpoints.go (around line 465) and injects the boolean into the Query struct passed to the indexer.
curl "http://localhost:4433/search?q=golang+concurrency&semantic=1"
Go Client Library
If using the official Go client (client/search.go), set Semantic: true on the SearchOptions struct. The client serializes this as the semantic query parameter:
package main
import (
"context"
"fmt"
"github.com/asciimoo/hister/client"
)
func main() {
c, _ := client.New("http://localhost:4433")
opts := client.SearchOptions{
Query: "golang concurrency patterns",
Semantic: true,
}
resp, err := c.Search(context.Background(), opts)
if err != nil {
panic(err)
}
fmt.Printf("Found %d results\n", len(resp.Results))
}
How the Search Engine Processes Semantic Queries
When semantic=true is passed to (*Indexer) search in server/indexer/indexer.go (lines 1704-1717), the engine executes the following pipeline:
- Query Preparation: The system strips wildcards that would corrupt embedding text using
querybuilder.RemoveStandaloneWildcards. - Vector Generation: The indexer calls
i.embedder.EmbedQuery(), which POSTs to the configured endpoint with thequery_prefixprepended. - Vector Retrieval: The resulting vector queries the underlying vector store (SQLite or PostgreSQL) for nearest neighbors.
- Filtering: Results below
similarity_thresholdare discarded, and the list is truncated toresult_limit. - Fusion: Semantic scores are blended with keyword BM25 scores using
semantic_weightto produce the final ranked list.
The Embedder itself is instantiated via NewEmbedder in server/vectorstore/embedder.go (lines 81-106) when the indexer initializes. It handles request batching (respecting max_embedding_batch_size), exponential backoff for HTTP 429/5xx errors, and document chunking with ChunkText to respect max_context_length and chunk_overlap.
Complete Configuration Examples
Production-Ready YAML Setup
# ~/.config/hister/config.yml
semantic_search:
enable: true
embedding_endpoint: "http://localhost:11434/v1/embeddings"
embedding_model: "qwen3-embedding:8b"
embedding_timeout: 300
api_key: ""
headers:
X-Custom-Auth: "bearer-token"
dimensions: 4096
max_context_length: 512
chunk_overlap: 64
max_embedding_batch_size: 8
query_prefix: "query: "
document_prefix: "passage: "
similarity_threshold: 0.15
result_limit: 100
semantic_weight: 0.5
max_embedding_concurrency: 2
Bash One-Liner with Semantic Flag
curl -G "http://localhost:4433/search" \
--data-urlencode "q=error handling in rust" \
-d "semantic=1" \
-d "limit=20"
Summary
- Enable the feature by setting
semantic_search.enable: truein~/.config/hister/config.ymland restart Hister to load theSemanticSearchstruct fromconfig/config.go. - Configure the
embedding_endpointto point at any OpenAI-compatible/v1/embeddingsservice; tunesimilarity_thresholdandsemantic_weightto balance precision and recall. - Toggle semantic search at runtime via
Ctrl+Ein the TUI, thesemantic=1HTTP parameter (parsed inserver/endpoints.go), or theSemanticfield in the Go client (client/search.go). - Understand that the indexer in
server/indexer/indexer.gocalls theEmbedder(created inserver/vectorstore/embedder.go) to vectorize queries, retrieve nearest neighbors, and fuse results with traditional keyword hits.
Frequently Asked Questions
What embedding services are compatible with Hister's semantic search?
Hister requires an endpoint implementing the OpenAI /v1/embeddings API format. Compatible providers include Ollama (default localhost:11434), OpenAI, Azure OpenAI, and any custom proxy exposing the same JSON schema. The embedding_model parameter is passed directly to the endpoint's model field.
How does Hister handle large documents that exceed the embedding model's context window?
The Embedder in server/vectorstore/embedder.go automatically chunks documents using the ChunkText function, respecting the max_context_length and chunk_overlap configuration values. Each chunk is embedded separately, and the system aggregates results to ensure no content is dropped due to token limits.
Can I adjust how much semantic results influence the final ranking?
Yes. Modify the semantic_weight parameter in config.yml (default 0.4). This value controls the interpolation between keyword (BM25) and semantic (cosine similarity) scores in the final ranking calculation performed by (*Indexer) search in server/indexer/indexer.go. Set to 1.0 for pure semantic search, 0.0 for pure keyword search.
Why am I getting timeout errors when enabling semantic search?
The embedding_timeout field (default 300 seconds) controls the HTTP client timeout for embedding requests. If your endpoint is slow or you are processing large batches (controlled by max_embedding_batch_size), increase this value or reduce max_embedding_concurrency to limit parallel load on the embedding service.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →