How to Enable and Configure Semantic Search in Hister

Enable semantic search in Hister by setting semantic_search.enable: true in ~/.config/hister/config.yml and pointing embedding_endpoint to any OpenAI-compatible /v1/embeddings service.

Hister augments its traditional keyword search with vector-based semantic search to improve result relevance. This optional feature embeds queries and documents into high-dimensional vectors, allowing the engine to match conceptually similar content even without exact keyword overlap. The implementation resides in the asciimoo/hister repository and integrates via a pluggable embedder interface.

Configure Semantic Search in config.yml

The primary method to enable the feature is editing the Hister configuration YAML (default path ~/.config/hister/config.yml). Add the semantic_search block at the root level:

semantic_search:
  enable: true
  embedding_endpoint: "http://localhost:11434/v1/embeddings"
  embedding_model: "qwen3-embedding:8b"
  embedding_timeout: 300
  api_key: ""
  headers: {}
  dimensions: 4096
  max_context_length: 512
  chunk_overlap: 64
  max_embedding_batch_size: 8
  query_prefix: "query: "
  document_prefix: ""
  similarity_threshold: 0.1
  result_limit: 50
  semantic_weight: 0.4
  max_embedding_concurrency: 2

The SemanticSearch struct is defined in config/config.go (lines 68-86) and validates these fields at startup. After modifying the file, restart Hister or execute hister reload to apply changes. When enable is false, all other fields are ignored and the system falls back to pure keyword search.

Critical Configuration Parameters

  • embedding_endpoint: Must implement the OpenAI /v1/embeddings API. The default points to a local Ollama server, but any compatible service (OpenAI, Azure, or custom) works.
  • max_context_length: Specifies the token ceiling for document chunking. The ChunkAndEmbed function in server/vectorstore/embedder.go uses this to split large documents before vectorization.
  • similarity_threshold: Minimum cosine similarity (0.0-1.0) required for a vector match to be considered a hit. Adjust this to filter noise in dense retrieval.
  • semantic_weight: Controls the blending ratio between keyword (BM25) and semantic (cosine) scores in the final ranking algorithm implemented in server/indexer/indexer.go.

Toggle Semantic Search at Runtime

Once enabled in the configuration, you can control semantic search activation per query through three interfaces.

TUI Hot-Key

The terminal UI binds Ctrl+E to ActionToggleSemantic (defined in config/config.go, lines 162-168). Pressing this key flips the semantic flag for the current session, and the status bar displays "Semantic search: enabled" or "disabled".

HTTP API Parameter

When calling the /search endpoint, append semantic=1 (or semantic=true) as a query parameter. The server parses this in server/endpoints.go (around line 465) and injects the boolean into the Query struct passed to the indexer.

curl "http://localhost:4433/search?q=golang+concurrency&semantic=1"

Go Client Library

If using the official Go client (client/search.go), set Semantic: true on the SearchOptions struct. The client serializes this as the semantic query parameter:

package main

import (
    "context"
    "fmt"
    "github.com/asciimoo/hister/client"
)

func main() {
    c, _ := client.New("http://localhost:4433")
    opts := client.SearchOptions{
        Query:   "golang concurrency patterns",
        Semantic: true,
    }
    resp, err := c.Search(context.Background(), opts)
    if err != nil {
        panic(err)
    }
    fmt.Printf("Found %d results\n", len(resp.Results))
}

How the Search Engine Processes Semantic Queries

When semantic=true is passed to (*Indexer) search in server/indexer/indexer.go (lines 1704-1717), the engine executes the following pipeline:

  1. Query Preparation: The system strips wildcards that would corrupt embedding text using querybuilder.RemoveStandaloneWildcards.
  2. Vector Generation: The indexer calls i.embedder.EmbedQuery(), which POSTs to the configured endpoint with the query_prefix prepended.
  3. Vector Retrieval: The resulting vector queries the underlying vector store (SQLite or PostgreSQL) for nearest neighbors.
  4. Filtering: Results below similarity_threshold are discarded, and the list is truncated to result_limit.
  5. Fusion: Semantic scores are blended with keyword BM25 scores using semantic_weight to produce the final ranked list.

The Embedder itself is instantiated via NewEmbedder in server/vectorstore/embedder.go (lines 81-106) when the indexer initializes. It handles request batching (respecting max_embedding_batch_size), exponential backoff for HTTP 429/5xx errors, and document chunking with ChunkText to respect max_context_length and chunk_overlap.

Complete Configuration Examples

Production-Ready YAML Setup


# ~/.config/hister/config.yml

semantic_search:
  enable: true
  embedding_endpoint: "http://localhost:11434/v1/embeddings"
  embedding_model: "qwen3-embedding:8b"
  embedding_timeout: 300
  api_key: ""
  headers:
    X-Custom-Auth: "bearer-token"
  dimensions: 4096
  max_context_length: 512
  chunk_overlap: 64
  max_embedding_batch_size: 8
  query_prefix: "query: "
  document_prefix: "passage: "
  similarity_threshold: 0.15
  result_limit: 100
  semantic_weight: 0.5
  max_embedding_concurrency: 2

Bash One-Liner with Semantic Flag

curl -G "http://localhost:4433/search" \
  --data-urlencode "q=error handling in rust" \
  -d "semantic=1" \
  -d "limit=20"

Summary

  • Enable the feature by setting semantic_search.enable: true in ~/.config/hister/config.yml and restart Hister to load the SemanticSearch struct from config/config.go.
  • Configure the embedding_endpoint to point at any OpenAI-compatible /v1/embeddings service; tune similarity_threshold and semantic_weight to balance precision and recall.
  • Toggle semantic search at runtime via Ctrl+E in the TUI, the semantic=1 HTTP parameter (parsed in server/endpoints.go), or the Semantic field in the Go client (client/search.go).
  • Understand that the indexer in server/indexer/indexer.go calls the Embedder (created in server/vectorstore/embedder.go) to vectorize queries, retrieve nearest neighbors, and fuse results with traditional keyword hits.

Frequently Asked Questions

Hister requires an endpoint implementing the OpenAI /v1/embeddings API format. Compatible providers include Ollama (default localhost:11434), OpenAI, Azure OpenAI, and any custom proxy exposing the same JSON schema. The embedding_model parameter is passed directly to the endpoint's model field.

How does Hister handle large documents that exceed the embedding model's context window?

The Embedder in server/vectorstore/embedder.go automatically chunks documents using the ChunkText function, respecting the max_context_length and chunk_overlap configuration values. Each chunk is embedded separately, and the system aggregates results to ensure no content is dropped due to token limits.

Can I adjust how much semantic results influence the final ranking?

Yes. Modify the semantic_weight parameter in config.yml (default 0.4). This value controls the interpolation between keyword (BM25) and semantic (cosine similarity) scores in the final ranking calculation performed by (*Indexer) search in server/indexer/indexer.go. Set to 1.0 for pure semantic search, 0.0 for pure keyword search.

The embedding_timeout field (default 300 seconds) controls the HTTP client timeout for embedding requests. If your endpoint is slow or you are processing large batches (controlled by max_embedding_batch_size), increase this value or reduce max_embedding_concurrency to limit parallel load on the embedding service.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →