# How Hister Combines Keyword and Semantic Search Results

> Discover how Hister merges keyword and semantic search results. Learn its unique approach to combining full-text and vector searches for comprehensive results.

- Repository: [Adam Tauber/hister](https://github.com/asciimoo/hister)
- Tags: deep-dive
- Published: 2026-09-01

---

**Hister merges keyword and semantic results by executing full-text and vector searches independently, converting each vector hit into a Document with a truncated preview, concatenating both result sets, and adjusting the total count to include semantic-only matches.**

The open-source search engine **Hister** (asciimoo/hister) implements a hybrid retrieval pipeline that unifies traditional inverted-index matching with dense vector similarity search. Understanding how Hister combines keyword and semantic search results reveals the architecture behind its federated ranking system.

## Dual-Path Search Execution

Hister’s search logic operates along two parallel tracks within [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go). First, it executes a classic full-text query against the inverted index to collect keyword matches. Simultaneously, if semantic search is enabled and the embedding subsystem is healthy, it prepares a second query stream.

The semantic path begins by sanitizing the user’s input. The indexer calls `querybuilder.RemoveStandaloneWildcards(expression.Text)` to strip lone wildcard characters that degrade embedding quality, then validates that the remaining string is not empty before proceeding to vectorization.

## Processing the Semantic Query

### Generating Query Embeddings

With the cleaned text in hand, Hister invokes `i.embedder.EmbedQuery(ctx, semanticText)` to create a dense vector representation of the query. This embedding captures the semantic meaning of the query rather than literal token matches.

### Vector Store Retrieval

The embedding is passed to `i.vectorStore.Search(ctx, vec, threshold, resultLimit)`, which queries the underlying vector store—defined in [`server/vectorstore/vectorstore.go`](https://github.com/asciimoo/hister/blob/main/server/vectorstore/vectorstore.go)—using the configured **semantic_threshold** (minimum similarity score) and **resultLimit** (maximum hits to return). According to the source, if the embedding step returns an error, Hister silently disables semantic search for that request and falls back to keyword-only results [【L1766-L1773】](https://github.com/asciimoo/hister/blob/master/server/indexer/indexer.go#L1766-L1773).

## Merging Result Sets

### Converting Vector Hits to Documents

Each semantic hit returned by the vector store must be normalized into the common `Document` type before merging. In [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go), Hister iterates over semantic results and constructs `Document` objects where the `MatchedChunk` field contains a truncated preview of the matching text segment. The truncation is hard-limited to **semanticTextPreviewLen** (500 runes) via the `truncateText` helper to prevent oversized payloads [【L1817-L1823】](https://github.com/asciimoo/hister/blob/master/server/indexer/indexer.go#L1817-L1823).

```go
// Convert semantic hits to Documents with truncated previews
for _, hit := range semanticHits {
    d := Document{
        URL:          hit.URL,
        MatchedChunk: truncateText(hit.ChunkText, semanticTextPreviewLen),
    }
    results.Documents = append(results.Documents, d)
    semanticOnlyCnt++
}

```

### Concatenation and Count Adjustment

Rather than interleaving scores, Hister concatenates the keyword and semantic result lists. After appending all semantic Documents, it recalculates the total hit count to account for matches that appear only in the semantic set. The logic sets `results.Total` to the maximum of the existing keyword total and the combined length of the Document slice plus the semantic-only counter [【L1833-L1841】](https://github.com/asciimoo/hister/blob/master/server/indexer/indexer.go#L1833-L1841).

```go
// Merge totals to include semantic-only hits
results.Total = max(results.Total, uint64(len(results.Documents))+semanticOnlyCnt)

```

### Graceful Degradation

If the embedding subsystem is unavailable or returns an error during the `EmbedQuery` call, Hister catches the failure early and skips the vector search phase entirely, ensuring the user still receives keyword results without interruption [【L1766-L1773】](https://github.com/asciimoo/hister/blob/master/server/indexer/indexer.go#L1766-L1773).

## Client-Side Configuration

The merging behavior is governed by two tunable parameters: **semantic_weight** (influencing ranking when scores are combined) and **semantic_threshold** (filtering minimum vector similarity). These are exposed to the client through the `/search` endpoint defined in [`server/mcp.go`](https://github.com/asciimoo/hister/blob/main/server/mcp.go), which also returns a boolean `semantic_enabled` flag indicating whether the response includes semantic matches [【L227-L229】](https://github.com/asciimoo/hister/blob/master/server/mcp.go#L227-L229).

From the client side, [`client/search.go`](https://github.com/asciimoo/hister/blob/main/client/search.go) constructs the search request, conditionally including the semantic flag based on CLI or TUI settings, then sends the payload to the server’s hybrid search handler.

## Summary

- **Parallel execution**: Hister runs keyword and semantic searches as independent operations inside [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go).
- **Query sanitization**: Standalone wildcards are removed via `querybuilder.RemoveStandaloneWildcards` before embedding generation.
- **Preview truncation**: Semantic matches are limited to 500 runes using `truncateText` to keep payloads concise.
- **Concatenation strategy**: Result lists are appended, and `results.Total` is adjusted upward to count semantic-only documents.
- **Fault tolerance**: Embedding failures trigger an immediate fallback to keyword-only search without user-visible errors.
- **Configuration**: `semantic_weight`, `semantic_threshold`, and the `semantic_enabled` flag control hybrid behavior and client display.

## Frequently Asked Questions

### How does Hister handle wildcard characters in semantic queries?

Hister strips standalone wildcard characters from the query string before generating embeddings. It calls `querybuilder.RemoveStandaloneWildcards(expression.Text)` to clean the input, ensuring that wildcard symbols do not pollute the vector representation. If the cleaned string is empty, the semantic phase is skipped.

### What is the maximum preview length for semantic search matches?

Semantic search previews are truncated to **500 runes** via the `semanticTextPreviewLen` constant. When converting vector store hits into `Document` objects, Hister applies `truncateText(hit.ChunkText, semanticTextPreviewLen)` to limit the `MatchedChunk` field, preventing large text blocks from being returned to the client.

### How does Hister respond if the embedding service is unavailable?

If `i.embedder.EmbedQuery` returns an error, Hister immediately aborts the semantic search phase and falls back to keyword-only results. This graceful degradation ensures that temporary embedding subsystem failures do not break the search experience, as implemented in the error handling block at lines 1766–1773 of [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go).

### What parameters control the influence of semantic results?

Two primary configuration options govern semantic behavior: **semantic_threshold**, which sets the minimum similarity score for vector matches, and **semantic_weight**, which influences how semantic scores combine with keyword relevance. The server exposes these to the client alongside a `semantic_enabled` boolean flag in [`server/mcp.go`](https://github.com/asciimoo/hister/blob/main/server/mcp.go), allowing the frontend to adjust ranking displays accordingly.