# How Semantic Search Works in Hister: A Deep Dive into the Vector Pipeline

> Explore Hister's semantic search pipeline: convert queries to vectors, retrieve similar documents from vector stores, and get diversified results. Learn how it works.

- Repository: [Adam Tauber/hister](https://github.com/asciimoo/hister)
- Tags: deep-dive
- Published: 2026-09-01

---

**Semantic search in Hister uses a three-stage pipeline that converts queries into dense vectors via OpenAI-compatible APIs, retrieves similar document chunks from SQLite or PostgreSQL vector stores, and applies post-processing logic to return diversified, relevant results.**

Hister, an open-source search engine by asciimoo/hister, implements semantic search through a modular architecture that separates embedding generation, vector storage, and result curation. The system transforms natural language queries into mathematical representations and performs approximate nearest neighbor searches against pre-indexed document chunks. Understanding this implementation reveals how modern vector databases integrate with traditional search architectures.

## Overview of the Semantic Search Pipeline

The implementation follows a predictable data flow from query to result. First, the `Embedder` component converts text into high-dimensional vectors. Next, the `VectorStore` interface executes similarity searches against these vectors. Finally, post-processing functions merge and diversify results to prevent duplicate documents from dominating the output. This pipeline is accessible via the Model Context Protocol (MCP) tool, CLI commands, and internal API calls.

## Stage 1: Query Embedding with the Embedder

The embedding stage converts textual queries into `[]float32` vectors suitable for mathematical comparison. This occurs in [`server/vectorstore/embedder.go`](https://github.com/asciimoo/hister/blob/main/server/vectorstore/embedder.go), where the `Embedder` struct manages connections to external embedding services.

### Configuration and Initialization

The `NewEmbedder` function initializes the component from the semantic-search configuration block. It supports OpenAI-compatible `/v1/embeddings` endpoints, allowing integration with various providers (OpenAI, local models via Ollama, or other API-compatible services).

```go
// Create the embedder from the configuration
embedder := vectorstore.NewEmbedder(&cfg.SemanticSearch)

```

### The Embedding Request Flow

When processing a search query, the `EmbedQuery` method handles the transformation. It applies an optional `QueryPrefix` (configured per model requirements), then delegates to `Embed` and `doEmbeddingRequest`. The latter manages HTTP retries, concurrency limits, and context-length validation to ensure robust operation under load.

The core embedding logic resides in:

```go
func (e *Embedder) EmbedQuery(ctx context.Context, text string) ([]float32, error)

```

This method returns a dense vector representation that captures the semantic meaning of the input text, enabling similarity comparisons against stored document chunks.

## Stage 2: Vector Storage and Retrieval

Once vectorized, queries are matched against a database of pre-computed document embeddings. Hister abstracts storage behind a unified interface while supporting multiple backend implementations.

### The VectorStore Interface

The `VectorStore` interface in [`server/vectorstore/vectorstore.go`](https://github.com/asciimoo/hister/blob/main/server/vectorstore/vectorstore.go) defines the contract for all storage backends:

```go
type VectorStore interface {
    Search(vector []float32, topK int, threshold float64, userID uint) ([]Result, error)
}

```

The `Search` method accepts the query vector, a maximum result count (`topK`), a similarity threshold, and a user ID for access control. It returns a slice of `Result` structs containing `DocID`, `ChunkIdx`, `ChunkText`, and `Similarity` scores.

### SQLite and PostgreSQL Backends

Hister ships with two production-ready backends:

- **SQLite**: Uses the `vector` extension for efficient similarity searches
- **PostgreSQL**: Leverages the `pgvector` extension for enterprise deployments

Both implementations share the same interface but optimize for their respective storage engines. The SQLite entry point in [`server/vectorstore/sqlite.go`](https://github.com/asciimoo/hister/blob/main/server/vectorstore/sqlite.go) demonstrates the search implementation:

```go
func (s *sqliteVectorStore) Search(vector []float32, topK int, threshold float64, userID uint) ([]Result, error)

```

Document indexing involves splitting content into chunks via `ChunkText` (defined in [`tokenizer.go`](https://github.com/asciimoo/hister/blob/main/tokenizer.go)), computing embeddings for each chunk, and storing them as rows in the selected vector database.

## Stage 3: Result Processing and Diversification

Raw vector search often returns multiple chunks from the same document, creating redundancy. Hister implements a two-step post-processing pipeline to curate the final result set.

### Merging Duplicate Hits

The `mergeSearchResults` function (lines 31-71 in [`vectorstore.go`](https://github.com/asciimoo/hister/blob/main/vectorstore.go)) consolidates multiple chunk hits from the same document. It aggregates similarity scores and metadata, ensuring each document appears only once in the intermediate results while preserving the highest-scoring chunks.

### Diversifying Results

The `diversifySearchResults` function (lines 81-107) applies business logic to create a balanced result set. It enforces two critical constraints:

- `maxChunksPerDocument`: Limits how many chunks from a single document can appear
- `documentLimit`: Caps the total number of unique documents returned

This prevents a single highly-relevant document from monopolizing the search results, ensuring users see diverse sources.

## Integration with the MCP Tool and CLI

The semantic search pipeline integrates deeply with Hister's Model Context Protocol (MCP) implementation in [`server/mcp.go`](https://github.com/asciimoo/hister/blob/main/server/mcp.go). The `mcpToolSearch` function orchestrates the full workflow:

1. Calls `doSearch` to process the query
2. Executes `vectorstore.Search` to retrieve candidates
3. Applies `diversifySearchResults` to curate output
4. Formats results via `mcpBuildSearchResult`

For command-line users, the `hister search` command abstracts this complexity:

```bash
hister search "best practices for Go testing"

```

This command internally parses the text, executes the embedding, queries the vector store, and prints diversified results to the terminal.

## Complete Code Example

The following example demonstrates the full semantic search workflow programmatically:

```go
// 1. Initialize the embedder
embedder := vectorstore.NewEmbedder(&cfg.SemanticSearch)

// 2. Vectorize the user query
vec, err := embedder.EmbedQuery(ctx, "how to reset my password")
if err != nil {
    log.Fatal(err)
}

// 3. Query the vector store (SQLite example)
store, _ := vectorstore.New(&cfg)
results, err := store.Search(vec, 10, 0.75, 0) // top-10, similarity ≥0.75
if err != nil {
    log.Fatal(err)
}

// 4. Results are already diversified - iterate directly
for _, r := range results {
    fmt.Printf("Doc %s (chunk %d) – %.2f similarity\n%s\n\n",
        r.DocID, r.ChunkIdx, r.Similarity, r.ChunkText)
}

```

## Summary

- **Hister implements semantic search** through a three-stage pipeline involving query embedding, vector retrieval, and result diversification.
- **The `Embedder` component** in [`server/vectorstore/embedder.go`](https://github.com/asciimoo/hister/blob/main/server/vectorstore/embedder.go) handles OpenAI-compatible API communication with retry logic and concurrency controls.
- **Two storage backends** (SQLite with `vector` extension and PostgreSQL with `pgvector`) implement the `VectorStore` interface defined in [`vectorstore.go`](https://github.com/asciimoo/hister/blob/main/vectorstore.go).
- **Post-processing functions** `mergeSearchResults` and `diversifySearchResults` prevent result redundancy and ensure diverse document coverage.
- **MCP integration** exposes semantic search to web interfaces and external tools, while CLI commands provide direct terminal access.

## Frequently Asked Questions

### What embedding models does Hister support?

Hister supports any embedding model accessible via an OpenAI-compatible `/v1/embeddings` endpoint. This includes OpenAI's text-embedding models, local deployments through Ollama, or any custom API server implementing the same protocol. Configuration occurs through the `SemanticSearch` config block when initializing the `NewEmbedder`.

### How does Hister handle document chunking?

During indexing, Hister splits documents into manageable chunks using the `ChunkText` function in [`server/tokenizer.go`](https://github.com/asciimoo/hister/blob/main/server/tokenizer.go). Each chunk is embedded separately and stored with metadata including document ID and chunk index. This chunking strategy allows precise retrieval of relevant passages rather than requiring entire documents to match the query.

### Can I use PostgreSQL instead of SQLite for vector storage?

Yes. Hister provides both `SQLiteVectorStore` and `PostgresVectorStore` implementations of the `VectorStore` interface. The PostgreSQL backend utilizes the `pgvector` extension and is selectable through configuration. Both backends support the same `Search` method signature and diversification logic, allowing seamless swapping based on deployment requirements.

### What is the purpose of the diversification step?

The diversification step (`diversifySearchResults`) ensures that search results provide broad coverage rather than overwhelming users with multiple chunks from the same document. By enforcing `maxChunksPerDocument` and `documentLimit` constraints, Hister prevents scenarios where one highly-scored document dominates the results, improving the discovery of varied sources.