How Semantic Search Works in Hister: A Deep Dive into the Vector Pipeline

Semantic search in Hister uses a three-stage pipeline that converts queries into dense vectors via OpenAI-compatible APIs, retrieves similar document chunks from SQLite or PostgreSQL vector stores, and applies post-processing logic to return diversified, relevant results.

Hister, an open-source search engine by asciimoo/hister, implements semantic search through a modular architecture that separates embedding generation, vector storage, and result curation. The system transforms natural language queries into mathematical representations and performs approximate nearest neighbor searches against pre-indexed document chunks. Understanding this implementation reveals how modern vector databases integrate with traditional search architectures.

Overview of the Semantic Search Pipeline

The implementation follows a predictable data flow from query to result. First, the Embedder component converts text into high-dimensional vectors. Next, the VectorStore interface executes similarity searches against these vectors. Finally, post-processing functions merge and diversify results to prevent duplicate documents from dominating the output. This pipeline is accessible via the Model Context Protocol (MCP) tool, CLI commands, and internal API calls.

Stage 1: Query Embedding with the Embedder

The embedding stage converts textual queries into []float32 vectors suitable for mathematical comparison. This occurs in server/vectorstore/embedder.go, where the Embedder struct manages connections to external embedding services.

Configuration and Initialization

The NewEmbedder function initializes the component from the semantic-search configuration block. It supports OpenAI-compatible /v1/embeddings endpoints, allowing integration with various providers (OpenAI, local models via Ollama, or other API-compatible services).

// Create the embedder from the configuration
embedder := vectorstore.NewEmbedder(&cfg.SemanticSearch)

The Embedding Request Flow

When processing a search query, the EmbedQuery method handles the transformation. It applies an optional QueryPrefix (configured per model requirements), then delegates to Embed and doEmbeddingRequest. The latter manages HTTP retries, concurrency limits, and context-length validation to ensure robust operation under load.

The core embedding logic resides in:

func (e *Embedder) EmbedQuery(ctx context.Context, text string) ([]float32, error)

This method returns a dense vector representation that captures the semantic meaning of the input text, enabling similarity comparisons against stored document chunks.

Stage 2: Vector Storage and Retrieval

Once vectorized, queries are matched against a database of pre-computed document embeddings. Hister abstracts storage behind a unified interface while supporting multiple backend implementations.

The VectorStore Interface

The VectorStore interface in server/vectorstore/vectorstore.go defines the contract for all storage backends:

type VectorStore interface {
    Search(vector []float32, topK int, threshold float64, userID uint) ([]Result, error)
}

The Search method accepts the query vector, a maximum result count (topK), a similarity threshold, and a user ID for access control. It returns a slice of Result structs containing DocID, ChunkIdx, ChunkText, and Similarity scores.

SQLite and PostgreSQL Backends

Hister ships with two production-ready backends:

  • SQLite: Uses the vector extension for efficient similarity searches
  • PostgreSQL: Leverages the pgvector extension for enterprise deployments

Both implementations share the same interface but optimize for their respective storage engines. The SQLite entry point in server/vectorstore/sqlite.go demonstrates the search implementation:

func (s *sqliteVectorStore) Search(vector []float32, topK int, threshold float64, userID uint) ([]Result, error)

Document indexing involves splitting content into chunks via ChunkText (defined in tokenizer.go), computing embeddings for each chunk, and storing them as rows in the selected vector database.

Stage 3: Result Processing and Diversification

Raw vector search often returns multiple chunks from the same document, creating redundancy. Hister implements a two-step post-processing pipeline to curate the final result set.

Merging Duplicate Hits

The mergeSearchResults function (lines 31-71 in vectorstore.go) consolidates multiple chunk hits from the same document. It aggregates similarity scores and metadata, ensuring each document appears only once in the intermediate results while preserving the highest-scoring chunks.

Diversifying Results

The diversifySearchResults function (lines 81-107) applies business logic to create a balanced result set. It enforces two critical constraints:

  • maxChunksPerDocument: Limits how many chunks from a single document can appear
  • documentLimit: Caps the total number of unique documents returned

This prevents a single highly-relevant document from monopolizing the search results, ensuring users see diverse sources.

Integration with the MCP Tool and CLI

The semantic search pipeline integrates deeply with Hister's Model Context Protocol (MCP) implementation in server/mcp.go. The mcpToolSearch function orchestrates the full workflow:

  1. Calls doSearch to process the query
  2. Executes vectorstore.Search to retrieve candidates
  3. Applies diversifySearchResults to curate output
  4. Formats results via mcpBuildSearchResult

For command-line users, the hister search command abstracts this complexity:

hister search "best practices for Go testing"

This command internally parses the text, executes the embedding, queries the vector store, and prints diversified results to the terminal.

Complete Code Example

The following example demonstrates the full semantic search workflow programmatically:

// 1. Initialize the embedder
embedder := vectorstore.NewEmbedder(&cfg.SemanticSearch)

// 2. Vectorize the user query
vec, err := embedder.EmbedQuery(ctx, "how to reset my password")
if err != nil {
    log.Fatal(err)
}

// 3. Query the vector store (SQLite example)
store, _ := vectorstore.New(&cfg)
results, err := store.Search(vec, 10, 0.75, 0) // top-10, similarity ≥0.75
if err != nil {
    log.Fatal(err)
}

// 4. Results are already diversified - iterate directly
for _, r := range results {
    fmt.Printf("Doc %s (chunk %d) – %.2f similarity\n%s\n\n",
        r.DocID, r.ChunkIdx, r.Similarity, r.ChunkText)
}

Summary

  • Hister implements semantic search through a three-stage pipeline involving query embedding, vector retrieval, and result diversification.
  • The Embedder component in server/vectorstore/embedder.go handles OpenAI-compatible API communication with retry logic and concurrency controls.
  • Two storage backends (SQLite with vector extension and PostgreSQL with pgvector) implement the VectorStore interface defined in vectorstore.go.
  • Post-processing functions mergeSearchResults and diversifySearchResults prevent result redundancy and ensure diverse document coverage.
  • MCP integration exposes semantic search to web interfaces and external tools, while CLI commands provide direct terminal access.

Frequently Asked Questions

What embedding models does Hister support?

Hister supports any embedding model accessible via an OpenAI-compatible /v1/embeddings endpoint. This includes OpenAI's text-embedding models, local deployments through Ollama, or any custom API server implementing the same protocol. Configuration occurs through the SemanticSearch config block when initializing the NewEmbedder.

How does Hister handle document chunking?

During indexing, Hister splits documents into manageable chunks using the ChunkText function in server/tokenizer.go. Each chunk is embedded separately and stored with metadata including document ID and chunk index. This chunking strategy allows precise retrieval of relevant passages rather than requiring entire documents to match the query.

Can I use PostgreSQL instead of SQLite for vector storage?

Yes. Hister provides both SQLiteVectorStore and PostgresVectorStore implementations of the VectorStore interface. The PostgreSQL backend utilizes the pgvector extension and is selectable through configuration. Both backends support the same Search method signature and diversification logic, allowing seamless swapping based on deployment requirements.

What is the purpose of the diversification step?

The diversification step (diversifySearchResults) ensures that search results provide broad coverage rather than overwhelming users with multiple chunks from the same document. By enforcing maxChunksPerDocument and documentLimit constraints, Hister prevents scenarios where one highly-scored document dominates the results, improving the discovery of varied sources.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →