# How to Use Semantic Search in Hyperresearch: A Complete Guide to Embedding-Based Retrieval

> Unlock semantic search in Hyperresearch with this guide. Learn to generate vector embeddings and enhance retrieval using Reciprocal Rank Fusion for smarter research.

- Repository: [Jordan Gibbs/hyperresearch](https://github.com/jordan-gibbs/hyperresearch)
- Tags: how-to-guide
- Published: 2026-09-13

---

**Hyperresearch enables semantic search by generating vector embeddings for your notes and fusing cosine similarity results with full-text rankings using Reciprocal Rank Fusion.**

Semantic search in Hyperresearch adds an embedding-based retrieval layer on top of its existing full-text engine, allowing you to find conceptually related notes even when they don't share exact keywords. This feature is implemented in the `jordan-gibbs/hyperresearch` repository and operates through a three-stage pipeline: provider configuration, vector generation, and fused retrieval.

## Configure an Embeddings Provider

Before generating vectors, you must designate an embeddings provider in your vault's configuration. Hyperresearch supports three options: `"none"` (default), `"voyage"` (Voyage AI), and `"openai"` (OpenAI).

Update your vault's [`config.toml`](https://github.com/jordan-gibbs/hyperresearch/blob/main/config.toml) under the `[embeddings]` table:

```toml
[embeddings]
provider = "voyage"          # or "openai"

model    = "voyage-3-lite"   # optional – defaults to a sensible model

body_chars = 500            # chars of a note body used for embedding

```

When `provider` is set to `"none"`, semantic search is disabled and Hyperresearch operates in full-text-only mode. The `body_chars` parameter controls how much of each note's content is vectorized, combining the title, summary, and first N characters of the body.

## Generate Embeddings for Your Notes

Once configured, you must populate the SQLite database with vector representations. The `embed_sync()` function in [`src/hyperresearch/core/embed.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/embed.py) orchestrates this process:

1. **Text extraction**: Constructs a plain-text representation via `_note_text()` (title + summary + first N body characters).
2. **Vector generation**: Calls `_http_embed()` to send text to the provider's API endpoint.
3. **Storage**: Packs vectors as `float32` blobs using `_pack()` and persists them to the `embeddings` table.
4. **Unpacking**: Retrieves vectors using `_unpack()` during search operations.

Run the sync command via CLI:

```bash
hyperresearch embed sync

```

This command walks every note in your vault, generates embeddings for new or changed content, and stores them as binary blobs in SQLite. According to the source code in [`src/hyperresearch/core/embed.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/embed.py), the system maintains a consistent schema where each embedding maps directly to a note ID for efficient lookup.

## Run Semantic-Enhanced Search

With vectors stored, activate semantic blending during search using the `--semantic` flag. The CLI entry point in [`src/hyperresearch/cli/search.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/search.py) (lines 16‑38) handles this functionality:

```bash
hyperresearch search "concurrency patterns" --semantic --limit 10

```

When `--semantic` is active, Hyperresearch executes `semantic_search()` from [`src/hyperresearch/core/embed.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/embed.py) (lines 63‑79) to perform a brute-force cosine similarity comparison between your query's embedding and stored note vectors. Results are then fused with standard full-text rankings using `reciprocal_rank_fusion()` (lines 82‑90), which applies reciprocal rank scoring to combine both result sets into a unified ranking.

Add `--json` to receive structured output containing the fused `score` field for each result.

## Programmatic Usage

You can also invoke the embedding pipeline directly from Python without using the CLI:

```python
from hyperresearch.core.vault import Vault
from hyperresearch.core.embed import embed_sync, semantic_search

# Open the vault (auto-sync will embed new notes)

vault = Vault.discover()
vault.auto_sync()               # creates embeddings if needed

# Ensure all notes have up-to-date vectors

embed_sync(vault)

# Perform a raw semantic lookup (pure cosine similarity)

hits = semantic_search(vault, "concurrency patterns", limit=5)
for hit in hits:
    print(f"Note {hit['id']} – similarity: {hit['score']:.3f}")

```

This approach gives you direct access to the `semantic_search()` function for pure vector similarity without the reciprocal rank fusion applied by the CLI layer.

## Summary

- **Provider setup**: Configure `[embeddings]` in [`config.toml`](https://github.com/jordan-gibbs/hyperresearch/blob/main/config.toml) with `provider` set to `"voyage"` or `"openai"` to enable the feature.
- **Vector storage**: The `embed_sync()` function in [`src/hyperresearch/core/embed.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/embed.py) generates `float32` vectors and stores them as blobs in the SQLite `embeddings` table.
- **CLI activation**: Use `hyperresearch search <query> --semantic` to trigger cosine similarity search fused with full-text results via `reciprocal_rank_fusion()`.
- **Python API**: Import `Vault` and `semantic_search` from `hyperresearch.core` to embed search functionality directly into your scripts.

## Frequently Asked Questions

### What embedding models does Hyperresearch support?

Hyperresearch supports Voyage AI and OpenAI providers. You can specify models like `voyage-3-lite` or any OpenAI embedding model in the [`config.toml`](https://github.com/jordan-gibbs/hyperresearch/blob/main/config.toml) file. If no model is specified, the system defaults to a sensible provider-specific option.

### How does Hyperresearch combine semantic and full-text search results?

The system uses **Reciprocal Rank Fusion** as implemented in [`src/hyperresearch/core/embed.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/embed.py). This algorithm takes the ranked lists from both the cosine similarity search (`semantic_search()`) and the traditional full-text search, then computes a fused score based on the reciprocal of each result's rank in both lists.

### Where are the embedding vectors stored?

Vectors are stored as `float32` binary blobs in the SQLite `embeddings` table within your vault database. The `_pack()` and `_unpack()` helper functions in [`src/hyperresearch/core/embed.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/embed.py) handle serialization and deserialization of these vectors.

### Do I need to regenerate embeddings after editing notes?

Yes. Run `hyperresearch embed sync` after modifying notes to update their vector representations. If using the Python API, calling `embed_sync(vault)` ensures all current notes have fresh embeddings before performing semantic searches.