# How to Synchronize Embeddings in Hyperresearch: CLI and API Guide

> Learn to synchronize embeddings in Hyperresearch using the CLI and API. Update your research notes' vector representations for fast semantic search with Voyage AI or OpenAI.

- Repository: [Jordan Gibbs/hyperresearch](https://github.com/jordan-gibbs/hyperresearch)
- Tags: how-to-guide
- Published: 2026-09-13

---

**Embedding synchronization in Hyperresearch scans your vault for modified notes, generates vector representations via Voyage AI or OpenAI, and updates the SQLite `embeddings` table to enable fast semantic search across your research.**

Hyperresearch is an open-source research management tool that stores a semantic-search index by embedding each note's content into high-dimensional vectors. To keep search results accurate, you must synchronize embeddings whenever note content changes, ensuring the vector database reflects the current state of your vault.

## How Embedding Synchronization Works

The synchronization process compares each note's current content hash against the stored embedding metadata to identify "dirty" notes requiring reprocessing.

### Architecture Overview

The embedding system consists of four core components defined in the `jordan-gibbs/hyperresearch` repository:

- **`hyperresearch.core.embed`** – Core logic residing in [`src/hyperresearch/core/embed.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/embed.py) that detects stale embeddings, batches texts, and calls provider APIs via `_http_embed`.
- **`hyperresearch.core.config.EmbeddingSettings`** – Configuration schema in [`src/hyperresearch/core/config.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/config.py) (line 136) that stores the provider name, model, and truncation settings.
- **SQLite `embeddings` table** – Persistence layer defined in [`src/hyperresearch/core/db.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/db.py) (line 109) storing `note_id`, `model` identifier, `dimensions`, binary `vector` BLOBs, and timestamps.
- **`hyperresearch.cli.embed_cmd`** – CLI entry point in [`src/hyperresearch/cli/embed_cmd.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/embed_cmd.py) (line 14) exposing the `hyperresearch embed sync` command.

### The Synchronization Pipeline

According to the source code in [`embed.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/embed.py), the `embed_sync` function executes the following steps:

1. **Load configuration** – Reads `vault.config.embeddings` to obtain the provider, model, and `body_chars` limit.
2. **Select candidates** – SQL joins the `notes`, `note_content`, and `embeddings` tables to find all notes.
3. **Detect dirty notes** – Computes `stamp = f"{model}@{row['content_hash']}"` for each note. If the stored `embedded_model` differs from this stamp, the note is queued for re-embedding.
4. **Batch processing** – Groups notes into batches of 32, constructing text from title, summary, and the first N body characters via `_note_text`.
5. **Provider API call** – Sends batched texts to `_http_embed`, which dispatches to Voyage AI or OpenAI endpoints.
6. **Vector persistence** – Packs vectors using `struct.pack` (via `_pack`) and writes them to the `embeddings` table with the current timestamp.
7. **Return summary** – Yields a dictionary containing `embedded` count, `skipped` count, and provider metadata.

If `embeddings.provider` is set to `"none"`, the function raises `EmbeddingError` and aborts.

## Configuring Embedding Providers

Before synchronizing, configure your provider credentials in `~/.hyperresearch/config.toml`. The `EmbeddingSettings` class validates these values:

```toml
[embeddings]
provider = "voyage"  # or "openai"

model = "voyage-3-lite"
body_chars = 1000

```

You can also set these via CLI:

```bash
hyperresearch config set embeddings.provider voyage
hyperresearch config set embeddings.model voyage-3-lite

```

## Running Embedding Synchronization

Hyperresearch offers three interfaces for triggering synchronization: the command-line tool, the Python API, and low-level provider access.

### CLI Method

The fastest way to synchronize embeddings across your entire vault uses the `embed sync` subcommand:

```bash
hyperresearch embed sync

```

This command invokes `embed_sync` from [`src/hyperresearch/cli/embed_cmd.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/embed_cmd.py) and prints a summary:

```

[green]Embedded:[/] 42 notes (provider: voyage, model: voyage-3-lite)

```

### Programmatic API

For custom scripts or plugins, import `embed_sync` directly from the core module:

```python
from hyperresearch.core.embed import embed_sync, semantic_search, EmbeddingError

# Assume `vault` is a hyperresearch.Vault instance

try:
    result = embed_sync(vault)
    print(f"Embedded {result['embedded']} notes, skipped {result['skipped']}")
except EmbeddingError as exc:
    print(f"Embedding failed: {exc}")

# Query after sync

hits = semantic_search(vault, "machine learning research", limit=5)
for hit in hits:
    print(hit["id"], f"{hit['score']:.3f}")

```

### Low-Level Provider Access

To debug provider responses or embed arbitrary text without the vault logic, use the internal `_http_embed` function:

```python
from hyperresearch.core.embed import _http_embed

vectors = _http_embed(
    provider="openai",
    model="text-embedding-3-small",
    texts=["What is vector search?"]
)

```

This function requires the appropriate environment variable (`OPENAI_API_KEY` or `VOYAGE_API_KEY`) and returns raw vector lists.

## Summary

- **Synchronization** compares content hashes to detect notes needing re-embedding.
- **Configuration** happens in `~/.hyperresearch/config.toml` via `EmbeddingSettings`.
- **Core logic** resides in [`src/hyperresearch/core/embed.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/embed.py), specifically the `embed_sync` function.
- **Storage** uses SQLite BLOBs in the `embeddings` table with model-specific stamps.
- **Interfaces** include CLI (`hyperresearch embed sync`), Python API (`embed_sync`), and low-level `_http_embed`.

## Frequently Asked Questions

### What triggers a note to be re-embedded?

A note is flagged as dirty when its computed stamp—formatted as `f"{model}@{content_hash}"`—differs from the `embedded_model` value stored in the `embeddings` table. This occurs when the note content changes or when you switch to a different embedding model.

### Which embedding providers does Hyperresearch support?

Hyperresearch supports **Voyage AI** and **OpenAI**. Configure your choice via `embeddings.provider` in [`config.toml`](https://github.com/jordan-gibbs/hyperresearch/blob/main/config.toml). If set to `"none"`, the system raises an `EmbeddingError` and refuses to sync.

### How are embeddings stored in the database?

Vectors are stored as binary BLOBs in the `embeddings` table defined in [`src/hyperresearch/core/db.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/db.py). Each row contains `note_id`, a model identifier string (including hash), `dimensions`, the packed vector data, and a `created_at` timestamp.

### Can I adjust how much note content is embedded?

Yes. The `body_chars` setting in `EmbeddingSettings` controls how many characters from the note body are included alongside the title and summary. Reducing this value speeds up synchronization and reduces API costs for long documents.