How to Synchronize Embeddings in Hyperresearch: CLI and API Guide

Embedding synchronization in Hyperresearch scans your vault for modified notes, generates vector representations via Voyage AI or OpenAI, and updates the SQLite embeddings table to enable fast semantic search across your research.

Hyperresearch is an open-source research management tool that stores a semantic-search index by embedding each note's content into high-dimensional vectors. To keep search results accurate, you must synchronize embeddings whenever note content changes, ensuring the vector database reflects the current state of your vault.

How Embedding Synchronization Works

The synchronization process compares each note's current content hash against the stored embedding metadata to identify "dirty" notes requiring reprocessing.

Architecture Overview

The embedding system consists of four core components defined in the jordan-gibbs/hyperresearch repository:

  • hyperresearch.core.embed – Core logic residing in src/hyperresearch/core/embed.py that detects stale embeddings, batches texts, and calls provider APIs via _http_embed.
  • hyperresearch.core.config.EmbeddingSettings – Configuration schema in src/hyperresearch/core/config.py (line 136) that stores the provider name, model, and truncation settings.
  • SQLite embeddings table – Persistence layer defined in src/hyperresearch/core/db.py (line 109) storing note_id, model identifier, dimensions, binary vector BLOBs, and timestamps.
  • hyperresearch.cli.embed_cmd – CLI entry point in src/hyperresearch/cli/embed_cmd.py (line 14) exposing the hyperresearch embed sync command.

The Synchronization Pipeline

According to the source code in embed.py, the embed_sync function executes the following steps:

  1. Load configuration – Reads vault.config.embeddings to obtain the provider, model, and body_chars limit.
  2. Select candidates – SQL joins the notes, note_content, and embeddings tables to find all notes.
  3. Detect dirty notes – Computes stamp = f"{model}@{row['content_hash']}" for each note. If the stored embedded_model differs from this stamp, the note is queued for re-embedding.
  4. Batch processing – Groups notes into batches of 32, constructing text from title, summary, and the first N body characters via _note_text.
  5. Provider API call – Sends batched texts to _http_embed, which dispatches to Voyage AI or OpenAI endpoints.
  6. Vector persistence – Packs vectors using struct.pack (via _pack) and writes them to the embeddings table with the current timestamp.
  7. Return summary – Yields a dictionary containing embedded count, skipped count, and provider metadata.

If embeddings.provider is set to "none", the function raises EmbeddingError and aborts.

Configuring Embedding Providers

Before synchronizing, configure your provider credentials in ~/.hyperresearch/config.toml. The EmbeddingSettings class validates these values:

[embeddings]
provider = "voyage"  # or "openai"

model = "voyage-3-lite"
body_chars = 1000

You can also set these via CLI:

hyperresearch config set embeddings.provider voyage
hyperresearch config set embeddings.model voyage-3-lite

Running Embedding Synchronization

Hyperresearch offers three interfaces for triggering synchronization: the command-line tool, the Python API, and low-level provider access.

CLI Method

The fastest way to synchronize embeddings across your entire vault uses the embed sync subcommand:

hyperresearch embed sync

This command invokes embed_sync from src/hyperresearch/cli/embed_cmd.py and prints a summary:


[green]Embedded:[/] 42 notes (provider: voyage, model: voyage-3-lite)

Programmatic API

For custom scripts or plugins, import embed_sync directly from the core module:

from hyperresearch.core.embed import embed_sync, semantic_search, EmbeddingError

# Assume `vault` is a hyperresearch.Vault instance

try:
    result = embed_sync(vault)
    print(f"Embedded {result['embedded']} notes, skipped {result['skipped']}")
except EmbeddingError as exc:
    print(f"Embedding failed: {exc}")

# Query after sync

hits = semantic_search(vault, "machine learning research", limit=5)
for hit in hits:
    print(hit["id"], f"{hit['score']:.3f}")

Low-Level Provider Access

To debug provider responses or embed arbitrary text without the vault logic, use the internal _http_embed function:

from hyperresearch.core.embed import _http_embed

vectors = _http_embed(
    provider="openai",
    model="text-embedding-3-small",
    texts=["What is vector search?"]
)

This function requires the appropriate environment variable (OPENAI_API_KEY or VOYAGE_API_KEY) and returns raw vector lists.

Summary

  • Synchronization compares content hashes to detect notes needing re-embedding.
  • Configuration happens in ~/.hyperresearch/config.toml via EmbeddingSettings.
  • Core logic resides in src/hyperresearch/core/embed.py, specifically the embed_sync function.
  • Storage uses SQLite BLOBs in the embeddings table with model-specific stamps.
  • Interfaces include CLI (hyperresearch embed sync), Python API (embed_sync), and low-level _http_embed.

Frequently Asked Questions

What triggers a note to be re-embedded?

A note is flagged as dirty when its computed stamp—formatted as f"{model}@{content_hash}"—differs from the embedded_model value stored in the embeddings table. This occurs when the note content changes or when you switch to a different embedding model.

Which embedding providers does Hyperresearch support?

Hyperresearch supports Voyage AI and OpenAI. Configure your choice via embeddings.provider in config.toml. If set to "none", the system raises an EmbeddingError and refuses to sync.

How are embeddings stored in the database?

Vectors are stored as binary BLOBs in the embeddings table defined in src/hyperresearch/core/db.py. Each row contains note_id, a model identifier string (including hash), dimensions, the packed vector data, and a created_at timestamp.

Can I adjust how much note content is embedded?

Yes. The body_chars setting in EmbeddingSettings controls how many characters from the note body are included alongside the title and summary. Reducing this value speeds up synchronization and reduces API costs for long documents.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →