How to Synchronize Embeddings in Hyperresearch: CLI and API Guide
Embedding synchronization in Hyperresearch scans your vault for modified notes, generates vector representations via Voyage AI or OpenAI, and updates the SQLite embeddings table to enable fast semantic search across your research.
Hyperresearch is an open-source research management tool that stores a semantic-search index by embedding each note's content into high-dimensional vectors. To keep search results accurate, you must synchronize embeddings whenever note content changes, ensuring the vector database reflects the current state of your vault.
How Embedding Synchronization Works
The synchronization process compares each note's current content hash against the stored embedding metadata to identify "dirty" notes requiring reprocessing.
Architecture Overview
The embedding system consists of four core components defined in the jordan-gibbs/hyperresearch repository:
hyperresearch.core.embed– Core logic residing insrc/hyperresearch/core/embed.pythat detects stale embeddings, batches texts, and calls provider APIs via_http_embed.hyperresearch.core.config.EmbeddingSettings– Configuration schema insrc/hyperresearch/core/config.py(line 136) that stores the provider name, model, and truncation settings.- SQLite
embeddingstable – Persistence layer defined insrc/hyperresearch/core/db.py(line 109) storingnote_id,modelidentifier,dimensions, binaryvectorBLOBs, and timestamps. hyperresearch.cli.embed_cmd– CLI entry point insrc/hyperresearch/cli/embed_cmd.py(line 14) exposing thehyperresearch embed synccommand.
The Synchronization Pipeline
According to the source code in embed.py, the embed_sync function executes the following steps:
- Load configuration – Reads
vault.config.embeddingsto obtain the provider, model, andbody_charslimit. - Select candidates – SQL joins the
notes,note_content, andembeddingstables to find all notes. - Detect dirty notes – Computes
stamp = f"{model}@{row['content_hash']}"for each note. If the storedembedded_modeldiffers from this stamp, the note is queued for re-embedding. - Batch processing – Groups notes into batches of 32, constructing text from title, summary, and the first N body characters via
_note_text. - Provider API call – Sends batched texts to
_http_embed, which dispatches to Voyage AI or OpenAI endpoints. - Vector persistence – Packs vectors using
struct.pack(via_pack) and writes them to theembeddingstable with the current timestamp. - Return summary – Yields a dictionary containing
embeddedcount,skippedcount, and provider metadata.
If embeddings.provider is set to "none", the function raises EmbeddingError and aborts.
Configuring Embedding Providers
Before synchronizing, configure your provider credentials in ~/.hyperresearch/config.toml. The EmbeddingSettings class validates these values:
[embeddings]
provider = "voyage" # or "openai"
model = "voyage-3-lite"
body_chars = 1000
You can also set these via CLI:
hyperresearch config set embeddings.provider voyage
hyperresearch config set embeddings.model voyage-3-lite
Running Embedding Synchronization
Hyperresearch offers three interfaces for triggering synchronization: the command-line tool, the Python API, and low-level provider access.
CLI Method
The fastest way to synchronize embeddings across your entire vault uses the embed sync subcommand:
hyperresearch embed sync
This command invokes embed_sync from src/hyperresearch/cli/embed_cmd.py and prints a summary:
[green]Embedded:[/] 42 notes (provider: voyage, model: voyage-3-lite)
Programmatic API
For custom scripts or plugins, import embed_sync directly from the core module:
from hyperresearch.core.embed import embed_sync, semantic_search, EmbeddingError
# Assume `vault` is a hyperresearch.Vault instance
try:
result = embed_sync(vault)
print(f"Embedded {result['embedded']} notes, skipped {result['skipped']}")
except EmbeddingError as exc:
print(f"Embedding failed: {exc}")
# Query after sync
hits = semantic_search(vault, "machine learning research", limit=5)
for hit in hits:
print(hit["id"], f"{hit['score']:.3f}")
Low-Level Provider Access
To debug provider responses or embed arbitrary text without the vault logic, use the internal _http_embed function:
from hyperresearch.core.embed import _http_embed
vectors = _http_embed(
provider="openai",
model="text-embedding-3-small",
texts=["What is vector search?"]
)
This function requires the appropriate environment variable (OPENAI_API_KEY or VOYAGE_API_KEY) and returns raw vector lists.
Summary
- Synchronization compares content hashes to detect notes needing re-embedding.
- Configuration happens in
~/.hyperresearch/config.tomlviaEmbeddingSettings. - Core logic resides in
src/hyperresearch/core/embed.py, specifically theembed_syncfunction. - Storage uses SQLite BLOBs in the
embeddingstable with model-specific stamps. - Interfaces include CLI (
hyperresearch embed sync), Python API (embed_sync), and low-level_http_embed.
Frequently Asked Questions
What triggers a note to be re-embedded?
A note is flagged as dirty when its computed stamp—formatted as f"{model}@{content_hash}"—differs from the embedded_model value stored in the embeddings table. This occurs when the note content changes or when you switch to a different embedding model.
Which embedding providers does Hyperresearch support?
Hyperresearch supports Voyage AI and OpenAI. Configure your choice via embeddings.provider in config.toml. If set to "none", the system raises an EmbeddingError and refuses to sync.
How are embeddings stored in the database?
Vectors are stored as binary BLOBs in the embeddings table defined in src/hyperresearch/core/db.py. Each row contains note_id, a model identifier string (including hash), dimensions, the packed vector data, and a created_at timestamp.
Can I adjust how much note content is embedded?
Yes. The body_chars setting in EmbeddingSettings controls how many characters from the note body are included alongside the title and summary. Reducing this value speeds up synchronization and reduces API costs for long documents.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →