How to Use Turbovec with the Haystack DocumentStore: Setup, Retrieval, and Persistence
Turbovec provides a drop-in Haystack 2.x DocumentStore implementation called TurboQuantDocumentStore that wraps a Rust-backed quantized vector index for high-performance embedding retrieval and storage.
If you are building retrieval pipelines with Haystack 2.x and need a high-performance, memory-efficient document store, the turbovec project by RyanCodrai offers a purpose-built solution. The TurboQuantDocumentStore class in the turbovec.haystack module replaces generic in-memory backends with a quantized vector index that preserves search speed while dramatically reducing memory use. In this guide, you will learn how to integrate turbovec with the Haystack DocumentStore, run filtered vector searches, and persist your index between sessions.
Quick Start: Initializing TurboQuantDocumentStore
The TurboQuantDocumentStore is imported directly from turbovec.haystack. You instantiate it with the target vector dimension and quantization bit width, then write standard Haystack Document objects.
from turbovec.haystack import TurboQuantDocumentStore
from haystack import Document
# Create a store for 1536-dim vectors with 4-bit quantization
store = TurboQuantDocumentStore(dim=1536, bit_width=4)
# Write a batch of documents (embeddings must be pre-computed)
docs = [
Document(id="doc-1", content="Hello world", embedding=[0.1]*1536, meta={"source": "demo"}),
Document(id="doc-2", content="Another text", embedding=[0.2]*1536, meta={"source": "demo"}),
]
store.write_documents(docs)
# Simple similarity search (top-3)
hits = store.embedding_retrieval(query_embedding=[0.15]*1536, top_k=3)
for hit in hits:
print(hit.id, hit.score, hit.meta)
As implemented in RyanCodrai/turbovec, the store immediately quantizes incoming vectors, so you should pre-compute embeddings using your preferred Haystack embedder before writing.
How TurboQuantDocumentStore Works Internally
Understanding the internal layout helps explain why turbovec is fast and memory-efficient.
Rust-Backed Quantized Index
In turbovec-python/python/turbovec/haystack.py, the TurboQuantDocumentStore constructor initializes a quantized IdMapIndex exposed by turbovec-python/python/turbovec/_turbovec.py. This Rust-backed index compresses vectors into 2–4 bit representations. Because the index discards full-precision embeddings after quantization, all retrieval operations work directly on the compressed data, which slashes RAM usage compared to a standard InMemoryDocumentStore.
Document ID Mapping and Metadata Storage
To bridge Haystack’s string-based IDs with the index’s numeric handles, the store maintains two private maps inside haystack.py:
_str_to_u64— maps Haystack document IDs to internalu64handles._u64_to_doc— holds the persisted document fields, includingcontent,meta,blob, and sparse embeddings.
These maps live in ordinary Python dictionaries, which makes metadata scanning and filtering fast and predictable.
Writing Documents with write_documents
When you call write_documents, the method validates the input list, assigns a fresh internal handle via _issue_handle, and forwards the embedding batch to the underlying index through IdMapIndex.add_with_ids. According to the source in turbovec-python/python/turbovec/haystack.py (lines 79–82 and 112–124), this batching step keeps Python-to-Rust overhead minimal.
Because the quantized index does not retain full-precision vectors, the return_embedding flag exists only for API parity with the Haystack protocol; fetched documents always have embedding=None (lines 55–64 and 63–65).
Embedding Retrieval and Search Options
Similarity search is handled by embedding_retrieval. This method normalizes the query vector when the store is configured for cosine similarity, then executes a kernel search against the quantized IdMapIndex (lines 71–84).
Metadata Filtering During Retrieval
If you supply a filters argument, embedding_retrieval builds an allow-list of valid handle IDs by scanning the in-memory _u64_to_doc table. It then passes that allow-list into the index search, preventing unnecessary distance computations against excluded documents (lines 101–112).
Scaling Similarity Scores
Set scale_score=True to rescale raw similarities into a comparable range. The implementation in haystack.py (lines 99–108) mirrors Haystack’s InMemoryDocumentStore logic:
- Cosine scores are linearly mapped to
[0, 1]. - Dot-product scores are passed through a sigmoid.
The example below combines filtering and score scaling:
filtered_hits = store.embedding_retrieval(
query_embedding=[0.15]*1536,
top_k=5,
filters={"field": "meta.source", "operator": "==", "value": "demo"},
scale_score=True,
)
# scores are now in [0, 1] for cosine mode
for hit in filtered_hits:
print(hit.id, hit.score)
Asynchronous Usage
All async APIs—such as write_documents_async and embedding_retrieval_async—are thin wrappers that execute their synchronous counterparts inside a one-worker ThreadPoolExecutor. This pattern mirrors Haystack’s to_thread approach (lines 81–95 in haystack.py) and is useful when you want non-blocking storage calls inside async web handlers.
import asyncio
async def async_demo():
store = TurboQuantDocumentStore(dim=1536, bit_width=4)
await store.write_documents_async(docs)
results = await store.embedding_retrieval_async(
query_embedding=[0.15]*1536, top_k=2
)
for doc in results:
print("async:", doc.id)
asyncio.run(async_demo())
Persisting and Reloading the DocumentStore
Long-running pipelines need durable storage. The turbovec Haystack integration handles persistence through save_to_disk and load_from_disk. Calling save_to_disk writes the quantized index to index.tvim and a human-readable JSON side-car named docstore.json that stores metadata, handle mappings, and store configuration (lines 108–123). When reloading, load_from_disk reconstructs the store from these files (lines 155–165). Compatibility shims in the codebase also support loading older schema versions (v1–v2) that lack blob or sparse-embedding fields.
import pathlib
path = pathlib.Path("./my_store")
store.save_to_disk(path) # writes index.tvim + docstore.json
restored = TurboQuantDocumentStore.load_from_disk(path)
print("restored count:", restored.count_documents())
Thread Safety and Concurrent Access
The store is safe to share across threads. Concurrent reads operate without blocking, while all mutations—writes, updates, and deletes—are serialized behind a re-entrant lock (_write_lock) defined in turbovec-python/python/turbovec/haystack.py (lines 72–79). This design lets you serve multiple retrieval threads from the same quantized index while avoiding race conditions during ingest.
Summary
TurboQuantDocumentStoreis the drop-inturbovecHaystack DocumentStore implementation, living inturbovec.haystack.- It wraps a Rust-backed
IdMapIndexthat stores vectors in 2–4 bit quantized form. - Writing batches documents via
write_documentsand assigns internal handles with_issue_handlebefore callingIdMapIndex.add_with_ids. - Retrieval uses
embedding_retrieval, supports cosine normalization, metadata filtering by scanning_u64_to_doc, and optional score scaling. - Async variants run sync logic in a
ThreadPoolExecutorfor API compatibility. - Persistence produces
index.tvimanddocstore.json, with backward-compatible loaders for older schemas. - A re-entrant write lock keeps mutation safe without blocking concurrent reads.
Frequently Asked Questions
What versions of Haystack does TurboQuantDocumentStore support?
TurboQuantDocumentStore targets Haystack 2.x and implements the DocumentStore protocol defined by that major version. It is tested for parity with Haystack’s reference behaviors, including score scaling and filtering semantics.
Why do retrieved documents always have embedding=None?
Because the underlying IdMapIndex keeps vectors only in 2–4 bit quantized form, the original full-precision embeddings are discarded after ingest. The store therefore returns embedding=None on retrieval regardless of the return_embedding flag, which exists solely for API compatibility.
Can I use metadata filters with TurboQuantDocumentStore?
Yes. When you pass a filters dictionary to embedding_retrieval, the method scans the internal _u64_to_doc map to build an allow-list of document handles. It then restricts the quantized index search to that subset, giving you fast filtered vector search.
Is TurboQuantDocumentStore thread-safe for concurrent workloads?
Yes. The store allows concurrent read operations across threads. All write operations are protected by a re-entrant _write_lock, so you can safely ingest documents in one thread while serving queries from many others.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →