# How to Use Turbovec with the Haystack DocumentStore: Setup, Retrieval, and Persistence

> Learn to use Turbovec with Haystack DocumentStore for fast embedding retrieval and storage. Explore setup, retrieval, and persistence with TurboQuantDocumentStore.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: how-to-guide
- Published: 2026-07-27

---

**Turbovec provides a drop-in Haystack 2.x `DocumentStore` implementation called `TurboQuantDocumentStore` that wraps a Rust-backed quantized vector index for high-performance embedding retrieval and storage.**

If you are building retrieval pipelines with **Haystack 2.x** and need a high-performance, memory-efficient document store, the `turbovec` project by **RyanCodrai** offers a purpose-built solution. The **`TurboQuantDocumentStore`** class in the `turbovec.haystack` module replaces generic in-memory backends with a quantized vector index that preserves search speed while dramatically reducing memory use. In this guide, you will learn how to integrate `turbovec` with the Haystack DocumentStore, run filtered vector searches, and persist your index between sessions.

## Quick Start: Initializing TurboQuantDocumentStore

The `TurboQuantDocumentStore` is imported directly from `turbovec.haystack`. You instantiate it with the target vector dimension and quantization bit width, then write standard Haystack `Document` objects.

```python
from turbovec.haystack import TurboQuantDocumentStore
from haystack import Document

# Create a store for 1536-dim vectors with 4-bit quantization

store = TurboQuantDocumentStore(dim=1536, bit_width=4)

# Write a batch of documents (embeddings must be pre-computed)

docs = [
    Document(id="doc-1", content="Hello world", embedding=[0.1]*1536, meta={"source": "demo"}),
    Document(id="doc-2", content="Another text", embedding=[0.2]*1536, meta={"source": "demo"}),
]
store.write_documents(docs)

# Simple similarity search (top-3)

hits = store.embedding_retrieval(query_embedding=[0.15]*1536, top_k=3)
for hit in hits:
    print(hit.id, hit.score, hit.meta)

```

As implemented in `RyanCodrai/turbovec`, the store immediately quantizes incoming vectors, so you should pre-compute embeddings using your preferred Haystack embedder before writing.

## How TurboQuantDocumentStore Works Internally

Understanding the internal layout helps explain why `turbovec` is fast and memory-efficient.

### Rust-Backed Quantized Index

In [`turbovec-python/python/turbovec/haystack.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/haystack.py), the `TurboQuantDocumentStore` constructor initializes a quantized **`IdMapIndex`** exposed by [`turbovec-python/python/turbovec/_turbovec.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/_turbovec.py). This Rust-backed index compresses vectors into **2–4 bit** representations. Because the index discards full-precision embeddings after quantization, all retrieval operations work directly on the compressed data, which slashes RAM usage compared to a standard `InMemoryDocumentStore`.

### Document ID Mapping and Metadata Storage

To bridge Haystack’s string-based IDs with the index’s numeric handles, the store maintains two private maps inside [`haystack.py`](https://github.com/RyanCodrai/turbovec/blob/main/haystack.py):

- **`_str_to_u64`** — maps Haystack document IDs to internal `u64` handles.
- **`_u64_to_doc`** — holds the persisted document fields, including `content`, `meta`, `blob`, and sparse embeddings.

These maps live in ordinary Python dictionaries, which makes metadata scanning and filtering fast and predictable.

## Writing Documents with write_documents

When you call `write_documents`, the method validates the input list, assigns a fresh internal handle via `_issue_handle`, and forwards the embedding batch to the underlying index through **`IdMapIndex.add_with_ids`**. According to the source in [`turbovec-python/python/turbovec/haystack.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/haystack.py) (lines 79–82 and 112–124), this batching step keeps Python-to-Rust overhead minimal.

Because the quantized index does not retain full-precision vectors, the `return_embedding` flag exists only for API parity with the Haystack protocol; fetched documents always have `embedding=None` (lines 55–64 and 63–65).

## Embedding Retrieval and Search Options

Similarity search is handled by **`embedding_retrieval`**. This method normalizes the query vector when the store is configured for **cosine** similarity, then executes a kernel search against the quantized `IdMapIndex` (lines 71–84).

### Metadata Filtering During Retrieval

If you supply a `filters` argument, `embedding_retrieval` builds an allow-list of valid handle IDs by scanning the in-memory `_u64_to_doc` table. It then passes that allow-list into the index search, preventing unnecessary distance computations against excluded documents (lines 101–112).

### Scaling Similarity Scores

Set **`scale_score=True`** to rescale raw similarities into a comparable range. The implementation in [`haystack.py`](https://github.com/RyanCodrai/turbovec/blob/main/haystack.py) (lines 99–108) mirrors Haystack’s `InMemoryDocumentStore` logic:

- **Cosine** scores are linearly mapped to `[0, 1]`.
- **Dot-product** scores are passed through a sigmoid.

The example below combines filtering and score scaling:

```python
filtered_hits = store.embedding_retrieval(
    query_embedding=[0.15]*1536,
    top_k=5,
    filters={"field": "meta.source", "operator": "==", "value": "demo"},
    scale_score=True,
)

# scores are now in [0, 1] for cosine mode

for hit in filtered_hits:
    print(hit.id, hit.score)

```

## Asynchronous Usage

All async APIs—such as **`write_documents_async`** and **`embedding_retrieval_async`**—are thin wrappers that execute their synchronous counterparts inside a one-worker `ThreadPoolExecutor`. This pattern mirrors Haystack’s `to_thread` approach (lines 81–95 in [`haystack.py`](https://github.com/RyanCodrai/turbovec/blob/main/haystack.py)) and is useful when you want non-blocking storage calls inside async web handlers.

```python
import asyncio

async def async_demo():
    store = TurboQuantDocumentStore(dim=1536, bit_width=4)
    await store.write_documents_async(docs)
    results = await store.embedding_retrieval_async(
        query_embedding=[0.15]*1536, top_k=2
    )
    for doc in results:
        print("async:", doc.id)

asyncio.run(async_demo())

```

## Persisting and Reloading the DocumentStore

Long-running pipelines need durable storage. The `turbovec` Haystack integration handles persistence through **`save_to_disk`** and **`load_from_disk`**. Calling `save_to_disk` writes the quantized index to `index.tvim` and a human-readable JSON side-car named [`docstore.json`](https://github.com/RyanCodrai/turbovec/blob/main/docstore.json) that stores metadata, handle mappings, and store configuration (lines 108–123). When reloading, `load_from_disk` reconstructs the store from these files (lines 155–165). Compatibility shims in the codebase also support loading older schema versions (v1–v2) that lack `blob` or sparse-embedding fields.

```python
import pathlib

path = pathlib.Path("./my_store")
store.save_to_disk(path)  # writes index.tvim + docstore.json

restored = TurboQuantDocumentStore.load_from_disk(path)
print("restored count:", restored.count_documents())

```

## Thread Safety and Concurrent Access

The store is safe to share across threads. Concurrent reads operate without blocking, while all mutations—writes, updates, and deletes—are serialized behind a **re-entrant lock** (`_write_lock`) defined in [`turbovec-python/python/turbovec/haystack.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/haystack.py) (lines 72–79). This design lets you serve multiple retrieval threads from the same quantized index while avoiding race conditions during ingest.

## Summary

- **`TurboQuantDocumentStore`** is the drop-in `turbovec` Haystack DocumentStore implementation, living in `turbovec.haystack`.
- It wraps a **Rust-backed `IdMapIndex`** that stores vectors in 2–4 bit quantized form.
- **Writing** batches documents via `write_documents` and assigns internal handles with `_issue_handle` before calling `IdMapIndex.add_with_ids`.
- **Retrieval** uses `embedding_retrieval`, supports cosine normalization, metadata filtering by scanning `_u64_to_doc`, and optional score scaling.
- **Async** variants run sync logic in a `ThreadPoolExecutor` for API compatibility.
- **Persistence** produces `index.tvim` and [`docstore.json`](https://github.com/RyanCodrai/turbovec/blob/main/docstore.json), with backward-compatible loaders for older schemas.
- A **re-entrant write lock** keeps mutation safe without blocking concurrent reads.

## Frequently Asked Questions

### What versions of Haystack does TurboQuantDocumentStore support?

`TurboQuantDocumentStore` targets **Haystack 2.x** and implements the `DocumentStore` protocol defined by that major version. It is tested for parity with Haystack’s reference behaviors, including score scaling and filtering semantics.

### Why do retrieved documents always have embedding=None?

Because the underlying `IdMapIndex` keeps vectors only in **2–4 bit quantized** form, the original full-precision embeddings are discarded after ingest. The store therefore returns `embedding=None` on retrieval regardless of the `return_embedding` flag, which exists solely for API compatibility.

### Can I use metadata filters with TurboQuantDocumentStore?

Yes. When you pass a `filters` dictionary to `embedding_retrieval`, the method scans the internal `_u64_to_doc` map to build an allow-list of document handles. It then restricts the quantized index search to that subset, giving you fast filtered vector search.

### Is TurboQuantDocumentStore thread-safe for concurrent workloads?

Yes. The store allows concurrent read operations across threads. All write operations are protected by a re-entrant `_write_lock`, so you can safely ingest documents in one thread while serving queries from many others.