# Integrating turbovec with LangChain, LlamaIndex, Haystack, or Agno: A Complete Guide

> Integrate Turbovec with LangChain LlamaIndex Haystack or Agno and quantize embeddings to 2-4 bits. This guide offers full API compatibility for efficient vector search.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: how-to-guide
- Published: 2026-06-16

---

**Turbovec provides drop-in vector store adapters for LangChain, LlamaIndex, Haystack, and Agno that quantize embeddings to 2–4 bits while maintaining full API compatibility with each framework.**

The `RyanCodrai/turbovec` repository ships Python wrappers that wrap the Rust-powered `IdMapIndex` core, letting you substitute turbovec for standard in-memory vector stores without rewriting your retrieval pipelines. Each adapter lives in its own module under `turbovec-python/python/turbovec/` and implements the host framework's expected interface while handling quantization, ID mapping, and persistence automatically.

## Core Architecture of the Turbovec Adapters

All four integrations share a common design built on `turbovec._turbovec.IdMapIndex`. The wrappers manage the transition between high-level framework objects (Documents, Nodes, Texts) and the compressed binary representation stored in Rust.

### Lazy Index Creation and Dimensionality Inference

Each adapter initializes an `IdMapIndex` without specifying dimensionality (`dim=None`). The first batch of embeddings locks the vector size, matching the "no-arg" constructor pattern used by LangChain and LlamaIndex native stores.

In [`turbovec/langchain.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/langchain.py) (line 48), [`turbovec/llama_index.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/llama_index.py) (line 100), and [`turbovec/haystack.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/haystack.py) (line 65), the constructors defer index creation until the first `add` operation, automatically inferring dimensions from the incoming NumPy arrays.

### Quantization and Memory Efficiency

Vectors are quantized to 2–4 bits per dimension before reaching the Rust index. The wrapper calls `IdMapIndex.add_with_ids` internally, then discards the full-precision floats, keeping only a side-car mapping of `id → (text, metadata)` in Python dictionaries.

This design choice means **turbovec uses significantly less RAM** than float32 stores, but operations requiring original embeddings—such as MMR reranking—are deliberately unsupported and raise `NotImplementedError`.

### ID Mapping and Handle Management

Turbovec assigns each document a fresh `u64` handle via `_issue_handle`. The adapters maintain bidirectional dictionaries (`_str_to_u64` and `_u64_to_str`) to map between framework string IDs and internal numeric handles. These mappings are persisted alongside the binary index so handles remain consistent across restarts.

In [`turbovec/langchain.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/langchain.py) and [`turbovec/llama_index.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/llama_index.py), you'll find these maps initialized in `__init__` and validated during `load` operations via [`turbovec/_persist.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/_persist.py).

### Persistence and Schema Validation

Each adapter writes two files: a binary `*.tvim` containing the compressed index, and a JSON side-car storing the text, metadata, and handle mappings. Loading validates the schema version and reconstructs the bidirectional maps using `check_persisted_handles` from [`turbovec/_persist.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/_persist.py).

- **LangChain**: Uses `dump` and `load` (lines 86–115)
- **LlamaIndex**: Uses `persist` and `from_persist_path` (lines 78–106)
- **Haystack**: Uses `save_to_disk` and `load_from_disk` (lines 95–110)

## Framework-Specific Integration Guides

### LangChain Integration (TurboQuantVectorStore)

The `TurboQuantVectorStore` class in [`turbovec/langchain.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/langchain.py) implements the standard LangChain vector store interface. It accepts an `Embeddings` instance and handles quantization transparently.

```python
from turbovec.langchain import TurboQuantVectorStore
from langchain_core.embeddings import Embeddings
import numpy as np

class DummyEmbeddings(Embeddings):
    def embed_documents(self, texts):
        return np.random.randn(len(texts), 1536).tolist()
    
    def embed_query(self, text):
        return np.random.randn(1536).tolist()

emb = DummyEmbeddings()
store = TurboQuantVectorStore(embedding=emb, bit_width=4)

# Add texts with metadata

store.add_texts(
    ["hello world", "turbovec is fast"],
    metadatas=[{"source": "demo"}] * 2
)

# Search

results = store.similarity_search("fast vector store", k=2)

```

The search logic resides in `_search_vector` (line 302), which builds an allow-list of handles for filtered queries before calling the Rust kernel.

### LlamaIndex Integration (TurboQuantVectorStore)

LlamaIndex users import from [`turbovec/llama_index.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/llama_index.py). The wrapper exposes `add` for nodes and `query` for retrieval, matching the `VectorStore` protocol.

```python
from turbovec.llama_index import TurboQuantVectorStore
from llama_index.core import VectorStoreIndex
from llama_index.core.schema import TextNode

vector_store = TurboQuantVectorStore(bit_width=4)
node = TextNode(text="turbovec integrates easily")
vector_store.add([node])

index = VectorStoreIndex.from_vector_store(vector_store)
response = index.as_query_engine().query("integration")

```

Key methods include `add` (line 38) and `query` (line 640), with duplicate resolution handled by `resolve_duplicates` from [`turbovec/_dedup.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/_dedup.py) (lines 42–58).

### Haystack Integration (TurboQuantDocumentStore)

For Haystack 2.x pipelines, use `TurboQuantDocumentStore` from [`turbovec/haystack.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/haystack.py). It implements `write_documents` and `embedding_retrieval` with batch-level duplicate handling.

```python
from turbovec.haystack import TurboQuantDocumentStore
from haystack import Document
import numpy as np

store = TurboQuantDocumentStore(bit_width=4)

docs = [
    Document(
        content="first doc",
        embedding=np.random.randn(1536).tolist(),
        meta={"type": "demo"}
    ),
    Document(
        content="second doc",
        embedding=np.random.randn(1536).tolist(),
        meta={"type": "demo"}
    ),
]

store.write_documents(docs)
results = store.embedding_retrieval(
    np.random.randn(1536).tolist(),
    top_k=2
)

```

The `embedding_retrieval` method (line 542) implements the same allow-list filtering pattern as the other adapters, scanning metadata first to reduce unnecessary distance calculations.

### Agno Integration (AgnoQuantVectorStore)

The `AgnoQuantVectorStore` in [`turbovec/agno.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/agno.py) provides a minimal, framework-agnostic interface for custom pipelines or testing. It has no external dependencies beyond NumPy.

```python
from turbovec.agno import AgnoQuantVectorStore
import numpy as np

store = AgnoQuantVectorStore(dim=1536, bit_width=4)

vectors = np.random.randn(5, 1536).astype(np.float32)
store.add_vectors(
    vectors,
    ids=[f"id_{i}" for i in range(5)]
)

query = np.random.randn(1, 1536).astype(np.float32)
scores, handles = store.search(query, k=3)

```

This class exposes `add_vectors` (line 58) and `search` (line 97) directly, bypassing the text-wrapping logic required by the other frameworks.

## Duplicate Handling and Data Safety

When adding documents with existing IDs, turbovec implements a two-phase commit to prevent data loss. The wrapper first issues new handles and adds the vectors, then removes the old handles only after successful insertion.

LangChain and LlamaIndex use the `resolve_duplicates` utility from [`turbovec/_dedup.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/_dedup.py), supporting policies like `keep_last`, `fail`, and `overwrite`. Haystack implements its own batch-level duplicate checking in `write_documents` (lines 80–118).

## Search Implementation and Filtering

All adapters follow the same search pattern in their query methods:

1. Convert the query embedding to a NumPy float32 vector
2. If filters are present, scan the side-car dictionary to build an allow-list of valid `u64` handles
3. Call `IdMapIndex.search` with the allow-list, skipping distance calculations for filtered-out vectors
4. Map returned handles back to framework objects using the bidirectional dictionaries

This filtered search appears in [`langchain.py`](https://github.com/RyanCodrai/turbovec/blob/main/langchain.py) (lines 302–332), [`llama_index.py`](https://github.com/RyanCodrai/turbovec/blob/main/llama_index.py) (lines 644–673), and [`haystack.py`](https://github.com/RyanCodrai/turbovec/blob/main/haystack.py) (lines 542–571).

## Limitations and Unsupported Features

Because turbovec discards full-precision vectors after quantization, certain retrieval augmentations are unavailable:

- **Maximal Marginal Relevance (MMR)**: Raises `NotImplementedError` in LangChain (lines 367–387) and LlamaIndex (lines 60–71)
- **Hybrid scoring**: Not implemented; only cosine similarity search is supported
- **Embedding retrieval**: Cannot return original float32 vectors, only quantized approximations

These limitations are inherent to the memory-efficient design and are clearly documented in the source code.

## Summary

- **Turbovec adapters** provide quantized, drop-in replacements for LangChain, LlamaIndex, Haystack, and Agno vector stores
- **Lazy initialization** allows dimensionality inference from the first batch of embeddings
- **2–4 bit quantization** happens in `IdMapIndex.add_with_ids`, dramatically reducing memory footprint
- **Bidirectional ID mapping** uses `u64` handles with persistent JSON side-cars for consistency across restarts
- **Filtered search** builds handle allow-lists in Python before calling the Rust kernel, optimizing query performance
- **Duplicate handling** follows a safe two-phase commit pattern to prevent partial-failure data corruption

## Frequently Asked Questions

### Does turbovec support Maximal Marginal Relevance (MMR) in LangChain?

No. Because turbovec quantizes vectors to 2–4 bits and does not retain full-precision embeddings, it cannot compute the diversity scores required for MMR. The LangChain adapter explicitly raises `NotImplementedError` with a clear message if you attempt to use MMR mode, as seen in [`turbovec/langchain.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/langchain.py) lines 367–387.

### How does turbovec handle duplicate document IDs?

The adapters resolve duplicates according to the host framework's policy before insertion. For LangChain and LlamaIndex, turbovec uses the `resolve_duplicates` utility from [`turbovec/_dedup.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/_dedup.py), supporting `fail`, `skip`, `overwrite`, and `keep_last` strategies. Haystack implements batch-level duplicate detection in `write_documents`. In all cases, turbovec uses a two-phase commit: new vectors are added before old vectors are removed, ensuring failed writes never corrupt existing data.

### Can I migrate existing data from FAISS or Chroma to turbovec?

Yes. You can extract vectors and metadata from your existing store, then add them to turbovec using the standard `add_texts` (LangChain), `add` (LlamaIndex), or `write_documents` (Haystack) methods. The turbovec wrapper will quantize the embeddings during insertion. Note that you cannot recover the original full-precision vectors after migration, so retain your source data if you need exact float32 values later.

### What bit width should I use for turbovec quantization?

The `bit_width` parameter accepts values of 2, 3, or 4. Use **4 bits** for highest recall accuracy with moderate memory savings, or **2 bits** for maximum compression when approximate results are acceptable. The choice depends on your embedding model and recall requirements; the adapter defaults to 4 bits if unspecified. You set this once at initialization in `TurboQuantVectorStore` or `AgnoQuantVectorStore` and cannot change it for an existing index.