Integrating turbovec with LangChain, LlamaIndex, Haystack, or Agno: A Complete Guide
Turbovec provides drop-in vector store adapters for LangChain, LlamaIndex, Haystack, and Agno that quantize embeddings to 2–4 bits while maintaining full API compatibility with each framework.
The RyanCodrai/turbovec repository ships Python wrappers that wrap the Rust-powered IdMapIndex core, letting you substitute turbovec for standard in-memory vector stores without rewriting your retrieval pipelines. Each adapter lives in its own module under turbovec-python/python/turbovec/ and implements the host framework's expected interface while handling quantization, ID mapping, and persistence automatically.
Core Architecture of the Turbovec Adapters
All four integrations share a common design built on turbovec._turbovec.IdMapIndex. The wrappers manage the transition between high-level framework objects (Documents, Nodes, Texts) and the compressed binary representation stored in Rust.
Lazy Index Creation and Dimensionality Inference
Each adapter initializes an IdMapIndex without specifying dimensionality (dim=None). The first batch of embeddings locks the vector size, matching the "no-arg" constructor pattern used by LangChain and LlamaIndex native stores.
In turbovec/langchain.py (line 48), turbovec/llama_index.py (line 100), and turbovec/haystack.py (line 65), the constructors defer index creation until the first add operation, automatically inferring dimensions from the incoming NumPy arrays.
Quantization and Memory Efficiency
Vectors are quantized to 2–4 bits per dimension before reaching the Rust index. The wrapper calls IdMapIndex.add_with_ids internally, then discards the full-precision floats, keeping only a side-car mapping of id → (text, metadata) in Python dictionaries.
This design choice means turbovec uses significantly less RAM than float32 stores, but operations requiring original embeddings—such as MMR reranking—are deliberately unsupported and raise NotImplementedError.
ID Mapping and Handle Management
Turbovec assigns each document a fresh u64 handle via _issue_handle. The adapters maintain bidirectional dictionaries (_str_to_u64 and _u64_to_str) to map between framework string IDs and internal numeric handles. These mappings are persisted alongside the binary index so handles remain consistent across restarts.
In turbovec/langchain.py and turbovec/llama_index.py, you'll find these maps initialized in __init__ and validated during load operations via turbovec/_persist.py.
Persistence and Schema Validation
Each adapter writes two files: a binary *.tvim containing the compressed index, and a JSON side-car storing the text, metadata, and handle mappings. Loading validates the schema version and reconstructs the bidirectional maps using check_persisted_handles from turbovec/_persist.py.
- LangChain: Uses
dumpandload(lines 86–115) - LlamaIndex: Uses
persistandfrom_persist_path(lines 78–106) - Haystack: Uses
save_to_diskandload_from_disk(lines 95–110)
Framework-Specific Integration Guides
LangChain Integration (TurboQuantVectorStore)
The TurboQuantVectorStore class in turbovec/langchain.py implements the standard LangChain vector store interface. It accepts an Embeddings instance and handles quantization transparently.
from turbovec.langchain import TurboQuantVectorStore
from langchain_core.embeddings import Embeddings
import numpy as np
class DummyEmbeddings(Embeddings):
def embed_documents(self, texts):
return np.random.randn(len(texts), 1536).tolist()
def embed_query(self, text):
return np.random.randn(1536).tolist()
emb = DummyEmbeddings()
store = TurboQuantVectorStore(embedding=emb, bit_width=4)
# Add texts with metadata
store.add_texts(
["hello world", "turbovec is fast"],
metadatas=[{"source": "demo"}] * 2
)
# Search
results = store.similarity_search("fast vector store", k=2)
The search logic resides in _search_vector (line 302), which builds an allow-list of handles for filtered queries before calling the Rust kernel.
LlamaIndex Integration (TurboQuantVectorStore)
LlamaIndex users import from turbovec/llama_index.py. The wrapper exposes add for nodes and query for retrieval, matching the VectorStore protocol.
from turbovec.llama_index import TurboQuantVectorStore
from llama_index.core import VectorStoreIndex
from llama_index.core.schema import TextNode
vector_store = TurboQuantVectorStore(bit_width=4)
node = TextNode(text="turbovec integrates easily")
vector_store.add([node])
index = VectorStoreIndex.from_vector_store(vector_store)
response = index.as_query_engine().query("integration")
Key methods include add (line 38) and query (line 640), with duplicate resolution handled by resolve_duplicates from turbovec/_dedup.py (lines 42–58).
Haystack Integration (TurboQuantDocumentStore)
For Haystack 2.x pipelines, use TurboQuantDocumentStore from turbovec/haystack.py. It implements write_documents and embedding_retrieval with batch-level duplicate handling.
from turbovec.haystack import TurboQuantDocumentStore
from haystack import Document
import numpy as np
store = TurboQuantDocumentStore(bit_width=4)
docs = [
Document(
content="first doc",
embedding=np.random.randn(1536).tolist(),
meta={"type": "demo"}
),
Document(
content="second doc",
embedding=np.random.randn(1536).tolist(),
meta={"type": "demo"}
),
]
store.write_documents(docs)
results = store.embedding_retrieval(
np.random.randn(1536).tolist(),
top_k=2
)
The embedding_retrieval method (line 542) implements the same allow-list filtering pattern as the other adapters, scanning metadata first to reduce unnecessary distance calculations.
Agno Integration (AgnoQuantVectorStore)
The AgnoQuantVectorStore in turbovec/agno.py provides a minimal, framework-agnostic interface for custom pipelines or testing. It has no external dependencies beyond NumPy.
from turbovec.agno import AgnoQuantVectorStore
import numpy as np
store = AgnoQuantVectorStore(dim=1536, bit_width=4)
vectors = np.random.randn(5, 1536).astype(np.float32)
store.add_vectors(
vectors,
ids=[f"id_{i}" for i in range(5)]
)
query = np.random.randn(1, 1536).astype(np.float32)
scores, handles = store.search(query, k=3)
This class exposes add_vectors (line 58) and search (line 97) directly, bypassing the text-wrapping logic required by the other frameworks.
Duplicate Handling and Data Safety
When adding documents with existing IDs, turbovec implements a two-phase commit to prevent data loss. The wrapper first issues new handles and adds the vectors, then removes the old handles only after successful insertion.
LangChain and LlamaIndex use the resolve_duplicates utility from turbovec/_dedup.py, supporting policies like keep_last, fail, and overwrite. Haystack implements its own batch-level duplicate checking in write_documents (lines 80–118).
Search Implementation and Filtering
All adapters follow the same search pattern in their query methods:
- Convert the query embedding to a NumPy float32 vector
- If filters are present, scan the side-car dictionary to build an allow-list of valid
u64handles - Call
IdMapIndex.searchwith the allow-list, skipping distance calculations for filtered-out vectors - Map returned handles back to framework objects using the bidirectional dictionaries
This filtered search appears in langchain.py (lines 302–332), llama_index.py (lines 644–673), and haystack.py (lines 542–571).
Limitations and Unsupported Features
Because turbovec discards full-precision vectors after quantization, certain retrieval augmentations are unavailable:
- Maximal Marginal Relevance (MMR): Raises
NotImplementedErrorin LangChain (lines 367–387) and LlamaIndex (lines 60–71) - Hybrid scoring: Not implemented; only cosine similarity search is supported
- Embedding retrieval: Cannot return original float32 vectors, only quantized approximations
These limitations are inherent to the memory-efficient design and are clearly documented in the source code.
Summary
- Turbovec adapters provide quantized, drop-in replacements for LangChain, LlamaIndex, Haystack, and Agno vector stores
- Lazy initialization allows dimensionality inference from the first batch of embeddings
- 2–4 bit quantization happens in
IdMapIndex.add_with_ids, dramatically reducing memory footprint - Bidirectional ID mapping uses
u64handles with persistent JSON side-cars for consistency across restarts - Filtered search builds handle allow-lists in Python before calling the Rust kernel, optimizing query performance
- Duplicate handling follows a safe two-phase commit pattern to prevent partial-failure data corruption
Frequently Asked Questions
Does turbovec support Maximal Marginal Relevance (MMR) in LangChain?
No. Because turbovec quantizes vectors to 2–4 bits and does not retain full-precision embeddings, it cannot compute the diversity scores required for MMR. The LangChain adapter explicitly raises NotImplementedError with a clear message if you attempt to use MMR mode, as seen in turbovec/langchain.py lines 367–387.
How does turbovec handle duplicate document IDs?
The adapters resolve duplicates according to the host framework's policy before insertion. For LangChain and LlamaIndex, turbovec uses the resolve_duplicates utility from turbovec/_dedup.py, supporting fail, skip, overwrite, and keep_last strategies. Haystack implements batch-level duplicate detection in write_documents. In all cases, turbovec uses a two-phase commit: new vectors are added before old vectors are removed, ensuring failed writes never corrupt existing data.
Can I migrate existing data from FAISS or Chroma to turbovec?
Yes. You can extract vectors and metadata from your existing store, then add them to turbovec using the standard add_texts (LangChain), add (LlamaIndex), or write_documents (Haystack) methods. The turbovec wrapper will quantize the embeddings during insertion. Note that you cannot recover the original full-precision vectors after migration, so retain your source data if you need exact float32 values later.
What bit width should I use for turbovec quantization?
The bit_width parameter accepts values of 2, 3, or 4. Use 4 bits for highest recall accuracy with moderate memory savings, or 2 bits for maximum compression when approximate results are acceptable. The choice depends on your embedding model and recall requirements; the adapter defaults to 4 bits if unspecified. You set this once at initialization in TurboQuantVectorStore or AgnoQuantVectorStore and cannot change it for an existing index.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →