How to Integrate Turbovec with LlamaIndex for Vector Storage
To integrate Turbovec with LlamaIndex for vector storage, import TurboQuantVectorStore from turbovec.llama_index, optionally configure its bit-width and similarity mode, and pass the instance to StorageContext.from_defaults so that VectorStoreIndex automatically uses the quantized backend.
Integrating Turbovec with LlamaIndex for vector storage replaces the default in-memory store with a high-performance, quantized index. The RyanCodrai/turbovec repository exposes this capability through a single compatibility class that mirrors the API of LlamaIndex's native SimpleVectorStore.
How the Turbovec-LlamaIndex Integration Works
TurboQuantVectorStore Class
The bridge between the two libraries is the TurboQuantVectorStore class, defined beginning at line 62 in turbovec-python/python/turbovec/llama_index.py. This class implements LlamaIndex's BasePydanticVectorStore protocol, which means it exposes the same public methods as LlamaIndex's built-in SimpleVectorStore.
Under the hood, each instance wraps a Turbovec IdMapIndex (constructed lazily at line 31 in the same file) that compresses vectors to 2–4 bits per dimension. The class also maintains a side-car JSON file containing node text and metadata, so the full LlamaIndex node model remains intact.
Similarity Modes and Thread Safety
Turbovec supports two immutable similarity modes. "cosine" is the default; it L2-normalizes embeddings at insertion and query time. "dot_product" keeps raw vectors and returns inner-product scores. The chosen mode is set at initialization and persisted with the side-car.
Thread safety is built in by design. Read operations such as query and get_nodes run lock-free, while all mutations—including add, delete, clear, and persist—are serialized behind a per-store re-entrant lock. This architecture allows concurrent readers to scale across threads without blocking.
Basic Integration Example
Getting started requires no changes to existing LlamaIndex logic beyond the initial import and storage setup.
from llama_index.core import VectorStoreIndex, StorageContext
from turbovec.llama_index import TurboQuantVectorStore
# Create the Turbovec-backed store (lazy construction)
vector_store = TurboQuantVectorStore()
# Build a storage context that LlamaIndex will use
storage_context = StorageContext.from_defaults(vector_store=vector_store)
# Index a collection of documents (any LlamaIndex Document objects)
index = VectorStoreIndex.from_documents(documents, storage_context=storage_context)
# Retrieve the top-5 most similar chunks for a query
retriever = index.as_retriever(similarity_top_k=5)
This pattern is documented in docs/integrations/llama_index.md and confirms that existing calls like VectorStoreIndex.from_documents and index.as_retriever work unchanged.
Configuring Quantization and Similarity
You can control compression and scoring behavior before nodes are added. The from_params factory accepts a bit_width argument (typically 2–4) and a similarity string.
# Explicitly choose 3-bit quantization and dot-product scoring
store = TurboQuantVectorStore.from_params(bit_width=3, similarity="dot_product")
Because the similarity mode is immutable for the lifetime of the store, it must be chosen at construction. These parameters are later restored when reloading from disk.
Persisting and Reloading Data
The store writes two files on demand: a binary {stem}.tvim file for the quantized IdMapIndex and a {stem}.nodes.json side-car for text and metadata. Use from_persist_dir or from_persist_path to restore a previous session.
# Persist the store to a directory (creates *.tvim and *.nodes.json)
storage_context.persist(persist_dir="./my_store")
# Later, load it back
vector_store = TurboQuantVectorStore.from_persist_dir(persist_dir="./my_store")
storage_context = StorageContext.from_defaults(
vector_store=vector_store,
persist_dir="./my_store"
)
As noted in docs/integrations/llama_index.md, the similarity mode and all node metadata are recovered automatically during reload.
Querying with Metadata Filters
TurboQuantVectorStore supports LlamaIndex's standard filter semantics. You can restrict results by metadata fields and optional node ID lists through a VectorStoreQuery.
from llama_index.core.vector_stores.types import (
MetadataFilter, MetadataFilters, FilterCondition, VectorStoreQuery,
)
filters = MetadataFilters(
filters=[
MetadataFilter(key="category", value="finance", operator=FilterOperator.EQ),
MetadataFilter(key="year", value=2023, operator=FilterOperator.GTE),
],
condition=FilterCondition.AND,
)
result = vector_store.query(
VectorStoreQuery(
query_embedding=my_embedding,
similarity_top_k=5,
filters=filters,
node_ids=["chunk-1", "chunk-2"], # optional restriction
)
)
Using the Async API
Every public method has an async counterpart, enabling seamless use in LlamaIndex's async pipelines. The available async methods include async_add, aquery, aget_nodes, and aclear.
await vector_store.async_add(nodes) # add nodes asynchronously
result = await vector_store.aquery(query) # async query
await vector_store.aclear() # async clear
Summary
TurboQuantVectorStoreinturbovec-python/python/turbovec/llama_index.pyimplements the LlamaIndexBasePydanticVectorStoreprotocol, making it a drop-in replacement forSimpleVectorStore.- To integrate, pass a
TurboQuantVectorStoreinstance toStorageContext.from_defaultsand proceed with standardVectorStoreIndexworkflows. - Quantization is handled by an internal
IdMapIndex; configure it viafrom_params(bit_width=..., similarity=...). - The store creates a binary
.tvimindex and a.nodes.jsonside-car onpersist, both restorable throughfrom_persist_dir. - Reads are lock-free and mutations are serialized by a re-entrant lock, ensuring safe concurrent access.
- Full async coverage—including
aqueryandasync_add—is provided for non-blocking LlamaIndex pipelines.
Frequently Asked Questions
What class connects Turbovec to LlamaIndex?
The TurboQuantVectorStore class, defined starting at line 62 in turbovec-python/python/turbovec/llama_index.py, serves as the integration layer. It subclasses LlamaIndex's BasePydanticVectorStore and delegates vector storage and search to a Turbovec IdMapIndex.
How is vector quantization configured in the LlamaIndex store?
Quantization is controlled through the bit_width parameter in the from_params factory method, supporting 2–4 bits per dimension. The internal IdMapIndex is built lazily if no existing index is supplied, as seen in the constructor logic around line 31 of turbovec-python/python/turbovec/llama_index.py.
Is the Turbovec LlamaIndex store thread-safe?
Yes. According to the RyanCodrai/turbovec source documentation, read paths such as query and get_nodes execute lock-free, while write paths—including add, delete, clear, and persist—are protected by a per-store re-entrant lock. This design guarantees consistent views for concurrent readers.
Which files are generated when persisting the store?
The persist method outputs two files: a binary {stem}.tvim file containing the quantized vector index and a {stem}.nodes.json side-car holding node text and metadata. These can be reloaded with from_persist_path or from_persist_dir, restoring both the similarity mode and all node data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →