How to Integrate Turbovec with LlamaIndex for Vector Storage

To integrate Turbovec with LlamaIndex for vector storage, import TurboQuantVectorStore from turbovec.llama_index, optionally configure its bit-width and similarity mode, and pass the instance to StorageContext.from_defaults so that VectorStoreIndex automatically uses the quantized backend.

Integrating Turbovec with LlamaIndex for vector storage replaces the default in-memory store with a high-performance, quantized index. The RyanCodrai/turbovec repository exposes this capability through a single compatibility class that mirrors the API of LlamaIndex's native SimpleVectorStore.

How the Turbovec-LlamaIndex Integration Works

TurboQuantVectorStore Class

The bridge between the two libraries is the TurboQuantVectorStore class, defined beginning at line 62 in turbovec-python/python/turbovec/llama_index.py. This class implements LlamaIndex's BasePydanticVectorStore protocol, which means it exposes the same public methods as LlamaIndex's built-in SimpleVectorStore.

Under the hood, each instance wraps a Turbovec IdMapIndex (constructed lazily at line 31 in the same file) that compresses vectors to 2–4 bits per dimension. The class also maintains a side-car JSON file containing node text and metadata, so the full LlamaIndex node model remains intact.

Similarity Modes and Thread Safety

Turbovec supports two immutable similarity modes. "cosine" is the default; it L2-normalizes embeddings at insertion and query time. "dot_product" keeps raw vectors and returns inner-product scores. The chosen mode is set at initialization and persisted with the side-car.

Thread safety is built in by design. Read operations such as query and get_nodes run lock-free, while all mutations—including add, delete, clear, and persist—are serialized behind a per-store re-entrant lock. This architecture allows concurrent readers to scale across threads without blocking.

Basic Integration Example

Getting started requires no changes to existing LlamaIndex logic beyond the initial import and storage setup.

from llama_index.core import VectorStoreIndex, StorageContext
from turbovec.llama_index import TurboQuantVectorStore

# Create the Turbovec-backed store (lazy construction)

vector_store = TurboQuantVectorStore()

# Build a storage context that LlamaIndex will use

storage_context = StorageContext.from_defaults(vector_store=vector_store)

# Index a collection of documents (any LlamaIndex Document objects)

index = VectorStoreIndex.from_documents(documents, storage_context=storage_context)

# Retrieve the top-5 most similar chunks for a query

retriever = index.as_retriever(similarity_top_k=5)

This pattern is documented in docs/integrations/llama_index.md and confirms that existing calls like VectorStoreIndex.from_documents and index.as_retriever work unchanged.

Configuring Quantization and Similarity

You can control compression and scoring behavior before nodes are added. The from_params factory accepts a bit_width argument (typically 2–4) and a similarity string.


# Explicitly choose 3-bit quantization and dot-product scoring

store = TurboQuantVectorStore.from_params(bit_width=3, similarity="dot_product")

Because the similarity mode is immutable for the lifetime of the store, it must be chosen at construction. These parameters are later restored when reloading from disk.

Persisting and Reloading Data

The store writes two files on demand: a binary {stem}.tvim file for the quantized IdMapIndex and a {stem}.nodes.json side-car for text and metadata. Use from_persist_dir or from_persist_path to restore a previous session.


# Persist the store to a directory (creates *.tvim and *.nodes.json)

storage_context.persist(persist_dir="./my_store")

# Later, load it back

vector_store = TurboQuantVectorStore.from_persist_dir(persist_dir="./my_store")
storage_context = StorageContext.from_defaults(
    vector_store=vector_store,
    persist_dir="./my_store"
)

As noted in docs/integrations/llama_index.md, the similarity mode and all node metadata are recovered automatically during reload.

Querying with Metadata Filters

TurboQuantVectorStore supports LlamaIndex's standard filter semantics. You can restrict results by metadata fields and optional node ID lists through a VectorStoreQuery.

from llama_index.core.vector_stores.types import (
    MetadataFilter, MetadataFilters, FilterCondition, VectorStoreQuery,
)

filters = MetadataFilters(
    filters=[
        MetadataFilter(key="category", value="finance", operator=FilterOperator.EQ),
        MetadataFilter(key="year", value=2023, operator=FilterOperator.GTE),
    ],
    condition=FilterCondition.AND,
)

result = vector_store.query(
    VectorStoreQuery(
        query_embedding=my_embedding,
        similarity_top_k=5,
        filters=filters,
        node_ids=["chunk-1", "chunk-2"],   # optional restriction

    )
)

Using the Async API

Every public method has an async counterpart, enabling seamless use in LlamaIndex's async pipelines. The available async methods include async_add, aquery, aget_nodes, and aclear.

await vector_store.async_add(nodes)               # add nodes asynchronously

result = await vector_store.aquery(query)         # async query

await vector_store.aclear()                       # async clear

Summary

  • TurboQuantVectorStore in turbovec-python/python/turbovec/llama_index.py implements the LlamaIndex BasePydanticVectorStore protocol, making it a drop-in replacement for SimpleVectorStore.
  • To integrate, pass a TurboQuantVectorStore instance to StorageContext.from_defaults and proceed with standard VectorStoreIndex workflows.
  • Quantization is handled by an internal IdMapIndex; configure it via from_params(bit_width=..., similarity=...).
  • The store creates a binary .tvim index and a .nodes.json side-car on persist, both restorable through from_persist_dir.
  • Reads are lock-free and mutations are serialized by a re-entrant lock, ensuring safe concurrent access.
  • Full async coverage—including aquery and async_add—is provided for non-blocking LlamaIndex pipelines.

Frequently Asked Questions

What class connects Turbovec to LlamaIndex?

The TurboQuantVectorStore class, defined starting at line 62 in turbovec-python/python/turbovec/llama_index.py, serves as the integration layer. It subclasses LlamaIndex's BasePydanticVectorStore and delegates vector storage and search to a Turbovec IdMapIndex.

How is vector quantization configured in the LlamaIndex store?

Quantization is controlled through the bit_width parameter in the from_params factory method, supporting 2–4 bits per dimension. The internal IdMapIndex is built lazily if no existing index is supplied, as seen in the constructor logic around line 31 of turbovec-python/python/turbovec/llama_index.py.

Is the Turbovec LlamaIndex store thread-safe?

Yes. According to the RyanCodrai/turbovec source documentation, read paths such as query and get_nodes execute lock-free, while write paths—including add, delete, clear, and persist—are protected by a per-store re-entrant lock. This design guarantees consistent views for concurrent readers.

Which files are generated when persisting the store?

The persist method outputs two files: a binary {stem}.tvim file containing the quantized vector index and a {stem}.nodes.json side-car holding node text and metadata. These can be reloaded with from_persist_path or from_persist_dir, restoring both the similarity mode and all node data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →