# How to Integrate Turbovec with LlamaIndex for Vector Storage

> Integrate Turbovec with LlamaIndex for efficient vector storage. Configure quantization and similarity modes using TurboQuantVectorStore for optimized performance.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: how-to-guide
- Published: 2026-07-27

---

**To integrate Turbovec with LlamaIndex for vector storage, import `TurboQuantVectorStore` from `turbovec.llama_index`, optionally configure its bit-width and similarity mode, and pass the instance to `StorageContext.from_defaults` so that `VectorStoreIndex` automatically uses the quantized backend.**

Integrating Turbovec with LlamaIndex for vector storage replaces the default in-memory store with a high-performance, quantized index. The `RyanCodrai/turbovec` repository exposes this capability through a single compatibility class that mirrors the API of LlamaIndex's native `SimpleVectorStore`.

## How the Turbovec-LlamaIndex Integration Works

### TurboQuantVectorStore Class

The bridge between the two libraries is the **`TurboQuantVectorStore`** class, defined beginning at line 62 in [`turbovec-python/python/turbovec/llama_index.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/llama_index.py). This class implements LlamaIndex's `BasePydanticVectorStore` protocol, which means it exposes the same public methods as LlamaIndex's built-in `SimpleVectorStore`.

Under the hood, each instance wraps a Turbovec **`IdMapIndex`** (constructed lazily at line 31 in the same file) that compresses vectors to 2–4 bits per dimension. The class also maintains a side-car JSON file containing node text and metadata, so the full LlamaIndex node model remains intact.

### Similarity Modes and Thread Safety

Turbovec supports two immutable similarity modes. **`"cosine"`** is the default; it L2-normalizes embeddings at insertion and query time. **`"dot_product"`** keeps raw vectors and returns inner-product scores. The chosen mode is set at initialization and persisted with the side-car.

Thread safety is built in by design. Read operations such as `query` and `get_nodes` run **lock-free**, while all mutations—including `add`, `delete`, `clear`, and `persist`—are serialized behind a per-store re-entrant lock. This architecture allows concurrent readers to scale across threads without blocking.

## Basic Integration Example

Getting started requires no changes to existing LlamaIndex logic beyond the initial import and storage setup.

```python
from llama_index.core import VectorStoreIndex, StorageContext
from turbovec.llama_index import TurboQuantVectorStore

# Create the Turbovec-backed store (lazy construction)

vector_store = TurboQuantVectorStore()

# Build a storage context that LlamaIndex will use

storage_context = StorageContext.from_defaults(vector_store=vector_store)

# Index a collection of documents (any LlamaIndex Document objects)

index = VectorStoreIndex.from_documents(documents, storage_context=storage_context)

# Retrieve the top-5 most similar chunks for a query

retriever = index.as_retriever(similarity_top_k=5)

```

This pattern is documented in [`docs/integrations/llama_index.md`](https://github.com/RyanCodrai/turbovec/blob/main/docs/integrations/llama_index.md) and confirms that existing calls like `VectorStoreIndex.from_documents` and `index.as_retriever` work unchanged.

## Configuring Quantization and Similarity

You can control compression and scoring behavior before nodes are added. The **`from_params`** factory accepts a `bit_width` argument (typically 2–4) and a `similarity` string.

```python

# Explicitly choose 3-bit quantization and dot-product scoring

store = TurboQuantVectorStore.from_params(bit_width=3, similarity="dot_product")

```

Because the similarity mode is immutable for the lifetime of the store, it must be chosen at construction. These parameters are later restored when reloading from disk.

## Persisting and Reloading Data

The store writes two files on demand: a binary **`{stem}.tvim`** file for the quantized `IdMapIndex` and a **`{stem}.nodes.json`** side-car for text and metadata. Use `from_persist_dir` or `from_persist_path` to restore a previous session.

```python

# Persist the store to a directory (creates *.tvim and *.nodes.json)

storage_context.persist(persist_dir="./my_store")

# Later, load it back

vector_store = TurboQuantVectorStore.from_persist_dir(persist_dir="./my_store")
storage_context = StorageContext.from_defaults(
    vector_store=vector_store,
    persist_dir="./my_store"
)

```

As noted in [`docs/integrations/llama_index.md`](https://github.com/RyanCodrai/turbovec/blob/main/docs/integrations/llama_index.md), the similarity mode and all node metadata are recovered automatically during reload.

## Querying with Metadata Filters

`TurboQuantVectorStore` supports LlamaIndex's standard filter semantics. You can restrict results by metadata fields and optional node ID lists through a `VectorStoreQuery`.

```python
from llama_index.core.vector_stores.types import (
    MetadataFilter, MetadataFilters, FilterCondition, VectorStoreQuery,
)

filters = MetadataFilters(
    filters=[
        MetadataFilter(key="category", value="finance", operator=FilterOperator.EQ),
        MetadataFilter(key="year", value=2023, operator=FilterOperator.GTE),
    ],
    condition=FilterCondition.AND,
)

result = vector_store.query(
    VectorStoreQuery(
        query_embedding=my_embedding,
        similarity_top_k=5,
        filters=filters,
        node_ids=["chunk-1", "chunk-2"],   # optional restriction

    )
)

```

## Using the Async API

Every public method has an async counterpart, enabling seamless use in LlamaIndex's async pipelines. The available async methods include **`async_add`**, **`aquery`**, **`aget_nodes`**, and **`aclear`**.

```python
await vector_store.async_add(nodes)               # add nodes asynchronously

result = await vector_store.aquery(query)         # async query

await vector_store.aclear()                       # async clear

```

## Summary

- **`TurboQuantVectorStore`** in [`turbovec-python/python/turbovec/llama_index.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/llama_index.py) implements the LlamaIndex `BasePydanticVectorStore` protocol, making it a drop-in replacement for `SimpleVectorStore`.
- To integrate, pass a `TurboQuantVectorStore` instance to **`StorageContext.from_defaults`** and proceed with standard `VectorStoreIndex` workflows.
- Quantization is handled by an internal **`IdMapIndex`**; configure it via `from_params(bit_width=..., similarity=...)`.
- The store creates a binary `.tvim` index and a [`.nodes.json`](https://github.com/RyanCodrai/turbovec/blob/main/.nodes.json) side-car on **`persist`**, both restorable through `from_persist_dir`.
- Reads are lock-free and mutations are serialized by a re-entrant lock, ensuring safe concurrent access.
- Full async coverage—including **`aquery`** and **`async_add`**—is provided for non-blocking LlamaIndex pipelines.

## Frequently Asked Questions

### What class connects Turbovec to LlamaIndex?

The **`TurboQuantVectorStore`** class, defined starting at line 62 in [`turbovec-python/python/turbovec/llama_index.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/llama_index.py), serves as the integration layer. It subclasses LlamaIndex's `BasePydanticVectorStore` and delegates vector storage and search to a Turbovec `IdMapIndex`.

### How is vector quantization configured in the LlamaIndex store?

Quantization is controlled through the `bit_width` parameter in the **`from_params`** factory method, supporting 2–4 bits per dimension. The internal `IdMapIndex` is built lazily if no existing index is supplied, as seen in the constructor logic around line 31 of [`turbovec-python/python/turbovec/llama_index.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/llama_index.py).

### Is the Turbovec LlamaIndex store thread-safe?

Yes. According to the `RyanCodrai/turbovec` source documentation, read paths such as `query` and `get_nodes` execute lock-free, while write paths—including `add`, `delete`, `clear`, and `persist`—are protected by a per-store re-entrant lock. This design guarantees consistent views for concurrent readers.

### Which files are generated when persisting the store?

The **`persist`** method outputs two files: a binary `{stem}.tvim` file containing the quantized vector index and a `{stem}.nodes.json` side-car holding node text and metadata. These can be reloaded with **`from_persist_path`** or **`from_persist_dir`**, restoring both the similarity mode and all node data.