# Migrating from FAISS to turbovec for Vector Search: A Complete Infrastructure Guide

> Migrate from FAISS to turbovec for vector search. Reduce RAM by 8x, eliminate rebuilds, and get direct SIMD filtered retrieval. Learn more in this infrastructure guide.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: migration-guide
- Published: 2026-07-18

---

**Migrating from FAISS to turbovec lets you replace offline-trained product quantization with an online, data-oblivious quantizer that cuts RAM usage by up to 8×, eliminates index rebuilds, and performs filtered top‑k retrieval directly inside SIMD kernels.**

If you are migrating from FAISS to turbovec for vector search infrastructure, you are moving from a batch-trained, memory-heavy `IndexPQ` pipeline to a Rust-based implementation of Google Research’s TurboQuant algorithm. The `turbovec` library—available in the RyanCodrai/turbovec repository—provides Python bindings and native Rust APIs that calibrate quantization parameters automatically during the first `add` call, allowing continuous streaming ingestion without a separate training phase.

## Why turbovec Replaces FAISS for Production Workloads

Compared with FAISS’s `IndexPQ` (and its FastScan variant), turbovec delivers the same or better recall with dramatically lower memory consumption. As documented in [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md), a 10‑million‑document corpus fits in roughly 4 GB with turbovec versus approximately 31 GB for FAISS, while query latency is typically faster on both ARM and x86 CPUs. Because quantization is **online**, you can add new vectors at any time without rebuilding the entire index—a common bottleneck in FAISS PQ pipelines.

## How turbovec Implements TurboQuant

The core algorithm described in [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md) consists of six stages that run automatically during ingestion and search.

### Normalization and Random Rotation

Each incoming vector is L2-normalized and multiplied by a single random orthogonal matrix. This transforms every coordinate into an identical Beta distribution— approximately normal with zero mean and variance `1/d`—regardless of the original data distribution. This step is the statistical foundation that makes TurboQuant data-oblivious.

### Per-Coordinate Calibration (TQ+)

On the first call to `add`, turbovec fits two scalars per dimension—**shift** and **scale**—to map empirical quantiles onto the canonical Beta marginal. This calibration eliminates the drift that low-bit quantization usually introduces, ensuring accuracy without a pre-computed training set.

### Lloyd-Max Codebook Generation

With the target distribution known analytically, optimal scalar quantization boundaries are pre-computed using Lloyd-Max iteration. The system supports **2‑bit** (4 buckets) and **4‑bit** (16 buckets) configurations without any data-driven training step.

### Bit-Packing Compression

Quantized integers are tightly packed into bytes, yielding **16× compression** for 2‑bit codes at `d = 1536` (1536 dimensions compress from 6144 bytes to 384 bytes per vector). This memory layout is what enables the small RAM footprint reported in benchmarks.

### Length-Renormalization Scoring

During search, a per-vector scalar—computed as `||v|| / ⟨u, x̂⟩`—corrects the systematic inner-product underestimation introduced by quantization. This makes final scores unbiased without requiring extra metadata storage per document.

### SIMD Search Kernels

Queries are rotated once, then scored directly against the codebook using hand-written **NEON** kernels on ARM and **AVX-512BW** kernels on x86. Blocked-wise allowlist filtering occurs inside the kernel itself, avoiding the post-filter over-fetch problem common in FAISS workflows.

## Choosing the Right Index Class

`turbovec` exposes two index classes in both Python and Rust. Selecting the correct class is the first migration decision.

### TurboQuantIndex (Positional IDs)

`TurboQuantIndex` is a slot-based positional index that offers the fastest ingestion and lowest overhead. It is ideal when you never delete vectors or can tolerate slot renumbering. Key methods include `add`, `search`, `swap_remove`, `write`, and `load`, documented in [`docs/api.md`](https://github.com/RyanCodrai/turbovec/blob/main/docs/api.md).

### IdMapIndex (Stable External IDs)

`IdMapIndex` wraps `TurboQuantIndex` with a hash-table-backed mapping from stable external `uint64` IDs to internal slots. Use this when external identifiers must survive deletions. It exposes `add_with_ids`, `search` with an optional `allowlist`, `remove`, and symmetric `write`/`load` operations.

## Step-by-Step FAISS to turbovec Migration

Follow these five steps to move an existing FAISS pipeline to turbovec.

1. **Replace FAISS index construction** with a turbovec index. Choose `TurboQuantIndex` for simple append-only pipelines or `IdMapIndex` when you need stable document IDs.
2. **Ingest vectors** using `add` for positional mode or `add_with_ids` for stable IDs. turbovec infers dimensionality automatically on the first add, so explicit `dim` arguments are unnecessary in Python.
3. **Persist the index** with `write("my_index.tv")` (`TurboQuantIndex`) or `write("my_index.tvim")` (`IdMapIndex`). Loading is symmetric via the class-level `load` method.
4. **Query** using `search(query, k)`. Optionally supply a boolean `mask` (positional) or a `uint64` `allowlist` (stable IDs) for hybrid retrieval or multi-tenant filtering; the filter is applied inside the SIMD kernel so only allowed candidates are scored.
5. **Delete** (if needed) with `swap_remove(slot)` (positional) or `remove(id)` (stable IDs); both operations are **O(1)**.

## Python and Rust Code Examples

### Python: Positional Index with TurboQuantIndex

```python
from turbovec import TurboQuantIndex
import numpy as np

# Create a 4-bit index (dimension inferred on first add)

index = TurboQuantIndex(bit_width=4)

# Add a batch of vectors (shape: n × dim)

vectors = np.random.randn(100_000, 1536).astype(np.float32)
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True)   # L2-normalize

index.add(vectors)

# Search

query = np.random.randn(1, 1536).astype(np.float32)
query /= np.linalg.norm(query)
scores, slots = index.search(query, k=10)

# Persist

index.write("my_index.tv")
loaded = TurboQuantIndex.load("my_index.tv")

```

*(see [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md) lines 31-38)*

### Python: Stable IDs and Allowlist Filtering with IdMapIndex

```python
from turbovec import IdMapIndex
import numpy as np

ids = np.arange(100_000, dtype=np.uint64)
vectors = np.random.randn(100_000, 1536).astype(np.float32)
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True)

index = IdMapIndex(bit_width=4)
index.add_with_ids(vectors, ids)

# Hybrid retrieval: restrict to a candidate set

candidate_ids = np.array([10, 20, 30, 40], dtype=np.uint64)
scores, ids_out = index.search(vectors[:5], k=5, allowlist=candidate_ids)

# Delete a document

index.remove(20)

# Save / load

index.write("my_index.tvim")
loaded = IdMapIndex.load("my_index.tvim")

```

*(see [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md) lines 44-58)*

### Rust: Positional Index

```rust
use turbovec::TurboQuantIndex;

let mut idx = TurboQuantIndex::new(1536, 4).unwrap();
idx.add(&vectors);                     // vectors: &[f32] shaped (n, dim)
let results = idx.search(&queries, 10);
idx.write("index.tv").unwrap();
let loaded = TurboQuantIndex::load("index.tv").unwrap();

```

*(see [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md) lines 94-100)*

### Rust: Stable ID Index

```rust
use turbovec::IdMapIndex;

let mut idx = IdMapIndex::new(1536, 4).unwrap();
idx.add_with_ids(&vectors, &[1001, 1002, 1003]).unwrap();
let (scores, ids) = idx.search(&queries, 10);
idx.remove(1002);
idx.write("index.tvim").unwrap();
let loaded = IdMapIndex::load("index.tvim").unwrap();

```

*(see [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md) lines 111-120)*

## Key Source Files and Integrations

Understanding the repository layout helps when debugging or extending turbovec:

- [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md) — High-level overview, benchmark summaries, and usage snippets.
- [`docs/api.md`](https://github.com/RyanCodrai/turbovec/blob/main/docs/api.md) — Full API reference for `TurboQuantIndex` and `IdMapIndex`.
- [`turbovec-python/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/src/lib.rs) — Python bindings implementation covering validation, search dispatch, and serialization.
- [`turbovec/Cargo.toml`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/Cargo.toml) — Native Rust crate definition and feature flags.
- [`turbovec-python/Cargo.toml`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/Cargo.toml) — Python package build configuration using Maturin.
- [`benchmarks/suite/recall_d1536_4bit.py`](https://github.com/RyanCodrai/turbovec/blob/main/benchmarks/suite/recall_d1536_4bit.py) — Reproducible recall benchmark comparing turbovec against FAISS.
- [`docs/integrations/langchain.md`](https://github.com/RyanCodrai/turbovec/blob/main/docs/integrations/langchain.md) — Drop-in LangChain integration example.

## Summary

- **Migrating from FAISS to turbovec** eliminates the offline training phase required by product-quantization indexes such as `IndexPQ`.
- turbovec’s TurboQuant algorithm achieves **16× compression** with 2-bit codes and near-optimal distortion via per-coordinate calibration and length-renormalization.
- Two index classes—`TurboQuantIndex` and `IdMapIndex`—cover both high-throughput positional workflows and stable-ID production systems.
- **Allowlist filtering** happens inside hand-written SIMD kernels (NEON/AVX-512BW), guaranteeing that top‑k results are drawn exclusively from the permitted set.
- Serialization is symmetric via `write` and `load`, with file extensions `.tv` and `.tvim` distinguishing the two index types.

## Frequently Asked Questions

### Does turbovec require a separate training phase before querying?

No. turbovec implements an **online** quantizer. When you call `add` for the first time, the index infers dimensionality and runs per-coordinate calibration automatically. This is a fundamental departure from FAISS `IndexPQ`, which requires an explicit `train` step on a representative vector sample before any vectors can be added.

### How does turbovec handle filtering compared to FAISS?

turbovec applies **allowlist filtering inside the SIMD kernel** itself. Whether you pass a boolean `mask` to `TurboQuantIndex.search` or a `uint64` `allowlist` to `IdMapIndex.search`, the kernel only scores and returns candidates from the permitted set. In many FAISS pipelines, filtering is performed after the search, which can lead to over-fetching and lower effective recall.

### Can I delete vectors from a turbovec index without rebuilding it?

Yes. `TurboQuantIndex` supports `swap_remove(slot)` and `IdMapIndex` supports `remove(id)`, both of which operate in **O(1)** time. Because `IdMapIndex` maintains a stable hash-table mapping, external identifiers remain valid after deletions, making it suitable for long-running production indexes that require document eviction.

### What bit widths does turbovec support, and how much memory does it save?

turbovec supports **2‑bit** and **4‑bit** quantization. At `d = 1536`, 2‑bit codes compress each vector from 6144 bytes to 384 bytes—a **16× reduction**. In benchmarked corpora of 10 million documents, this translates to approximately 4 GB of RAM for turbovec versus roughly 31 GB for an equivalent FAISS `IndexPQ` setup, while maintaining comparable or superior recall.