Migrating from FAISS to turbovec for Vector Search: A Complete Infrastructure Guide
Migrating from FAISS to turbovec lets you replace offline-trained product quantization with an online, data-oblivious quantizer that cuts RAM usage by up to 8×, eliminates index rebuilds, and performs filtered top‑k retrieval directly inside SIMD kernels.
If you are migrating from FAISS to turbovec for vector search infrastructure, you are moving from a batch-trained, memory-heavy IndexPQ pipeline to a Rust-based implementation of Google Research’s TurboQuant algorithm. The turbovec library—available in the RyanCodrai/turbovec repository—provides Python bindings and native Rust APIs that calibrate quantization parameters automatically during the first add call, allowing continuous streaming ingestion without a separate training phase.
Why turbovec Replaces FAISS for Production Workloads
Compared with FAISS’s IndexPQ (and its FastScan variant), turbovec delivers the same or better recall with dramatically lower memory consumption. As documented in README.md, a 10‑million‑document corpus fits in roughly 4 GB with turbovec versus approximately 31 GB for FAISS, while query latency is typically faster on both ARM and x86 CPUs. Because quantization is online, you can add new vectors at any time without rebuilding the entire index—a common bottleneck in FAISS PQ pipelines.
How turbovec Implements TurboQuant
The core algorithm described in README.md consists of six stages that run automatically during ingestion and search.
Normalization and Random Rotation
Each incoming vector is L2-normalized and multiplied by a single random orthogonal matrix. This transforms every coordinate into an identical Beta distribution— approximately normal with zero mean and variance 1/d—regardless of the original data distribution. This step is the statistical foundation that makes TurboQuant data-oblivious.
Per-Coordinate Calibration (TQ+)
On the first call to add, turbovec fits two scalars per dimension—shift and scale—to map empirical quantiles onto the canonical Beta marginal. This calibration eliminates the drift that low-bit quantization usually introduces, ensuring accuracy without a pre-computed training set.
Lloyd-Max Codebook Generation
With the target distribution known analytically, optimal scalar quantization boundaries are pre-computed using Lloyd-Max iteration. The system supports 2‑bit (4 buckets) and 4‑bit (16 buckets) configurations without any data-driven training step.
Bit-Packing Compression
Quantized integers are tightly packed into bytes, yielding 16× compression for 2‑bit codes at d = 1536 (1536 dimensions compress from 6144 bytes to 384 bytes per vector). This memory layout is what enables the small RAM footprint reported in benchmarks.
Length-Renormalization Scoring
During search, a per-vector scalar—computed as ||v|| / ⟨u, x̂⟩—corrects the systematic inner-product underestimation introduced by quantization. This makes final scores unbiased without requiring extra metadata storage per document.
SIMD Search Kernels
Queries are rotated once, then scored directly against the codebook using hand-written NEON kernels on ARM and AVX-512BW kernels on x86. Blocked-wise allowlist filtering occurs inside the kernel itself, avoiding the post-filter over-fetch problem common in FAISS workflows.
Choosing the Right Index Class
turbovec exposes two index classes in both Python and Rust. Selecting the correct class is the first migration decision.
TurboQuantIndex (Positional IDs)
TurboQuantIndex is a slot-based positional index that offers the fastest ingestion and lowest overhead. It is ideal when you never delete vectors or can tolerate slot renumbering. Key methods include add, search, swap_remove, write, and load, documented in docs/api.md.
IdMapIndex (Stable External IDs)
IdMapIndex wraps TurboQuantIndex with a hash-table-backed mapping from stable external uint64 IDs to internal slots. Use this when external identifiers must survive deletions. It exposes add_with_ids, search with an optional allowlist, remove, and symmetric write/load operations.
Step-by-Step FAISS to turbovec Migration
Follow these five steps to move an existing FAISS pipeline to turbovec.
- Replace FAISS index construction with a turbovec index. Choose
TurboQuantIndexfor simple append-only pipelines orIdMapIndexwhen you need stable document IDs. - Ingest vectors using
addfor positional mode oradd_with_idsfor stable IDs. turbovec infers dimensionality automatically on the first add, so explicitdimarguments are unnecessary in Python. - Persist the index with
write("my_index.tv")(TurboQuantIndex) orwrite("my_index.tvim")(IdMapIndex). Loading is symmetric via the class-levelloadmethod. - Query using
search(query, k). Optionally supply a booleanmask(positional) or auint64allowlist(stable IDs) for hybrid retrieval or multi-tenant filtering; the filter is applied inside the SIMD kernel so only allowed candidates are scored. - Delete (if needed) with
swap_remove(slot)(positional) orremove(id)(stable IDs); both operations are O(1).
Python and Rust Code Examples
Python: Positional Index with TurboQuantIndex
from turbovec import TurboQuantIndex
import numpy as np
# Create a 4-bit index (dimension inferred on first add)
index = TurboQuantIndex(bit_width=4)
# Add a batch of vectors (shape: n × dim)
vectors = np.random.randn(100_000, 1536).astype(np.float32)
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True) # L2-normalize
index.add(vectors)
# Search
query = np.random.randn(1, 1536).astype(np.float32)
query /= np.linalg.norm(query)
scores, slots = index.search(query, k=10)
# Persist
index.write("my_index.tv")
loaded = TurboQuantIndex.load("my_index.tv")
(see README.md lines 31-38)
Python: Stable IDs and Allowlist Filtering with IdMapIndex
from turbovec import IdMapIndex
import numpy as np
ids = np.arange(100_000, dtype=np.uint64)
vectors = np.random.randn(100_000, 1536).astype(np.float32)
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True)
index = IdMapIndex(bit_width=4)
index.add_with_ids(vectors, ids)
# Hybrid retrieval: restrict to a candidate set
candidate_ids = np.array([10, 20, 30, 40], dtype=np.uint64)
scores, ids_out = index.search(vectors[:5], k=5, allowlist=candidate_ids)
# Delete a document
index.remove(20)
# Save / load
index.write("my_index.tvim")
loaded = IdMapIndex.load("my_index.tvim")
(see README.md lines 44-58)
Rust: Positional Index
use turbovec::TurboQuantIndex;
let mut idx = TurboQuantIndex::new(1536, 4).unwrap();
idx.add(&vectors); // vectors: &[f32] shaped (n, dim)
let results = idx.search(&queries, 10);
idx.write("index.tv").unwrap();
let loaded = TurboQuantIndex::load("index.tv").unwrap();
(see README.md lines 94-100)
Rust: Stable ID Index
use turbovec::IdMapIndex;
let mut idx = IdMapIndex::new(1536, 4).unwrap();
idx.add_with_ids(&vectors, &[1001, 1002, 1003]).unwrap();
let (scores, ids) = idx.search(&queries, 10);
idx.remove(1002);
idx.write("index.tvim").unwrap();
let loaded = IdMapIndex::load("index.tvim").unwrap();
(see README.md lines 111-120)
Key Source Files and Integrations
Understanding the repository layout helps when debugging or extending turbovec:
README.md— High-level overview, benchmark summaries, and usage snippets.docs/api.md— Full API reference forTurboQuantIndexandIdMapIndex.turbovec-python/src/lib.rs— Python bindings implementation covering validation, search dispatch, and serialization.turbovec/Cargo.toml— Native Rust crate definition and feature flags.turbovec-python/Cargo.toml— Python package build configuration using Maturin.benchmarks/suite/recall_d1536_4bit.py— Reproducible recall benchmark comparing turbovec against FAISS.docs/integrations/langchain.md— Drop-in LangChain integration example.
Summary
- Migrating from FAISS to turbovec eliminates the offline training phase required by product-quantization indexes such as
IndexPQ. - turbovec’s TurboQuant algorithm achieves 16× compression with 2-bit codes and near-optimal distortion via per-coordinate calibration and length-renormalization.
- Two index classes—
TurboQuantIndexandIdMapIndex—cover both high-throughput positional workflows and stable-ID production systems. - Allowlist filtering happens inside hand-written SIMD kernels (NEON/AVX-512BW), guaranteeing that top‑k results are drawn exclusively from the permitted set.
- Serialization is symmetric via
writeandload, with file extensions.tvand.tvimdistinguishing the two index types.
Frequently Asked Questions
Does turbovec require a separate training phase before querying?
No. turbovec implements an online quantizer. When you call add for the first time, the index infers dimensionality and runs per-coordinate calibration automatically. This is a fundamental departure from FAISS IndexPQ, which requires an explicit train step on a representative vector sample before any vectors can be added.
How does turbovec handle filtering compared to FAISS?
turbovec applies allowlist filtering inside the SIMD kernel itself. Whether you pass a boolean mask to TurboQuantIndex.search or a uint64 allowlist to IdMapIndex.search, the kernel only scores and returns candidates from the permitted set. In many FAISS pipelines, filtering is performed after the search, which can lead to over-fetching and lower effective recall.
Can I delete vectors from a turbovec index without rebuilding it?
Yes. TurboQuantIndex supports swap_remove(slot) and IdMapIndex supports remove(id), both of which operate in O(1) time. Because IdMapIndex maintains a stable hash-table mapping, external identifiers remain valid after deletions, making it suitable for long-running production indexes that require document eviction.
What bit widths does turbovec support, and how much memory does it save?
turbovec supports 2‑bit and 4‑bit quantization. At d = 1536, 2‑bit codes compress each vector from 6144 bytes to 384 bytes—a 16× reduction. In benchmarked corpora of 10 million documents, this translates to approximately 4 GB of RAM for turbovec versus roughly 31 GB for an equivalent FAISS IndexPQ setup, while maintaining comparable or superior recall.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →