Migrating from FAISS to turbovec for Vector Search: A Complete Infrastructure Guide

Migrating from FAISS to turbovec lets you replace offline-trained product quantization with an online, data-oblivious quantizer that cuts RAM usage by up to 8×, eliminates index rebuilds, and performs filtered top‑k retrieval directly inside SIMD kernels.

If you are migrating from FAISS to turbovec for vector search infrastructure, you are moving from a batch-trained, memory-heavy IndexPQ pipeline to a Rust-based implementation of Google Research’s TurboQuant algorithm. The turbovec library—available in the RyanCodrai/turbovec repository—provides Python bindings and native Rust APIs that calibrate quantization parameters automatically during the first add call, allowing continuous streaming ingestion without a separate training phase.

Why turbovec Replaces FAISS for Production Workloads

Compared with FAISS’s IndexPQ (and its FastScan variant), turbovec delivers the same or better recall with dramatically lower memory consumption. As documented in README.md, a 10‑million‑document corpus fits in roughly 4 GB with turbovec versus approximately 31 GB for FAISS, while query latency is typically faster on both ARM and x86 CPUs. Because quantization is online, you can add new vectors at any time without rebuilding the entire index—a common bottleneck in FAISS PQ pipelines.

How turbovec Implements TurboQuant

The core algorithm described in README.md consists of six stages that run automatically during ingestion and search.

Normalization and Random Rotation

Each incoming vector is L2-normalized and multiplied by a single random orthogonal matrix. This transforms every coordinate into an identical Beta distribution— approximately normal with zero mean and variance 1/d—regardless of the original data distribution. This step is the statistical foundation that makes TurboQuant data-oblivious.

Per-Coordinate Calibration (TQ+)

On the first call to add, turbovec fits two scalars per dimension—shift and scale—to map empirical quantiles onto the canonical Beta marginal. This calibration eliminates the drift that low-bit quantization usually introduces, ensuring accuracy without a pre-computed training set.

Lloyd-Max Codebook Generation

With the target distribution known analytically, optimal scalar quantization boundaries are pre-computed using Lloyd-Max iteration. The system supports 2‑bit (4 buckets) and 4‑bit (16 buckets) configurations without any data-driven training step.

Bit-Packing Compression

Quantized integers are tightly packed into bytes, yielding 16× compression for 2‑bit codes at d = 1536 (1536 dimensions compress from 6144 bytes to 384 bytes per vector). This memory layout is what enables the small RAM footprint reported in benchmarks.

Length-Renormalization Scoring

During search, a per-vector scalar—computed as ||v|| / ⟨u, x̂⟩—corrects the systematic inner-product underestimation introduced by quantization. This makes final scores unbiased without requiring extra metadata storage per document.

SIMD Search Kernels

Queries are rotated once, then scored directly against the codebook using hand-written NEON kernels on ARM and AVX-512BW kernels on x86. Blocked-wise allowlist filtering occurs inside the kernel itself, avoiding the post-filter over-fetch problem common in FAISS workflows.

Choosing the Right Index Class

turbovec exposes two index classes in both Python and Rust. Selecting the correct class is the first migration decision.

TurboQuantIndex (Positional IDs)

TurboQuantIndex is a slot-based positional index that offers the fastest ingestion and lowest overhead. It is ideal when you never delete vectors or can tolerate slot renumbering. Key methods include add, search, swap_remove, write, and load, documented in docs/api.md.

IdMapIndex (Stable External IDs)

IdMapIndex wraps TurboQuantIndex with a hash-table-backed mapping from stable external uint64 IDs to internal slots. Use this when external identifiers must survive deletions. It exposes add_with_ids, search with an optional allowlist, remove, and symmetric write/load operations.

Step-by-Step FAISS to turbovec Migration

Follow these five steps to move an existing FAISS pipeline to turbovec.

  1. Replace FAISS index construction with a turbovec index. Choose TurboQuantIndex for simple append-only pipelines or IdMapIndex when you need stable document IDs.
  2. Ingest vectors using add for positional mode or add_with_ids for stable IDs. turbovec infers dimensionality automatically on the first add, so explicit dim arguments are unnecessary in Python.
  3. Persist the index with write("my_index.tv") (TurboQuantIndex) or write("my_index.tvim") (IdMapIndex). Loading is symmetric via the class-level load method.
  4. Query using search(query, k). Optionally supply a boolean mask (positional) or a uint64 allowlist (stable IDs) for hybrid retrieval or multi-tenant filtering; the filter is applied inside the SIMD kernel so only allowed candidates are scored.
  5. Delete (if needed) with swap_remove(slot) (positional) or remove(id) (stable IDs); both operations are O(1).

Python and Rust Code Examples

Python: Positional Index with TurboQuantIndex

from turbovec import TurboQuantIndex
import numpy as np

# Create a 4-bit index (dimension inferred on first add)

index = TurboQuantIndex(bit_width=4)

# Add a batch of vectors (shape: n × dim)

vectors = np.random.randn(100_000, 1536).astype(np.float32)
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True)   # L2-normalize

index.add(vectors)

# Search

query = np.random.randn(1, 1536).astype(np.float32)
query /= np.linalg.norm(query)
scores, slots = index.search(query, k=10)

# Persist

index.write("my_index.tv")
loaded = TurboQuantIndex.load("my_index.tv")

(see README.md lines 31-38)

Python: Stable IDs and Allowlist Filtering with IdMapIndex

from turbovec import IdMapIndex
import numpy as np

ids = np.arange(100_000, dtype=np.uint64)
vectors = np.random.randn(100_000, 1536).astype(np.float32)
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True)

index = IdMapIndex(bit_width=4)
index.add_with_ids(vectors, ids)

# Hybrid retrieval: restrict to a candidate set

candidate_ids = np.array([10, 20, 30, 40], dtype=np.uint64)
scores, ids_out = index.search(vectors[:5], k=5, allowlist=candidate_ids)

# Delete a document

index.remove(20)

# Save / load

index.write("my_index.tvim")
loaded = IdMapIndex.load("my_index.tvim")

(see README.md lines 44-58)

Rust: Positional Index

use turbovec::TurboQuantIndex;

let mut idx = TurboQuantIndex::new(1536, 4).unwrap();
idx.add(&vectors);                     // vectors: &[f32] shaped (n, dim)
let results = idx.search(&queries, 10);
idx.write("index.tv").unwrap();
let loaded = TurboQuantIndex::load("index.tv").unwrap();

(see README.md lines 94-100)

Rust: Stable ID Index

use turbovec::IdMapIndex;

let mut idx = IdMapIndex::new(1536, 4).unwrap();
idx.add_with_ids(&vectors, &[1001, 1002, 1003]).unwrap();
let (scores, ids) = idx.search(&queries, 10);
idx.remove(1002);
idx.write("index.tvim").unwrap();
let loaded = IdMapIndex::load("index.tvim").unwrap();

(see README.md lines 111-120)

Key Source Files and Integrations

Understanding the repository layout helps when debugging or extending turbovec:

Summary

  • Migrating from FAISS to turbovec eliminates the offline training phase required by product-quantization indexes such as IndexPQ.
  • turbovec’s TurboQuant algorithm achieves 16× compression with 2-bit codes and near-optimal distortion via per-coordinate calibration and length-renormalization.
  • Two index classes—TurboQuantIndex and IdMapIndex—cover both high-throughput positional workflows and stable-ID production systems.
  • Allowlist filtering happens inside hand-written SIMD kernels (NEON/AVX-512BW), guaranteeing that top‑k results are drawn exclusively from the permitted set.
  • Serialization is symmetric via write and load, with file extensions .tv and .tvim distinguishing the two index types.

Frequently Asked Questions

Does turbovec require a separate training phase before querying?

No. turbovec implements an online quantizer. When you call add for the first time, the index infers dimensionality and runs per-coordinate calibration automatically. This is a fundamental departure from FAISS IndexPQ, which requires an explicit train step on a representative vector sample before any vectors can be added.

How does turbovec handle filtering compared to FAISS?

turbovec applies allowlist filtering inside the SIMD kernel itself. Whether you pass a boolean mask to TurboQuantIndex.search or a uint64 allowlist to IdMapIndex.search, the kernel only scores and returns candidates from the permitted set. In many FAISS pipelines, filtering is performed after the search, which can lead to over-fetching and lower effective recall.

Can I delete vectors from a turbovec index without rebuilding it?

Yes. TurboQuantIndex supports swap_remove(slot) and IdMapIndex supports remove(id), both of which operate in O(1) time. Because IdMapIndex maintains a stable hash-table mapping, external identifiers remain valid after deletions, making it suitable for long-running production indexes that require document eviction.

What bit widths does turbovec support, and how much memory does it save?

turbovec supports 2‑bit and 4‑bit quantization. At d = 1536, 2‑bit codes compress each vector from 6144 bytes to 384 bytes—a 16× reduction. In benchmarked corpora of 10 million documents, this translates to approximately 4 GB of RAM for turbovec versus roughly 31 GB for an equivalent FAISS IndexPQ setup, while maintaining comparable or superior recall.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →