Migrating from FAISS IndexPQ to turbovec: A Complete Guide

Migrating from FAISS IndexPQ to turbovec eliminates the training step, enables 8×–16× compression with 2-bit or 4-bit quantization, and provides built-in allowlist filtering while maintaining FAISS-compatible Python and Rust APIs.

If you are currently using FAISS IndexPQ for approximate nearest neighbor (ANN) search, the RyanCodrai/turbovec repository offers a drop-in replacement with significant performance and usability improvements. This guide covers the architectural differences, migration steps, and implementation details for both Python and Rust codebases.

Why Switch from FAISS IndexPQ to turbovec?

Training-Free Quantization

FAISS IndexPQ requires a separate training step using k-means to learn codebook centroids on representative data. In turbovec/src/encode.rs, turbovec implements a data-oblivious approach that uses random rotation and a pre-computed Lloyd-Max codebook, allowing you to call add() immediately without any train() step.

Superior Compression and Memory Efficiency

While FAISS 8-bit PQ provides roughly 4× compression, turbovec offers configurable bit-widths:

  • 4-bit: 8× compression (16 centroids per sub-vector)
  • 2-bit: 16× compression (4 centroids per sub-vector)

For a 10 million document index with 1536-dimensional vectors, this reduces memory usage from approximately 60 GB (float32) to roughly 4 GB.

SIMD-Optimized Search with Filtering

The search kernels in turbovec/src/search.rs include hand-written NEON (ARM) and AVX-512BW (x86) implementations that operate directly on packed codes. These kernels beat FAISS FastScan by 12–20% on ARM and match or exceed performance on x86. Crucially, turbovec supports native hybrid filtering through an allowlist parameter that the kernel respects during search via the block_has_allowed logic (lines 13–21), guaranteeing exactly k results from your candidate set without post-filtering.

Stable External IDs and O(1) Deletion

Unlike FAISS, which returns internal indices that require manual mapping, turbovec's IdMapIndex stores stable uint64 external IDs per vector and supports remove(id) in O(1) time.

Step-by-Step Migration Guide

1. Installation

Replace your FAISS dependency with turbovec:


# Python

pip install turbovec

# Rust (Cargo.toml)

[dependencies]
turbovec = "0.1"

2. Replace Index Creation

Eliminate the training step entirely:


# Before: FAISS IndexPQ

import faiss
d = 1536
M = d // 8  # 8 sub-quantizers for 8-bit

index = faiss.IndexPQ(d, M, 8)
index.train(train_vectors)  # Required

index.add(db_vectors)

# After: turbovec TurboQuantIndex

from turbovec import TurboQuantIndex
index = TurboQuantIndex(dim=d, bit_width=4)  # 4-bit quantization

index.add(db_vectors)  # Immediate indexing, no training

3. Migrate External ID Mapping

Replace IndexIDMap2 with IdMapIndex:

from turbovec import IdMapIndex

idx = IdMapIndex(dim=1536, bit_width=4)
idx.add_with_ids(vectors, external_ids)  # external_ids: uint64 numpy array

# O(1) deletion by external ID

idx.remove(target_id)

4. Implement Hybrid Filtering

Pass candidate sets directly to the search kernel:

import numpy as np

# Candidate set from BM25 or SQL filter

allowed = np.array([123, 456, 789], dtype=np.uint64)
scores, ids = idx.search(query, k=10, allowlist=allowed)

The allowlist mask is processed inside the SIMD kernels (search_multi_query_avx2, search_multi_query_avx512bw) via the block_has_allowed check, ensuring only allowed IDs are considered during the distance computation.

5. Persistence

Save and load indices using the same pattern as FAISS:


# Save

index.write("my_index.tq")

# Load

loaded = TurboQuantIndex.load("my_index.tq")

Core Architecture Differences

The Encoding Pipeline (turbovec/src/encode.rs)

According to the turbovec source code, the encoding process follows this pipeline:

  1. Normalization – Vectors are normalized for inner product search
  2. Random Rotation – Data-oblivious rotation via Hadamard matrix
  3. TQ+ Calibration – Per-vector scale computation (norm / <u, x̂>)
  4. Lloyd-Max Quantization – Vector quantization using pre-computed codebooks
  5. Bit-Packing – Packing codes into 2-bit or 4-bit representations

The per-vector scale is stored as f32 (see lines 1–8 and 46–53 in encode.rs) to enable unbiased inner-product scoring during search.

SIMD Search Kernels (turbovec/src/search.rs)

The search implementation provides three optimized paths:

  • score_4bit_block_neon – ARM NEON intrinsics for mobile/edge devices
  • search_multi_query_avx2 – AVX2 implementation for x86_64
  • search_multi_query_avx512bw – AVX-512BW implementation for server CPUs

These kernels (lines 34–84, 150–170, and 290–320) decode packed codes on-the-fly and apply the per-vector scales computed during encoding, avoiding the need for separate codebook lookups.

Rust Implementation Example

For Rust developers, the API mirrors the Python interface:

use turbovec::{IdMapIndex, TurboQuantIndex};

fn main() -> anyhow::Result<()> {
    // Create 4-bit index
    let dim = 1536usize;
    let mut index = IdMapIndex::new(dim, 4);
    
    // Add with external IDs
    let vectors: Vec<f32> = vec![0.1; 10_000 * dim];
    let ids: Vec<u64> = (0..10_000).map(|i| i + 1_000_000).collect();
    index.add_with_ids(&vectors, &ids);
    
    // Search
    let query = vec![0.1f32; dim];
    let (scores, ext_ids) = index.search(&query, 10);
    
    // Persist
    index.write("index.tvim")?;
    let loaded = IdMapIndex::load("index.tvim")?;
    
    Ok(())
}

Performance Considerations and Best Practices

Consideration FAISS IndexPQ turbovec Recommendation
Training Data Requires large corpus for k-means convergence None needed Skip training entirely; add vectors immediately
Bit-Width 8-bit standard 2-bit or 4-bit Use 4-bit for best recall/speed trade-off; 2-bit for maximum compression
Vector Norms Stores codes only Stores f32 scale per vector Ensure input vectors are normalized; turbovec handles norm storage internally
GPU Support CUDA backends available CPU-only (SIMD) Use turbovec for low-latency CPU search on edge devices
Filtering Manual post-filtering Kernel-level allowlist Use turbovec when combining dense vectors with sparse retrieval (BM25/SQL)

Summary

  • Drop-in replacement: Swap faiss.IndexPQ with TurboQuantIndex or IdMapIndex and remove all train() calls
  • Higher compression: Achieve 8× (4-bit) or 16× (2-bit) compression versus FAISS's typical 4×
  • Zero training: Random rotation and Lloyd-Max codebooks eliminate the training step
  • Native filtering: The allowlist parameter performs kernel-level masking without post-processing
  • Stable IDs: IdMapIndex provides uint64 external IDs and O(1) deletion by ID

Frequently Asked Questions

Does turbovec require a training step like FAISS IndexPQ?

No. According to the implementation in turbovec/src/encode.rs, turbovec uses a data-oblivious random rotation and pre-computed Lloyd-Max codebooks. You can call add() or add_with_ids() immediately after instantiation without preparing training data or calling a train() method.

How do I migrate from FAISS IndexIDMap to turbovec?

Replace faiss.IndexIDMap2 with turbovec.IdMapIndex. The add_with_ids() method accepts a uint64 numpy array (Python) or slice (Rust) of external IDs. Unlike FAISS, which returns internal indices requiring mapping, IdMapIndex.search() returns your original external IDs directly, and supports remove(id) in O(1) time.

Can turbovec handle filtered search for hybrid retrieval?

Yes. turbovec natively supports hybrid filtering through the allowlist parameter in the search() method. As implemented in turbovec/src/search.rs (lines 13–21), the SIMD kernels check block_has_allowed during distance computation, ensuring only IDs from your candidate set are considered. This guarantees you receive exactly k results from the allowed set, eliminating the need for post-filtering and re-querying.

What bit-width should I choose when migrating?

Choose 4-bit (8× compression) for production workloads requiring the best recall-speed trade-off, as it provides 16 centroids per sub-vector. Choose 2-bit (16× compression) for aggressive memory reduction on resource-constrained devices. Both are implemented in turbovec/src/search.rs with optimized SIMD kernels for ARM and x86 architectures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →