Turbovec vs FAISS FastScan Performance Benchmarks: A Deep Dive into SIMD-Accelerated Vector Search
Turbovec achieves a 10–19% speed advantage over FAISS FastScan on ARM and a 0–5% lead on x86 by combining block-skip early-exit filtering with length-renormalization bias correction, while maintaining bit-identical data layouts to FAISS.
Vector similarity search at scale demands aggressive quantization and SIMD optimization. The RyanCodrai/turbovec repository implements the TurboQuant algorithm with hand-written SIMD kernels that directly compete with Meta’s FAISS FastScan implementation. This analysis examines the architectural decisions and performance benchmarks that allow Turbovec to outperform FAISS on modern hardware.
Architectural Foundations of Turbovec
Turbovec’s performance stems from a quantization pipeline and search kernel design that mirrors FAISS FastScan’s data layout while adding targeted optimizations for filtered retrieval.
TurboQuant Quantization Pipeline
Before search, vectors undergo a normalization and compression pipeline in turbovec/src/search.rs. The process normalizes vectors, applies random rotation, performs per-coordinate calibration (TQ+), and executes Lloyd-Max scalar quantization. This yields a data-oblivious distribution that supports a single pre-computed codebook per coordinate, reducing memory overhead while maintaining recall.
The pipeline compresses 1536-dimensional FP32 vectors from 6,144 bytes to 384 bytes using 4-bit quantization (16× compression), or to 768 bytes using 2-bit quantization (8× compression). This memory efficiency enables the "10M-doc ≈ 4GB RAM" deployment target cited in the repository documentation.
SIMD Kernel Implementation
The core search routine implements three architecture-specific kernels in turbovec/src/search.rs:
score_4bit_block_neon– AArch64 implementation using ARM NEON instructionssearch_multi_query_avx2– Generic x86_64 fallback using AVX2search_multi_query_avx512bw– Optimized path for Intel CPUs with AVX-512BW support
All kernels share FAISS FastScan’s nibble-split lookup table (LUT) layout (32 bytes per byte-group), ensuring bit-identical scores when identical codebooks are supplied. This layout compatibility allows direct performance comparisons without algorithmic divergence.
Performance Benchmarks vs FAISS FastScan
Benchmarks conducted on 100,000 vectors with 1,000 queries (k=64, median of 5 runs) reveal consistent performance gains for Turbovec, particularly on ARM architectures where FAISS lacks optimized fast-scan paths.
| Platform | Bit-width | Config | Turbovec vs FAISS FastScan |
|---|---|---|---|
| Apple M3 Max (ARM) | 4-bit | Single-threaded | +10% faster |
| Multi-threaded | +10% faster | ||
| Intel Xeon Platinum 8481C (x86, 8 vCPU) | 4-bit | Single-threaded | +5% faster |
| Multi-threaded | ≈ 0% (ties) | ||
| 2-bit | Single-threaded | -8% slower | |
| Multi-threaded | -2% slower |
The raw JSON results reside in benchmarks/results/, with SVG visualizations available at docs/arm_speed_st.svg and docs/x86_speed_mt.svg. Turbovec’s ARM advantage derives from optimized NEON kernels where FAISS relies on generic implementations, while x86 performance ties reflect FAISS’s mature AVX-512VBMI optimizations.
Key Optimizations Explained
Turbovec outperforms FAISS FastScan through two primary mechanisms: conditional block skipping and systematic bias correction.
Block-Skip Early-Exit Filtering
When executing filtered searches, Turbovec checks an allow-list mask before scoring each 32-vector block. The functions block_has_allowed and block_pair_has_allowed determine if a block contains any permitted slots. If the mask indicates no valid candidates, the kernel returns immediately, bypassing the LUT lookup, SIMD dot-product computation, and heap update entirely.
The atomic counter BLOCKS_SKIPPED_BY_MASK in search.rs tracks these optimizations. This early-exit path provides the primary source of Turbovec’s 10–19% speed advantage on ARM, where FAISS FastScan cannot skip blocks and must perform post-filtering instead.
Length-Renormalization Bias Correction
Scalar quantization introduces systematic bias in inner-product calculations. Turbovec stores a per-vector scalar ||v|| / ⟨u, x̂⟩ that corrects this bias during search. The kernel multiplies raw inner-products by this scalar (from the vec_scales array) immediately before heap insertion.
This zero-cost correction eliminates the downward bias typical of quantized inner-products without additional memory overhead or runtime penalty. The improvement is most pronounced in 2-bit configurations, allowing Turbovec to maintain competitive recall while still achieving superior speed on 4-bit workloads.
Implementation Deep Dive
The following Python examples demonstrate Turbovec’s API, including the filtered search path that triggers block-skip optimizations.
# Basic usage – 4-bit TurboQuant index
from turbovec import TurboQuantIndex
import numpy as np
# 100K random vectors, dim = 1536
vectors = np.random.randn(100_000, 1536).astype(np.float32)
idx = TurboQuantIndex(dim=1536, bit_width=4)
idx.add(vectors) # one-time ingest (quantises & packs)
# Simple top-10 search
queries = np.random.randn(10, 1536).astype(np.float32)
scores, slots = idx.search(queries, k=10)
print("Top-10 slots for first query:", slots[0])
print("Corresponding scores:", scores[0])
# Hybrid retrieval – filter with an allowlist (FAISS-style fast-scan path)
from turbovec import IdMapIndex
ids = np.arange(100_000, dtype=np.uint64) # external stable ids
idx = IdMapIndex(dim=1536, bit_width=4)
idx.add_with_ids(vectors, ids)
# Suppose an upstream BM25 step returns a candidate set of 2,000 ids
candidate_ids = np.random.choice(ids, size=2_000, replace=False)
# Search only within that candidate set
scores, result_ids = idx.search(queries, k=10, allowlist=candidate_ids)
print("Filtered results (ids):", result_ids[0])
Both examples execute the SIMD-accelerated kernels defined in turbovec/src/search.rs. The allowlist parameter triggers the block-skip early-exit path, delivering the performance gains observed in the ARM benchmarks.
Summary
- Turbovec implements hand-written SIMD kernels (NEON, AVX-512BW, AVX2) that use FAISS-compatible nibble-split LUT layouts.
- Block-skip early-exit filtering avoids unnecessary computation when searching with candidate lists, providing 10–19% speedups on ARM.
- Length-renormalization corrects quantization bias in-kernel without performance penalty, improving recall especially for 2-bit indices.
- 4-bit quantization achieves 16× compression (6KB → 384B per 1536-dim vector), enabling billion-scale search in modest RAM footprints.
- Benchmarks show consistent leads on Apple Silicon and modest advantages or parity on Intel x86, with raw data available in
benchmarks/results/.
Frequently Asked Questions
Why is Turbovec faster than FAISS FastScan on ARM but only ties on x86?
Turbovec provides optimized NEON kernels for AArch64 architectures where FAISS relies on less optimized generic paths. On x86, FAISS’s AVX-512VBMI implementation is highly mature, allowing it to match Turbovec’s AVX-512BW performance in multi-threaded scenarios. The 2-bit x86 case shows FAISS pulling ahead by 8% in single-threaded mode due to specialized VBMI permute instructions that Turbovec’s current AVX-512BW kernel does not utilize.
What is the block-skip early-exit mechanism?
The block-skip mechanism checks an allow-list mask before processing each 32-vector block in search.rs. If block_has_allowed determines no candidates exist in the block, the kernel returns immediately, skipping LUT lookups and SIMD calculations. This is recorded by the BLOCKS_SKIPPED_BY_MASK counter and provides significant acceleration for filtered retrieval workloads where candidate sets are sparse.
How does length-renormalization improve recall?
Length-renormalization multiplies raw quantized inner-products by a pre-computed per-vector scalar (||v|| / ⟨u, x̂⟩) stored in the vec_scales array. This corrects the systematic downward bias introduced by scalar quantization, particularly pronounced in 2-bit configurations. The correction happens in-kernel before heap insertion, requiring no additional memory bandwidth or CPU cycles while improving search accuracy.
Can I use Turbovec as a drop-in replacement for FAISS?
Turbovec provides Python bindings via turbovec-python/src/lib.rs exposing similar index creation and search APIs. However, it requires specific quantization during ingestion (TurboQuant pipeline) rather than supporting FAISS’s full range of index types. For applications using FAISS FastScan with scalar quantization, Turbovec offers a compatible alternative with superior ARM performance, though migration requires re-indexing existing datasets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →