Turbovec vs FAISS FastScan Performance Benchmarks: A Deep Dive into SIMD-Accelerated Vector Search

Turbovec achieves a 10–19% speed advantage over FAISS FastScan on ARM and a 0–5% lead on x86 by combining block-skip early-exit filtering with length-renormalization bias correction, while maintaining bit-identical data layouts to FAISS.

Vector similarity search at scale demands aggressive quantization and SIMD optimization. The RyanCodrai/turbovec repository implements the TurboQuant algorithm with hand-written SIMD kernels that directly compete with Meta’s FAISS FastScan implementation. This analysis examines the architectural decisions and performance benchmarks that allow Turbovec to outperform FAISS on modern hardware.

Architectural Foundations of Turbovec

Turbovec’s performance stems from a quantization pipeline and search kernel design that mirrors FAISS FastScan’s data layout while adding targeted optimizations for filtered retrieval.

TurboQuant Quantization Pipeline

Before search, vectors undergo a normalization and compression pipeline in turbovec/src/search.rs. The process normalizes vectors, applies random rotation, performs per-coordinate calibration (TQ+), and executes Lloyd-Max scalar quantization. This yields a data-oblivious distribution that supports a single pre-computed codebook per coordinate, reducing memory overhead while maintaining recall.

The pipeline compresses 1536-dimensional FP32 vectors from 6,144 bytes to 384 bytes using 4-bit quantization (16× compression), or to 768 bytes using 2-bit quantization (8× compression). This memory efficiency enables the "10M-doc ≈ 4GB RAM" deployment target cited in the repository documentation.

SIMD Kernel Implementation

The core search routine implements three architecture-specific kernels in turbovec/src/search.rs:

  • score_4bit_block_neon – AArch64 implementation using ARM NEON instructions
  • search_multi_query_avx2 – Generic x86_64 fallback using AVX2
  • search_multi_query_avx512bw – Optimized path for Intel CPUs with AVX-512BW support

All kernels share FAISS FastScan’s nibble-split lookup table (LUT) layout (32 bytes per byte-group), ensuring bit-identical scores when identical codebooks are supplied. This layout compatibility allows direct performance comparisons without algorithmic divergence.

Performance Benchmarks vs FAISS FastScan

Benchmarks conducted on 100,000 vectors with 1,000 queries (k=64, median of 5 runs) reveal consistent performance gains for Turbovec, particularly on ARM architectures where FAISS lacks optimized fast-scan paths.

Platform Bit-width Config Turbovec vs FAISS FastScan
Apple M3 Max (ARM) 4-bit Single-threaded +10% faster
Multi-threaded +10% faster
Intel Xeon Platinum 8481C (x86, 8 vCPU) 4-bit Single-threaded +5% faster
Multi-threaded ≈ 0% (ties)
2-bit Single-threaded -8% slower
Multi-threaded -2% slower

The raw JSON results reside in benchmarks/results/, with SVG visualizations available at docs/arm_speed_st.svg and docs/x86_speed_mt.svg. Turbovec’s ARM advantage derives from optimized NEON kernels where FAISS relies on generic implementations, while x86 performance ties reflect FAISS’s mature AVX-512VBMI optimizations.

Key Optimizations Explained

Turbovec outperforms FAISS FastScan through two primary mechanisms: conditional block skipping and systematic bias correction.

Block-Skip Early-Exit Filtering

When executing filtered searches, Turbovec checks an allow-list mask before scoring each 32-vector block. The functions block_has_allowed and block_pair_has_allowed determine if a block contains any permitted slots. If the mask indicates no valid candidates, the kernel returns immediately, bypassing the LUT lookup, SIMD dot-product computation, and heap update entirely.

The atomic counter BLOCKS_SKIPPED_BY_MASK in search.rs tracks these optimizations. This early-exit path provides the primary source of Turbovec’s 10–19% speed advantage on ARM, where FAISS FastScan cannot skip blocks and must perform post-filtering instead.

Length-Renormalization Bias Correction

Scalar quantization introduces systematic bias in inner-product calculations. Turbovec stores a per-vector scalar ||v|| / ⟨u, x̂⟩ that corrects this bias during search. The kernel multiplies raw inner-products by this scalar (from the vec_scales array) immediately before heap insertion.

This zero-cost correction eliminates the downward bias typical of quantized inner-products without additional memory overhead or runtime penalty. The improvement is most pronounced in 2-bit configurations, allowing Turbovec to maintain competitive recall while still achieving superior speed on 4-bit workloads.

Implementation Deep Dive

The following Python examples demonstrate Turbovec’s API, including the filtered search path that triggers block-skip optimizations.


# Basic usage – 4-bit TurboQuant index

from turbovec import TurboQuantIndex
import numpy as np

# 100K random vectors, dim = 1536

vectors = np.random.randn(100_000, 1536).astype(np.float32)

idx = TurboQuantIndex(dim=1536, bit_width=4)
idx.add(vectors)                     # one-time ingest (quantises & packs)

# Simple top-10 search

queries = np.random.randn(10, 1536).astype(np.float32)
scores, slots = idx.search(queries, k=10)

print("Top-10 slots for first query:", slots[0])
print("Corresponding scores:", scores[0])

# Hybrid retrieval – filter with an allowlist (FAISS-style fast-scan path)

from turbovec import IdMapIndex

ids = np.arange(100_000, dtype=np.uint64)      # external stable ids

idx = IdMapIndex(dim=1536, bit_width=4)
idx.add_with_ids(vectors, ids)

# Suppose an upstream BM25 step returns a candidate set of 2,000 ids

candidate_ids = np.random.choice(ids, size=2_000, replace=False)

# Search only within that candidate set

scores, result_ids = idx.search(queries, k=10, allowlist=candidate_ids)

print("Filtered results (ids):", result_ids[0])

Both examples execute the SIMD-accelerated kernels defined in turbovec/src/search.rs. The allowlist parameter triggers the block-skip early-exit path, delivering the performance gains observed in the ARM benchmarks.

Summary

  • Turbovec implements hand-written SIMD kernels (NEON, AVX-512BW, AVX2) that use FAISS-compatible nibble-split LUT layouts.
  • Block-skip early-exit filtering avoids unnecessary computation when searching with candidate lists, providing 10–19% speedups on ARM.
  • Length-renormalization corrects quantization bias in-kernel without performance penalty, improving recall especially for 2-bit indices.
  • 4-bit quantization achieves 16× compression (6KB → 384B per 1536-dim vector), enabling billion-scale search in modest RAM footprints.
  • Benchmarks show consistent leads on Apple Silicon and modest advantages or parity on Intel x86, with raw data available in benchmarks/results/.

Frequently Asked Questions

Why is Turbovec faster than FAISS FastScan on ARM but only ties on x86?

Turbovec provides optimized NEON kernels for AArch64 architectures where FAISS relies on less optimized generic paths. On x86, FAISS’s AVX-512VBMI implementation is highly mature, allowing it to match Turbovec’s AVX-512BW performance in multi-threaded scenarios. The 2-bit x86 case shows FAISS pulling ahead by 8% in single-threaded mode due to specialized VBMI permute instructions that Turbovec’s current AVX-512BW kernel does not utilize.

What is the block-skip early-exit mechanism?

The block-skip mechanism checks an allow-list mask before processing each 32-vector block in search.rs. If block_has_allowed determines no candidates exist in the block, the kernel returns immediately, skipping LUT lookups and SIMD calculations. This is recorded by the BLOCKS_SKIPPED_BY_MASK counter and provides significant acceleration for filtered retrieval workloads where candidate sets are sparse.

How does length-renormalization improve recall?

Length-renormalization multiplies raw quantized inner-products by a pre-computed per-vector scalar (||v|| / ⟨u, x̂⟩) stored in the vec_scales array. This corrects the systematic downward bias introduced by scalar quantization, particularly pronounced in 2-bit configurations. The correction happens in-kernel before heap insertion, requiring no additional memory bandwidth or CPU cycles while improving search accuracy.

Can I use Turbovec as a drop-in replacement for FAISS?

Turbovec provides Python bindings via turbovec-python/src/lib.rs exposing similar index creation and search APIs. However, it requires specific quantization during ingestion (TurboQuant pipeline) rather than supporting FAISS’s full range of index types. For applications using FAISS FastScan with scalar quantization, Turbovec offers a compatible alternative with superior ARM performance, though migration requires re-indexing existing datasets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →