Why Turbovec Outperforms FAISS on ARM but Lags on x86 for 2‑Bit Quantization

Turbovec beats FAISS on ARM by pairing a hand-written NEON kernel with a sequential byte layout, while its x86 AVX2 kernel trails FAISS by 5–8% on 2-bit quantization because FAISS leverages faster AVX-512VBMI instructions.

Understanding turbovec vs FAISS 2-bit quantization performance requires dissecting how each engine maps quantized vectors to SIMD-friendly memory layouts. In the RyanCodrai/turbovec repository, the turbovec/src/pack.rs and turbovec/src/search.rs modules implement architecture-specific kernels that explain the divergence. The ARM path is optimized for NEON's register model, whereas the x86 path targets AVX2 and cannot match FAISS's newer AVX-512VBMI FastScan implementation.

Architecture-Specific Data Layouts in pack.rs

The foundation of the performance gap lies in how bytes are packed before scoring. In turbovec/src/pack.rs, two conditional compilation paths produce completely different on-disk SIMD block layouts.

ARM Sequential Layout

On ARM, the pack_blocked function stores each vector's bytes in simple sequential order. This layout is consumed directly by the NEON kernel with no de-interleaving or permutation overhead.

#[cfg(not(target_arch = "x86_64"))]
fn pack_blocked(
    n: usize,
    n_blocks: usize,
    n_byte_groups: usize,
    blocked_size: usize,
    codes_flat: &[Vec<u8>],
    _perm0: &[usize; 16],
) -> Vec<u8> {
    // stores bytes sequentially, directly consumable by NEON
    // see turbovec/src/pack.rs lines 130-145
    …
}

Because _perm0 is unused, the ARM path avoids the interleaved hi/lo nibble shuffling that FAISS and the x86 path require.

x86 FAISS-Style Interleaved Layout

On x86, Turbovec adopts the same perm0-interleaved pattern used by FAISS to keep AVX2 lanes aligned. The pack_blocked implementation interleaves high and low nibbles as it writes the blocked codes.

#[cfg(target_arch = "x86_64")]
fn pack_blocked(
    n: usize,
    n_blocks: usize,
    n_byte_groups: usize,
    blocked_size: usize,
    codes_flat: &[Vec<u8>],
    perm0: &[usize; 16],
) -> Vec<u8> {
    // interleaves hi/lo nibbles per FAISS layout
    // see turbovec/src/pack.rs lines 78-92
    …
}

This layout is necessary for search_multi_query_avx2, but it adds structural complexity that AVX-512VBMI handles more efficiently in FAISS's native kernels.

SIMD Kernel Comparison

With the data packed, the scoring loop in turbovec/src/search.rs dispatches to hand-written SIMD kernels gated by target architecture.

ARM NEON Kernel

The ARM implementation calls score_4bit_block_neon, which operates on the sequential layout without any intermediate shuffling. The bytes are already in the order NEON expects, so the kernel maximizes throughput on every load.

unsafe fn score_4bit_block_neon(
    blocked_codes: &[u8],
    uint8_luts: &[u8],
    block_offset: usize,
    n_byte_groups: usize,
    scale: f32,
    bias: f32,
    vec_scales: &[f32],
    base_vec: usize,
    n_vectors: usize,
    out: &mut [f32; BLOCK],
) {
    // sequential layout – no interleaving needed
    // see turbovec/src/search.rs lines 46-58
    …
}

On ARM devices, this is the only high-throughput path available in Turbovec, and it outperforms FAISS because FAISS provides only a generic scalar fallback on that architecture.

x86 AVX2 Kernel

The x86 counterpart, search_multi_query_avx2, enables avx2 and fma target features and reads the interleaved layout produced by the x86 pack_blocked function.

#[cfg(target_arch = "x86_64")]
#[target_feature(enable = "avx2", enable = "fma")]
unsafe fn search_multi_query_avx2(
    blocked_codes: &[u8],
    luts: &[&[u8]],
    scales: &[f32],
    biases: &[f32],
    n_byte_groups: usize,
    vec_scales: &[f32],
    n_vectors: usize,
    nq: usize,
    k: usize,
    mask: Option<&[u64]>,
    …
) {
    // uses FAISS-style perm0-interleaved layout
    // see turbovec/src/search.rs lines 61-70
    …
}

While this kernel is highly optimized for AVX2, it sits one generation behind the AVX-512VBMI instructions that FAISS uses for 2-bit accumulation.

The AVX-512VBMI Advantage on x86

FAISS's IndexPQFastScan exploits AVX-512VBMI for 2-bit product quantization. This extension provides byte-level permutation and lookup capabilities that drastically accelerate the nibble-lookup/accumulate inner loop. Turbovec does not currently emit AVX-512 instructions; it stops at AVX2 and FMA. Consequently, for 2-bit quantization on x86, FAISS's newer instruction set yields a ~5–8% speed advantage over Turbovec according to the benchmarks in benchmarks/suite/.

The gap is specific to 2-bit codes because AVX-512VBMI excels at the fine-grained bit manipulation required for 2-bit packed data. In 4-bit workloads, Turbovec's AVX2 kernel and interleaved layout are competitive enough to win by up to ~5%.

Calibration and LUT Precision

Turbovec uses the same 8-bit LUT (LUT256) strategy as FAISS, but it adds a per-vector calibration step involving scale, bias, and vec_scales. This calibration improves recall with minimal runtime cost because the extra arithmetic is folded into the SIMD register operations. On ARM, the NEON pipeline absorbs these operations without stalling. On x86, the overhead is similarly negligible, but it cannot compensate for the absence of AVX-512VBMI in the 2-bit scoring path.

Benchmark Results

The README.md and scripts inside benchmarks/suite/ quantify the architecture-specific trade-offs:

  • ARM: Turbovec is 10–19% faster than FAISS across both 2-bit and 4-bit configurations.
  • x86 (4-bit): Turbovec wins by up to ~5%.
  • x86 (2-bit): Turbovec is ~5–8% slower than FAISS, reflecting the latter's AVX-512VBMI FastScan dominance.

The benchmarks/create_diagrams.py script generates the SVG charts (e.g., arm_speed_mt.svg) that visualize these results.

Summary

  • ARM sequential layout in turbovec/src/pack.rs eliminates shuffle overhead and feeds NEON contiguous bytes.
  • The score_4bit_block_neon kernel in turbovec/src/search.rs is the fastest path on ARM, while FAISS falls back to scalar code.
  • x86 interleaved layout matches FAISS's AVX2 expectations but introduces complexity that AVX-512VBMI handles more efficiently.
  • Turbovec's AVX2/FMA kernel cannot match FAISS's AVX-512VBMI nibble-accumulate loop for 2-bit quantization, creating a ~5–8% deficit.
  • Per-vector calibration (scale, bias, vec_scales) preserves Turbovec's recall without negating its architectural speedups.

Frequently Asked Questions

Why does FAISS lack a fast NEON path on ARM?

FAISS relies on a generic scalar fallback on ARM instead of a hand-optimized NEON kernel. As implemented in RyanCodrai/turbovec, this scalar path cannot match the throughput of score_4bit_block_neon, which is why Turbovec records a 10–19% speed advantage on every ARM configuration tested.

Does Turbovec's x86 lag disappear on CPUs without AVX-512VBMI?

On x86 CPUs that lack AVX-512VBMI, FAISS cannot execute its fastest 2-bit FastScan kernels and must fall back to AVX2 or scalar paths. In those scenarios, Turbovec's search_multi_query_avx2 implementation becomes far more competitive, though the exact margin depends on the specific microarchitecture.

How does the sequential ARM layout differ from the x86 interleaved layout?

In turbovec/src/pack.rs, the ARM pack_blocked stores each vector's bytes in sequential order, letting NEON load contiguous chunks with zero preprocessing. The x86 version permutes high and low nibbles into a FAISS-style interleaved pattern to satisfy AVX2 lane alignment, adding complexity that AVX-512VBMI handles more efficiently in FAISS's native implementation.

Why does Turbovec win on 4-bit x86 but lose on 2-bit x86?

Turbovec's AVX2 kernel and interleaved layout are well-suited to 4-bit quantization, yielding up to a 5% speedup over FAISS. However, 2-bit quantization magnifies the benefit of FAISS's AVX-512VBMI nibble-lookup and accumulate loop, which outpaces Turbovec's AVX2-only path by roughly 5–8%.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →