# Why Turbovec Outperforms FAISS on ARM but Lags on x86 for 2‑Bit Quantization

> Discover why Turbovec excels on ARM but trails FAISS on x86 for 2-bit quantization. Explore NEON vs AVX-512 performance differences and optimize your vector search.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: performance
- Published: 2026-07-27

---

**Turbovec beats FAISS on ARM by pairing a hand-written NEON kernel with a sequential byte layout, while its x86 AVX2 kernel trails FAISS by 5–8% on 2-bit quantization because FAISS leverages faster AVX-512VBMI instructions.**

Understanding turbovec vs FAISS 2-bit quantization performance requires dissecting how each engine maps quantized vectors to SIMD-friendly memory layouts. In the `RyanCodrai/turbovec` repository, the [`turbovec/src/pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/pack.rs) and [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) modules implement architecture-specific kernels that explain the divergence. The ARM path is optimized for NEON's register model, whereas the x86 path targets AVX2 and cannot match FAISS's newer AVX-512VBMI FastScan implementation.

## Architecture-Specific Data Layouts in [`pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/pack.rs)

The foundation of the performance gap lies in how bytes are packed before scoring. In [`turbovec/src/pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/pack.rs), two conditional compilation paths produce completely different on-disk SIMD block layouts.

### ARM Sequential Layout

On ARM, the `pack_blocked` function stores each vector's bytes in simple sequential order. This layout is consumed directly by the NEON kernel with no de-interleaving or permutation overhead.

```rust
#[cfg(not(target_arch = "x86_64"))]
fn pack_blocked(
    n: usize,
    n_blocks: usize,
    n_byte_groups: usize,
    blocked_size: usize,
    codes_flat: &[Vec<u8>],
    _perm0: &[usize; 16],
) -> Vec<u8> {
    // stores bytes sequentially, directly consumable by NEON
    // see turbovec/src/pack.rs lines 130-145
    …
}

```

Because `_perm0` is unused, the ARM path avoids the interleaved hi/lo nibble shuffling that FAISS and the x86 path require.

### x86 FAISS-Style Interleaved Layout

On x86, Turbovec adopts the same perm0-interleaved pattern used by FAISS to keep AVX2 lanes aligned. The `pack_blocked` implementation interleaves high and low nibbles as it writes the blocked codes.

```rust
#[cfg(target_arch = "x86_64")]
fn pack_blocked(
    n: usize,
    n_blocks: usize,
    n_byte_groups: usize,
    blocked_size: usize,
    codes_flat: &[Vec<u8>],
    perm0: &[usize; 16],
) -> Vec<u8> {
    // interleaves hi/lo nibbles per FAISS layout
    // see turbovec/src/pack.rs lines 78-92
    …
}

```

This layout is necessary for `search_multi_query_avx2`, but it adds structural complexity that AVX-512VBMI handles more efficiently in FAISS's native kernels.

## SIMD Kernel Comparison

With the data packed, the scoring loop in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) dispatches to hand-written SIMD kernels gated by target architecture.

### ARM NEON Kernel

The ARM implementation calls `score_4bit_block_neon`, which operates on the sequential layout without any intermediate shuffling. The bytes are already in the order NEON expects, so the kernel maximizes throughput on every load.

```rust
unsafe fn score_4bit_block_neon(
    blocked_codes: &[u8],
    uint8_luts: &[u8],
    block_offset: usize,
    n_byte_groups: usize,
    scale: f32,
    bias: f32,
    vec_scales: &[f32],
    base_vec: usize,
    n_vectors: usize,
    out: &mut [f32; BLOCK],
) {
    // sequential layout – no interleaving needed
    // see turbovec/src/search.rs lines 46-58
    …
}

```

On ARM devices, this is the only high-throughput path available in Turbovec, and it outperforms FAISS because FAISS provides only a generic scalar fallback on that architecture.

### x86 AVX2 Kernel

The x86 counterpart, `search_multi_query_avx2`, enables `avx2` and `fma` target features and reads the interleaved layout produced by the x86 `pack_blocked` function.

```rust
#[cfg(target_arch = "x86_64")]
#[target_feature(enable = "avx2", enable = "fma")]
unsafe fn search_multi_query_avx2(
    blocked_codes: &[u8],
    luts: &[&[u8]],
    scales: &[f32],
    biases: &[f32],
    n_byte_groups: usize,
    vec_scales: &[f32],
    n_vectors: usize,
    nq: usize,
    k: usize,
    mask: Option<&[u64]>,
    …
) {
    // uses FAISS-style perm0-interleaved layout
    // see turbovec/src/search.rs lines 61-70
    …
}

```

While this kernel is highly optimized for AVX2, it sits one generation behind the AVX-512VBMI instructions that FAISS uses for 2-bit accumulation.

## The AVX-512VBMI Advantage on x86

FAISS's `IndexPQFastScan` exploits **AVX-512VBMI** for 2-bit product quantization. This extension provides byte-level permutation and lookup capabilities that drastically accelerate the nibble-lookup/accumulate inner loop. Turbovec does not currently emit AVX-512 instructions; it stops at AVX2 and FMA. Consequently, for 2-bit quantization on x86, FAISS's newer instruction set yields a **~5–8% speed advantage** over Turbovec according to the benchmarks in `benchmarks/suite/`.

The gap is specific to 2-bit codes because AVX-512VBMI excels at the fine-grained bit manipulation required for 2-bit packed data. In 4-bit workloads, Turbovec's AVX2 kernel and interleaved layout are competitive enough to win by up to ~5%.

## Calibration and LUT Precision

Turbovec uses the same **8-bit LUT (LUT256)** strategy as FAISS, but it adds a per-vector calibration step involving `scale`, `bias`, and `vec_scales`. This calibration improves recall with minimal runtime cost because the extra arithmetic is folded into the SIMD register operations. On ARM, the NEON pipeline absorbs these operations without stalling. On x86, the overhead is similarly negligible, but it cannot compensate for the absence of AVX-512VBMI in the 2-bit scoring path.

## Benchmark Results

The [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md) and scripts inside `benchmarks/suite/` quantify the architecture-specific trade-offs:

- **ARM**: Turbovec is **10–19% faster** than FAISS across both 2-bit and 4-bit configurations.
- **x86 (4-bit)**: Turbovec wins by up to **~5%**.
- **x86 (2-bit)**: Turbovec is **~5–8% slower** than FAISS, reflecting the latter's AVX-512VBMI FastScan dominance.

The [`benchmarks/create_diagrams.py`](https://github.com/RyanCodrai/turbovec/blob/main/benchmarks/create_diagrams.py) script generates the SVG charts (e.g., `arm_speed_mt.svg`) that visualize these results.

## Summary

- **ARM sequential layout** in [`turbovec/src/pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/pack.rs) eliminates shuffle overhead and feeds NEON contiguous bytes.
- The **`score_4bit_block_neon`** kernel in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) is the fastest path on ARM, while FAISS falls back to scalar code.
- **x86 interleaved layout** matches FAISS's AVX2 expectations but introduces complexity that AVX-512VBMI handles more efficiently.
- Turbovec's **AVX2/FMA** kernel cannot match FAISS's **AVX-512VBMI** nibble-accumulate loop for 2-bit quantization, creating a **~5–8% deficit**.
- Per-vector **calibration** (`scale`, `bias`, `vec_scales`) preserves Turbovec's recall without negating its architectural speedups.

## Frequently Asked Questions

### Why does FAISS lack a fast NEON path on ARM?

FAISS relies on a generic scalar fallback on ARM instead of a hand-optimized NEON kernel. As implemented in `RyanCodrai/turbovec`, this scalar path cannot match the throughput of `score_4bit_block_neon`, which is why Turbovec records a **10–19% speed advantage** on every ARM configuration tested.

### Does Turbovec's x86 lag disappear on CPUs without AVX-512VBMI?

On x86 CPUs that lack AVX-512VBMI, FAISS cannot execute its fastest 2-bit FastScan kernels and must fall back to AVX2 or scalar paths. In those scenarios, Turbovec's `search_multi_query_avx2` implementation becomes far more competitive, though the exact margin depends on the specific microarchitecture.

### How does the sequential ARM layout differ from the x86 interleaved layout?

In [`turbovec/src/pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/pack.rs), the ARM `pack_blocked` stores each vector's bytes in sequential order, letting NEON load contiguous chunks with zero preprocessing. The x86 version permutes high and low nibbles into a FAISS-style interleaved pattern to satisfy AVX2 lane alignment, adding complexity that AVX-512VBMI handles more efficiently in FAISS's native implementation.

### Why does Turbovec win on 4-bit x86 but lose on 2-bit x86?

Turbovec's AVX2 kernel and interleaved layout are well-suited to 4-bit quantization, yielding up to a **5% speedup** over FAISS. However, 2-bit quantization magnifies the benefit of FAISS's AVX-512VBMI nibble-lookup and accumulate loop, which outpaces Turbovec's AVX2-only path by roughly **5–8%**.