# Why 4-Bit Quantization Offers Better Performance Consistency Than 2-Bit in TurboVect

> Discover why 4-bit quantization in TurboVect provides superior performance consistency over 2-bit. Learn how improved instruction scheduling and reduced memory latency enhance stability across architectures.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: performance
- Published: 2026-06-10

---

**TurboVect's 4-bit quantization delivers more stable performance than 2-bit because longer SIMD inner loops enable better instruction scheduling, while halving memory bandwidth pressure reduces cache-related latency variance across x86 and ARM architectures.**

TurboVect is a high-performance vector similarity search library that implements scalar quantization for approximate nearest neighbor search. When comparing bit-width configurations across the codebase, **4-bit quantization** demonstrates superior performance consistency compared to **2-bit**, particularly under multi-threaded workloads where instruction throughput and memory access patterns become critical bottlenecks.

## Kernel Length and Instruction Scheduling Limitations

The primary architectural factor favoring 4-bit stability lies in the length of the inner-accumulate loop within the SIMD kernels. In the x86 implementation, 2-bit encoding produces only four possible codes per sub-vector, creating a very tight loop that limits the compiler's ability to apply loop unrolling and optimize SIMD instruction scheduling.

According to the TurboVect source analysis in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs), the 2-bit multi-threaded (MT) benchmarks fall 2–4% behind FAISS because, as the code comments explain, *"the inner accumulate loop is too short for unrolling amortization to match FAISS’s AVX-512 VBMI path"* (lines 61–62). This constraint prevents the processor from fully exploiting instruction-level parallelism, causing performance volatility as the pipeline stalls between short iterations.

## SIMD Dispatch and Codebook Resolution

**4-bit quantization** increases the sub-vector codebook size to 16 centroids, enabling richer per-query lookup tables (LUTs) that operate on groups of 8 coordinates per byte. This larger codebook allows the AVX-512 VBMI path to remain fully vectorized throughout the search kernel, maintaining consistent throughput across different query distributions.

The implementation in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) builds these nibble-wise LUTs dynamically during query execution. With 16 possible values (4 bits), the kernel avoids the frequent branch mispredictions and partial register dependencies that plague the 2-bit path, where the limited centroid count forces more conditional execution paths.

## Memory Bandwidth and Cache Efficiency

A 4-bit vector stores exactly half the bytes of a 2-bit vector for equivalent dimensionality, directly reducing memory traffic during the bulk-copy and scoring phases. This compression advantage manifests as:

- **Reduced cache pressure**: Fewer cache lines needed per vector batch
- **Better prefetcher utilization**: More predictable access patterns during LUT construction
- **Lower latency variance**: Less sensitivity to memory subsystem stalls

The quantizer implementation in [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs) leverages this bandwidth reduction to maintain stable throughput during high-concurrency searches, whereas the 2-bit mode must saturate the memory bus with twice the data volume for equivalent recall targets.

## Calibration Stability with TQ+

TurboVect's **TQ+ calibration** (TurboQuant Plus) computes optimal scaling factors for length renormalization at both bit widths, but the higher-resolution 4-bit codebook reduces quantization noise significantly. This stabilization occurs in [`turbovec/src/codebook.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/codebook.rs), which generates Lloyd-Max optimal boundaries for the Beta distribution.

With 16 centroids instead of 4, the 4-bit representation preserves finer-grained distance information, making the final similarity scores less sensitive to coarse quantization errors. As noted in the repository documentation, this produces more predictable per-vector scaling factors and eliminates the performance cliffs observed in 2-bit mode when handling high-variance vector distributions.

## Performance Verification

The following Python example demonstrates the API differences between bit widths:

```python

# 2-bit index (fastest compression, higher variance in latency)

from turbovec import TurboQuantIndex
import numpy as np

vectors = np.random.rand(100_000, 1536).astype('float32')
idx_2 = TurboQuantIndex(dim=1536, bit_width=2)   # ← 2-bit

idx_2.add(vectors)                               # ingest

scores, ids = idx_2.search(vectors[0], k=10)     # query

print("2-bit top-1 score:", scores[0])

# 4-bit index (more stable performance)

idx_4 = TurboQuantIndex(dim=1536, bit_width=4)   # ← 4-bit

idx_4.add(vectors)
scores, ids = idx_4.search(vectors[0], k=10)
print("4-bit top-1 score:", scores[0])

```

To quantify the consistency difference empirically, measure latency variance across multiple query executions:

```bash

# Measure latency consistency (run several times)

python - <<'PY'
import time, numpy as np
from turbovec import TurboQuantIndex
vecs = np.random.rand(100_000, 1536).astype('float32')
for bw in (2, 4):
    idx = TurboQuantIndex(dim=1536, bit_width=bw)
    idx.add(vecs)
    q = vecs[0]
    times = []
    for _ in range(20):
        t0 = time.perf_counter()
        idx.search(q, k=10)
        times.append(time.perf_counter() - t0)
    print(f"{bw}-bit avg={sum(times)/len(times):.4f}s  std={np.std(times):.6f}s")
PY

```

The 4-bit configuration typically exhibits a lower standard deviation in query latency, confirming the architectural consistency advantages documented in the [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md) (lines 60–62), where 4-bit wins by 1–6% across all configurations while 2-bit MT lags 2–4% behind comparable FAISS implementations.

## Summary

- **Longer inner loops** in 4-bit kernels enable effective SIMD unrolling and instruction scheduling, eliminating the pipeline stalls present in 2-bit's truncated accumulate loops
- **50% reduction in memory bandwidth** compared to 2-bit encoding minimizes cache pressure and memory subsystem variance during bulk vector operations
- **16-centroid codebooks** (vs. 4) allow full AVX-512 VBMI vectorization with richer per-query LUTs, reducing branch mispredictions
- **TQ+ calibration** produces more stable length-renormalization corrections at 4-bit depth due to reduced quantization noise
- **Cross-platform consistency** applies to both x86 and ARM implementations, with 4-bit showing uniform gains while 2-bit exhibits platform-specific slowdowns

## Frequently Asked Questions

### What causes the performance variance in 2-bit quantization?

The variance stems from extremely short inner-accumulate loops—only four iterations per sub-vector—that prevent effective loop unrolling and SIMD instruction scheduling on modern CPUs. As implemented in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs), these micro-loops cannot amortize the overhead of AVX-512 VBMI dispatch, causing 2–4% throughput degradation under multi-threaded workloads compared to FAISS implementations.

### How does TurboVect's 4-bit implementation utilize AVX-512 instructions?

The 4-bit path constructs per-query nibble lookup tables (LUTs) that process 8 coordinates per byte using 16 possible centroids. This granularity keeps the AVX-512 VBMI pipeline fully occupied, as the inner loop length provides sufficient iterations for the compiler to apply vectorization optimizations without the amortization penalties that affect 2-bit mode.

### Is 4-bit quantization always better than 2-bit for every workload?

While 4-bit offers superior consistency and recall accuracy, 2-bit provides maximum compression (32× vs. 16×) and may suffice for memory-constrained environments where query latency variance is acceptable. The [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs) implementation supports both widths via the same TQ+ calibration pipeline, allowing users to trade compression ratio for performance stability based on deployment constraints.

### Which source files control the quantization and search behavior differences?

The key architectural distinctions reside in three core files: [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs) handles the rotation and Lloyd-Max quantization, [`turbovec/src/codebook.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/codebook.rs) generates the optimal Beta-distribution boundaries for both bit widths, and [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) contains the divergent SIMD kernel implementations where the inner-loop length constraints manifest as performance differences.