Why 4-Bit Quantization Offers Better Performance Consistency Than 2-Bit in TurboVect

TurboVect's 4-bit quantization delivers more stable performance than 2-bit because longer SIMD inner loops enable better instruction scheduling, while halving memory bandwidth pressure reduces cache-related latency variance across x86 and ARM architectures.

TurboVect is a high-performance vector similarity search library that implements scalar quantization for approximate nearest neighbor search. When comparing bit-width configurations across the codebase, 4-bit quantization demonstrates superior performance consistency compared to 2-bit, particularly under multi-threaded workloads where instruction throughput and memory access patterns become critical bottlenecks.

Kernel Length and Instruction Scheduling Limitations

The primary architectural factor favoring 4-bit stability lies in the length of the inner-accumulate loop within the SIMD kernels. In the x86 implementation, 2-bit encoding produces only four possible codes per sub-vector, creating a very tight loop that limits the compiler's ability to apply loop unrolling and optimize SIMD instruction scheduling.

According to the TurboVect source analysis in turbovec/src/search.rs, the 2-bit multi-threaded (MT) benchmarks fall 2–4% behind FAISS because, as the code comments explain, "the inner accumulate loop is too short for unrolling amortization to match FAISS’s AVX-512 VBMI path" (lines 61–62). This constraint prevents the processor from fully exploiting instruction-level parallelism, causing performance volatility as the pipeline stalls between short iterations.

SIMD Dispatch and Codebook Resolution

4-bit quantization increases the sub-vector codebook size to 16 centroids, enabling richer per-query lookup tables (LUTs) that operate on groups of 8 coordinates per byte. This larger codebook allows the AVX-512 VBMI path to remain fully vectorized throughout the search kernel, maintaining consistent throughput across different query distributions.

The implementation in turbovec/src/search.rs builds these nibble-wise LUTs dynamically during query execution. With 16 possible values (4 bits), the kernel avoids the frequent branch mispredictions and partial register dependencies that plague the 2-bit path, where the limited centroid count forces more conditional execution paths.

Memory Bandwidth and Cache Efficiency

A 4-bit vector stores exactly half the bytes of a 2-bit vector for equivalent dimensionality, directly reducing memory traffic during the bulk-copy and scoring phases. This compression advantage manifests as:

  • Reduced cache pressure: Fewer cache lines needed per vector batch
  • Better prefetcher utilization: More predictable access patterns during LUT construction
  • Lower latency variance: Less sensitivity to memory subsystem stalls

The quantizer implementation in turbovec/src/encode.rs leverages this bandwidth reduction to maintain stable throughput during high-concurrency searches, whereas the 2-bit mode must saturate the memory bus with twice the data volume for equivalent recall targets.

Calibration Stability with TQ+

TurboVect's TQ+ calibration (TurboQuant Plus) computes optimal scaling factors for length renormalization at both bit widths, but the higher-resolution 4-bit codebook reduces quantization noise significantly. This stabilization occurs in turbovec/src/codebook.rs, which generates Lloyd-Max optimal boundaries for the Beta distribution.

With 16 centroids instead of 4, the 4-bit representation preserves finer-grained distance information, making the final similarity scores less sensitive to coarse quantization errors. As noted in the repository documentation, this produces more predictable per-vector scaling factors and eliminates the performance cliffs observed in 2-bit mode when handling high-variance vector distributions.

Performance Verification

The following Python example demonstrates the API differences between bit widths:


# 2-bit index (fastest compression, higher variance in latency)

from turbovec import TurboQuantIndex
import numpy as np

vectors = np.random.rand(100_000, 1536).astype('float32')
idx_2 = TurboQuantIndex(dim=1536, bit_width=2)   # ← 2-bit

idx_2.add(vectors)                               # ingest

scores, ids = idx_2.search(vectors[0], k=10)     # query

print("2-bit top-1 score:", scores[0])

# 4-bit index (more stable performance)

idx_4 = TurboQuantIndex(dim=1536, bit_width=4)   # ← 4-bit

idx_4.add(vectors)
scores, ids = idx_4.search(vectors[0], k=10)
print("4-bit top-1 score:", scores[0])

To quantify the consistency difference empirically, measure latency variance across multiple query executions:


# Measure latency consistency (run several times)

python - <<'PY'
import time, numpy as np
from turbovec import TurboQuantIndex
vecs = np.random.rand(100_000, 1536).astype('float32')
for bw in (2, 4):
    idx = TurboQuantIndex(dim=1536, bit_width=bw)
    idx.add(vecs)
    q = vecs[0]
    times = []
    for _ in range(20):
        t0 = time.perf_counter()
        idx.search(q, k=10)
        times.append(time.perf_counter() - t0)
    print(f"{bw}-bit avg={sum(times)/len(times):.4f}s  std={np.std(times):.6f}s")
PY

The 4-bit configuration typically exhibits a lower standard deviation in query latency, confirming the architectural consistency advantages documented in the README.md (lines 60–62), where 4-bit wins by 1–6% across all configurations while 2-bit MT lags 2–4% behind comparable FAISS implementations.

Summary

  • Longer inner loops in 4-bit kernels enable effective SIMD unrolling and instruction scheduling, eliminating the pipeline stalls present in 2-bit's truncated accumulate loops
  • 50% reduction in memory bandwidth compared to 2-bit encoding minimizes cache pressure and memory subsystem variance during bulk vector operations
  • 16-centroid codebooks (vs. 4) allow full AVX-512 VBMI vectorization with richer per-query LUTs, reducing branch mispredictions
  • TQ+ calibration produces more stable length-renormalization corrections at 4-bit depth due to reduced quantization noise
  • Cross-platform consistency applies to both x86 and ARM implementations, with 4-bit showing uniform gains while 2-bit exhibits platform-specific slowdowns

Frequently Asked Questions

What causes the performance variance in 2-bit quantization?

The variance stems from extremely short inner-accumulate loops—only four iterations per sub-vector—that prevent effective loop unrolling and SIMD instruction scheduling on modern CPUs. As implemented in turbovec/src/search.rs, these micro-loops cannot amortize the overhead of AVX-512 VBMI dispatch, causing 2–4% throughput degradation under multi-threaded workloads compared to FAISS implementations.

How does TurboVect's 4-bit implementation utilize AVX-512 instructions?

The 4-bit path constructs per-query nibble lookup tables (LUTs) that process 8 coordinates per byte using 16 possible centroids. This granularity keeps the AVX-512 VBMI pipeline fully occupied, as the inner loop length provides sufficient iterations for the compiler to apply vectorization optimizations without the amortization penalties that affect 2-bit mode.

Is 4-bit quantization always better than 2-bit for every workload?

While 4-bit offers superior consistency and recall accuracy, 2-bit provides maximum compression (32× vs. 16×) and may suffice for memory-constrained environments where query latency variance is acceptable. The turbovec/src/encode.rs implementation supports both widths via the same TQ+ calibration pipeline, allowing users to trade compression ratio for performance stability based on deployment constraints.

Which source files control the quantization and search behavior differences?

The key architectural distinctions reside in three core files: turbovec/src/encode.rs handles the rotation and Lloyd-Max quantization, turbovec/src/codebook.rs generates the optimal Beta-distribution boundaries for both bit widths, and turbovec/src/search.rs contains the divergent SIMD kernel implementations where the inner-loop length constraints manifest as performance differences.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →