Turbovec Bit Widths: How 2, 3, and 4-Bit Quantization Impact Compression and Performance

TLDR: Turbovec supports exactly three bit widths — 2, 3, and 4 bits per vector coordinate — which determine codebook size, packed memory footprint, and recall, giving roughly 16× down to 8× compression with recall ranging from 0.40 to 0.70+ at R@10.

Turbovec is a Rust-based approximate nearest neighbor (ANN) library that compresses high-dimensional vectors using product quantization with bit width as the core tuning knob. According to the source code in RyanCodrai/turbovec, the TurboQuantIndex API only accepts bit widths of 2, 3, or 4, and that single number controls everything from codebook size to SIMD packing layout. This article breaks down how each turbovec bit width affects compression, memory traffic, CPU work, and recall — so you can pick the right value for your workload.

Supported Bit Widths and Validation Logic

The bit width is validated in the constructor TurboQuantIndex::new, located in turbovec/src/lib.rs. If you pass any value outside the inclusive range 2..=4, the library returns ConstructError::BitWidthOutOfRange and refuses to build the index.

Bit width Valid? Constructor behavior
2 bits Yes Accepted at lib.rs lines 626–632
3 bits Yes Accepted at lib.rs lines 626–632
4 bits Yes Accepted at lib.rs lines 626–632
1, 5+ No Returns ConstructError::BitWidthOutOfRange
use turbovec::TurboQuantIndex;

// Valid: 2, 3, or 4 bits
let idx = TurboQuantIndex::new(1536, 3).unwrap();

// Invalid: throws ConstructError::BitWidthOutOfRange
let err = TurboQuantIndex::new(1536, 5);
assert!(err.is_err());

The same range check is mirrored in the codebook builder. The helper function codebook::codebook(bit_width, dim) only computes centroids for these three widths, so the codebook — the set of Lloyd-Max quantization centroids used during encoding — is always consistent with the chosen bit width.

How Bit Width Affects Compression Ratio

The compression ratio is the most immediate consequence of choosing a bit width. Each vector coordinate is packed into bit_width bits instead of the original 32 bits of an f32.

Bit width Bits per coordinate Compression vs. f32 Packed size per 1536-dim vector
2 bits 2/32 = 1/16 ≈ 16× compression 1536 × 2 / 8 = 384 bytes
3 bits 3/32 ≈ 1/10.7 ≈ 10.7× compression 1536 × 3 / 8 = 576 bytes
4 bits 4/32 = 1/8 ≈ 8× compression 1536 × 4 / 8 = 768 bytes

The packed size is computed directly in TurboQuantIndex::packed as dim * self.bit_width / 8 bytes per vector. So a 2‑bit index of a 1536‑dim dataset uses half the memory of its 4‑bit counterpart.

Codebook Size and Quantization Granularity

The bit width determines how many quantization levels the codebook must contain. The formula is 1 << bit_width centroids per dimension, generated by codebook::codebook(bit_width, dim).

  • 2 bits → 4 centroids per dimension. Very coarse approximation; each coordinate can only map to one of 4 values.
  • 3 bits → 8 centroids. Medium fidelity.
  • 4 bits → 16 centroids. Finest resolution among the supported widths.

Smaller codebooks mean less storage for the centroids themselves, but the bigger win is in arithmetic efficiency. During encoding and search, the SIMD kernel resolves each coordinate against its codebook levels. With 2 bits there are only 4 lookups per coordinate, while 4 bits requires 16 — which directly increases the per-vector compute cost.

Memory Traffic and Cache Locality

Turbovec packs vectors into a blocked-cache layout for SIMD-friendly scanning. The function pack::blocked_geometry computes two critical values based on the bit width:

  • n_byte_groups — how many byte groups fit a packed row
  • block_size — the size of each SIMD-blocked row in the cache

As the raw analysis shows, larger bit widths increase n_byte_groups, which raises the memory footprint of each cache row. This directly impacts:

Aspect 2-bit 3-bit 4-bit
Packed row size dim × 2 / 8 dim × 3 / 8 dim × 4 / 8
RAM bandwidth per scan Lowest — fewest bytes per vector Intermediate Highest
Cache locality Best (more vectors fit in L2/L3) Moderate Reduced
Memory traffic ~16× less than f32 ~10.7× less ~8× less

For memory-bandwidth-bound workloads — which is the vast majority of ANN search at scale — reducing bits per coordinate directly boosts query throughput because the system moves fewer bytes from RAM to CPU.

Recall and Accuracy Trade-Offs

The library ships with empirically calibrated recall floors in turbovec/tests/recall_sanity.rs. These are the minimum recall expected at R@10 for each bit width:

  • 2-bit index — recall floor of ≈ 0.40
  • 3-bit index — recall floor of ≈ 0.58
  • 4-bit index — recall floor of ≈ 0.70 (often higher in practice, ≳ 0.70)

These are sanity thresholds, not absolute bounds; real values depend on your dataset (dimensionality, distribution, number of vectors). But the trend is consistent across SIFT-like data: every extra bit buys meaningful recall.

CPU Work and SIMD Efficiency

The per-vector computational cost scales with the codebook size:

  • 2-bit kernels address only 4 quantization levels — the cheapest path.
  • 3-bit kernels index 8 levels, adding more arithmetic.
  • 4-bit kernels index 16 levels — the most compute-heavy.

That said, the absolute overhead is modest because SIMD lanes process whole blocks in parallel regardless of bit width. The bit width changes the number of addressable levels per lane, not the vectorization width. So on CPUs with AVX2 or AVX-512, the 4-bit kernel still runs at near-parity with 2-bit — the dominant factor remains memory bandwidth, not compute.

Practical Guidance: Which Bit Width Should You Use?

Your choice reduces to a single trade-off between recall and speed.

Use 4 bits when you need the highest accuracy and can tolerate a larger index — e.g., a 1536-dim embedding store where recall must be ≥ 0.70 at R@10.

Use 3 bits as the middle ground for balanced compression and accuracy, offering ~10× compression with recall around 0.58.

Use 2 bits for latency-critical production systems where the index must fit in cache and you accept recall in the 0.40 range. This yields the smallest packed representation and the fastest scan throughput.

Code Example

use turbovec::TurboQuantIndex;

// 2-bit index — maximum compression, lower recall
let mut idx2 = TurboQuantIndex::new(1536, 2).unwrap();
idx2.add(&vec![0.0_f32; 1536 * 10]);          // add 10 vectors
let results2 = idx2.search(&vec![0.0_f32; 1536 * 2], 10);
println!("2-bit top-1 id: {}", results2.indices[0]);

// 4-bit index — higher accuracy
let mut idx4 = TurboQuantIndex::new(1536, 4).unwrap();
idx4.add(&vec![0.0_f32; 1536 * 10]);
let results4 = idx4.search(&vec![0.0_f32; 1536 * 2], 10);
println!("4-bit top-1 id: {}", results4.indices[0]);

Both snippets compile the same API; only the bit_width argument changes. The codebook is auto-generated for the chosen width, and the packed size follows dim * bit_width / 8 bytes per vector.

Key Source Files in RyanCodrai/turbovec

File Relevance to bit widths
turbovec/src/lib.rs Core public API (new, new_lazy), bit-width validation, and ConstructError::BitWidthOutOfRange
turbovec/src/codebook.rs codebook(bit_width, dim) — generates the Lloyd-Max codebook with 1 << bit_width centroids
turbovec/src/pack.rs blocked_geometry — computes n_byte_groups and block size based on bit_width, controlling cache layout
turbovec/tests/recall_sanity.rs Enforces recall floors (2-bit ≥ 0.40, 3-bit ≥ 0.58, 4-bit ≥ 0.70)

Summary

  • Turbovec supports exactly 2, 3, and 4 bits per coordinate — anything else triggers ConstructError::BitWidthOutOfRange.
  • Compression scales from 16× (2-bit) down to 8× (4-bit) — packed size is dim * bit_width / 8 bytes per vector.
  • The codebook size is 1 << bit_width centroids per dim, which drives CPU work and quantization precision.
  • Recall floors are 0.40 / 0.58 / 0.70 at R@10 for 2/3/4 bits, respectively.
  • Smaller bit widths win on memory bandwidth and cache locality but cost recall; 4-bit is the accuracy pick.

Frequently Asked Questions

What bit widths can I use in Turbovec?

Turbovec only accepts bit widths of 2, 3, or 4. The TurboQuantIndex::new method validates the input against the range 2..=4 and returns ConstructError::BitWidthOutOfRange for any other value, including 1 or 5 and above.

How does a lower bit width improve performance in Turbovec?

Using 2 bits instead of 4 halves the packed bytes per vector (dim × 2 / 8 vs dim × 4 / 8), which reduces RAM bandwidth during search and improves cache locality. Additionally, 2-bit kernels only address 4 codebook levels per coordinate instead of 16, slightly lowering per-vector CPU work.

What recall can I expect with a 4-bit Turbovec index?

Approximately 80% at R@10 on typical 1536-dim datasets, per the recall floors in turbovec/tests/recall_sanity.rs. Exactly recall depends on your data distribution, but 4-bit consistently outperforms 2-bit (≤0.40) and 3-bit (≈0.58) in the same test suite.

Does Turbovec support a bit width of 8 or 16?

No. The library intentionally restricts quantization to the low-bit regime of 2, 3, and 4 bits. For higher precision you would use a different quantization scheme, but that is outside current Turbovec's public API — which is validated in turbovec/src/lib.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →