# Turbovec Bit Widths: How 2, 3, and 4-Bit Quantization Impact Compression and Performance

> Explore Turbovec's 2, 3, and 4-bit quantization options to understand their impact on compression ratios and performance. Optimize your vector search with precise bit width selection.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: performance
- Published: 2026-08-22

---

**TLDR: Turbovec supports exactly three bit widths — 2, 3, and 4 bits per vector coordinate — which determine codebook size, packed memory footprint, and recall, giving roughly 16× down to 8× compression with recall ranging from 0.40 to 0.70+ at R@10.**

Turbovec is a Rust-based approximate nearest neighbor (ANN) library that compresses high-dimensional vectors using product quantization with bit width as the core tuning knob. According to the source code in `RyanCodrai/turbovec`, the `TurboQuantIndex` API only accepts bit widths of 2, 3, or 4, and that single number controls everything from codebook size to SIMD packing layout. This article breaks down how each turbovec bit width affects compression, memory traffic, CPU work, and recall — so you can pick the right value for your workload.

## Supported Bit Widths and Validation Logic

The bit width is validated in the constructor `TurboQuantIndex::new`, located in [`turbovec/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs). If you pass any value outside the inclusive range `2..=4`, the library returns `ConstructError::BitWidthOutOfRange` and refuses to build the index.

| Bit width | Valid? | Constructor behavior |
|-----------|--------|----------------------|
| 2 bits    | Yes    | Accepted at [`lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/lib.rs) lines 626–632 |
| 3 bits    | Yes    | Accepted at [`lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/lib.rs) lines 626–632 |
| 4 bits    | Yes    | Accepted at [`lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/lib.rs) lines 626–632 |
| 1, 5+    | No    | Returns `ConstructError::BitWidthOutOfRange` |

```rust
use turbovec::TurboQuantIndex;

// Valid: 2, 3, or 4 bits
let idx = TurboQuantIndex::new(1536, 3).unwrap();

// Invalid: throws ConstructError::BitWidthOutOfRange
let err = TurboQuantIndex::new(1536, 5);
assert!(err.is_err());

```

The same range check is mirrored in the **codebook builder**. The helper function `codebook::codebook(bit_width, dim)` only computes centroids for these three widths, so the codebook — the set of Lloyd-Max quantization centroids used during encoding — is always consistent with the chosen bit width.

## How Bit Width Affects Compression Ratio

The compression ratio is the most immediate consequence of choosing a bit width. Each vector coordinate is packed into `bit_width` bits instead of the original 32 bits of an `f32`.

| Bit width | Bits per coordinate | Compression vs. `f32` | Packed size per 1536-dim vector |
|-----------|--------------------|------------------------|--------------------------------|
| 2 bits    | `2/32 = 1/16`      | ≈ 16× compression       | `1536 × 2 / 8 = 384 bytes` |
| 3 bits    | `3/32 ≈ 1/10.7`    | ≈ 10.7× compression     | `1536 × 3 / 8 = 576 bytes` |
| 4 bits    | `4/32 = 1/8`       | ≈ 8× compression        | `1536 × 4 / 8 = 768 bytes` |

The packed size is computed directly in `TurboQuantIndex::packed` as `dim * self.bit_width / 8` bytes per vector. So a 2‑bit index of a 1536‑dim dataset uses half the memory of its 4‑bit counterpart.

## Codebook Size and Quantization Granularity

The bit width determines how many quantization levels the codebook must contain. The formula is `1 << bit_width` centroids per dimension, generated by `codebook::codebook(bit_width, dim)`.

- **2 bits** → 4 centroids per dimension. Very coarse approximation; each coordinate can only map to one of 4 values.
- **3 bits** → 8 centroids. Medium fidelity.
- **4 bits** → 16 centroids. Finest resolution among the supported widths.

Smaller codebooks mean less storage for the centroids themselves, but the bigger win is in **arithmetic efficiency**. During encoding and search, the SIMD kernel resolves each coordinate against its codebook levels. With 2 bits there are only 4 lookups per coordinate, while 4 bits requires 16 — which directly increases the per-vector compute cost.

## Memory Traffic and Cache Locality

Turbovec packs vectors into a blocked-cache layout for SIMD-friendly scanning. The function `pack::blocked_geometry` computes two critical values based on the bit width:

- `n_byte_groups` — how many byte groups fit a packed row
- `block_size` — the size of each SIMD-blocked row in the cache

As the raw analysis shows, larger bit widths increase `n_byte_groups`, which raises the memory footprint of each cache row. This directly impacts:

| Aspect | 2-bit | 3-bit | 4-bit |
|--------|-------|-------|-------|
| **Packed row size** | `dim × 2 / 8` | `dim × 3 / 8` | `dim × 4 / 8` |
| **RAM bandwidth per scan** | Lowest — fewest bytes per vector | Intermediate | Highest |
| **Cache locality** | Best (more vectors fit in L2/L3) | Moderate | Reduced |
| **Memory traffic** | ~16× less than `f32` | ~10.7× less | ~8× less |

For memory-bandwidth-bound workloads — which is the vast majority of ANN search at scale — reducing bits per coordinate directly boosts query throughput because the system moves fewer bytes from RAM to CPU.

## Recall and Accuracy Trade-Offs

The library ships with empirically calibrated recall floors in [`turbovec/tests/recall_sanity.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/recall_sanity.rs). These are the minimum recall expected at `R@10` for each bit width:

- **2-bit index** — recall floor of **≈ 0.40**
- **3-bit index** — recall floor of **≈ 0.58**
- **4-bit index** — recall floor of **≈ 0.70** (often higher in practice, ≳ 0.70)

These are sanity thresholds, not absolute bounds; real values depend on your dataset (dimensionality, distribution, number of vectors). But the trend is consistent across SIFT-like data: every extra bit buys meaningful recall.

## CPU Work and SIMD Efficiency

The per-vector computational cost scales with the codebook size:

- **2-bit kernels** address only 4 quantization levels — the cheapest path.
- **3-bit kernels** index 8 levels, adding more arithmetic.
- **4-bit kernels** index 16 levels — the most compute-heavy.

That said, the absolute overhead is modest because **SIMD lanes process whole blocks in parallel** regardless of bit width. The bit width changes the number of addressable levels per lane, not the vectorization width. So on CPUs with AVX2 or AVX-512, the 4-bit kernel still runs at near-parity with 2-bit — the dominant factor remains memory bandwidth, not compute.

## Practical Guidance: Which Bit Width Should You Use?

Your choice reduces to a single trade-off between recall and speed.

**Use 4 bits** when you need the highest accuracy and can tolerate a larger index — e.g., a 1536-dim embedding store where recall must be ≥ 0.70 at R@10.

**Use 3 bits** as the middle ground for balanced compression and accuracy, offering ~10× compression with recall around 0.58.

**Use 2 bits** for latency-critical production systems where the index must fit in cache and you accept recall in the 0.40 range. This yields the smallest packed representation and the fastest scan throughput.

## Code Example

```rust
use turbovec::TurboQuantIndex;

// 2-bit index — maximum compression, lower recall
let mut idx2 = TurboQuantIndex::new(1536, 2).unwrap();
idx2.add(&vec![0.0_f32; 1536 * 10]);          // add 10 vectors
let results2 = idx2.search(&vec![0.0_f32; 1536 * 2], 10);
println!("2-bit top-1 id: {}", results2.indices[0]);

// 4-bit index — higher accuracy
let mut idx4 = TurboQuantIndex::new(1536, 4).unwrap();
idx4.add(&vec![0.0_f32; 1536 * 10]);
let results4 = idx4.search(&vec![0.0_f32; 1536 * 2], 10);
println!("4-bit top-1 id: {}", results4.indices[0]);

```

Both snippets compile the same API; only the `bit_width` argument changes. The codebook is auto-generated for the chosen width, and the packed size follows `dim * bit_width / 8` bytes per vector.

## Key Source Files in RyanCodrai/turbovec

| File | Relevance to bit widths |
|------|-------------------------|
| [`turbovec/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs) | Core public API (`new`, `new_lazy`), bit-width validation, and `ConstructError::BitWidthOutOfRange` |
| [`turbovec/src/codebook.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/codebook.rs) | `codebook(bit_width, dim)` — generates the Lloyd-Max codebook with `1 << bit_width` centroids |
| [`turbovec/src/pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/pack.rs) | `blocked_geometry` — computes `n_byte_groups` and block size based on `bit_width`, controlling cache layout |
| [`turbovec/tests/recall_sanity.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/recall_sanity.rs) | Enforces recall floors (2-bit ≥ 0.40, 3-bit ≥ 0.58, 4-bit ≥ 0.70) |

## Summary

- Turbovec supports exactly **2, 3, and 4 bits** per coordinate — anything else triggers `ConstructError::BitWidthOutOfRange`.
- Compression scales from **16× (2-bit)** down to **8× (4-bit)** — packed size is `dim * bit_width / 8` bytes per vector.
- The codebook size is `1 << bit_width` centroids per dim, which drives CPU work and quantization precision.
- Recall floors are **0.40 / 0.58 / 0.70** at R@10 for 2/3/4 bits, respectively.
- Smaller bit widths win on memory bandwidth and cache locality but cost recall; 4-bit is the accuracy pick.

## Frequently Asked Questions

### What bit widths can I use in Turbovec?

Turbovec only accepts bit widths of `2`, `3`, or `4`. The `TurboQuantIndex::new` method validates the input against the range `2..=4` and returns `ConstructError::BitWidthOutOfRange` for any other value, including 1 or 5 and above.

### How does a lower bit width improve performance in Turbovec?

Using 2 bits instead of 4 halves the packed bytes per vector (`dim × 2 / 8` vs `dim × 4 / 8`), which reduces RAM bandwidth during search and improves cache locality. Additionally, 2-bit kernels only address 4 codebook levels per coordinate instead of 16, slightly lowering per-vector CPU work.

### What recall can I expect with a 4-bit Turbovec index?

Approximately 80% at R@10 on typical 1536-dim datasets, per the recall floors in [`turbovec/tests/recall_sanity.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/recall_sanity.rs). Exactly recall depends on your data distribution, but 4-bit consistently outperforms 2-bit (≤0.40) and 3-bit (≈0.58) in the same test suite.

### Does Turbovec support a bit width of 8 or 16?

No. The library intentionally restricts quantization to the low-bit regime of 2, 3, and 4 bits. For higher precision you would use a different quantization scheme, but that is outside current Turbovec's public API — which is validated in [`turbovec/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs).