turbovec Bit Width Options: Supported Values and Use Cases
turbovec supports exactly three fixed bit widths—2, 3, and 4 bits per coordinate—where 2-bit maximizes compression, 3-bit balances size and accuracy, and 4-bit delivers the highest recall with SIMD-accelerated performance.
The turbovec library (RyanCodrai/turbovec) is a Rust-based vector search engine that compresses high-dimensional float32 vectors into fixed-bit quantized representations. When constructing a TurboQuantIndex, you must select a bit width that dictates how many distinct levels each coordinate can represent. This choice directly impacts memory footprint, search latency, and result accuracy.
Supported Bit Width Values
turbovec restricts quantization to three specific bit depths, each packing coordinates into a fixed number of unsigned integer levels:
- 2-bit: Allocates 4 possible levels (0–3). This yields the highest compression ratio—approximately 16× reduction compared to 32-bit floats—but introduces the coarsest quantization, which may cause a modest drop in recall.
- 3-bit: Allocates 8 possible levels (0–7). This configuration sits between the extremes, offering a middle ground between compression efficiency and granularity.
- 4-bit: Allocates 16 possible levels (0–15). This provides the finest quantization among supported options, delivering the best recall. On x86 CPUs with AVX-512, the 4-bit kernel actually outperforms FAISS for equivalent operations.
Any other value is rejected at construction time.
Constructor Validation
The bit width constraint is enforced rigidly in the source code. Both TurboQuantIndex::new and TurboQuantIndex::new_lazy—defined in turbovec/src/lib.rs (lines 22–26 and 52–56)—validate the input and return a BitWidthOutOfRange error if the supplied value is not 2, 3, or 4. This validation ensures that downstream SIMD kernels and memory layouts receive only supported quantization densities.
Unit tests in turbovec/tests/from_parts.rs explicitly verify that unsupported widths are rejected, while turbovec/tests/kernel_correctness.rs runs end-to-end search validation against all three legal bit widths to ensure numerical correctness.
Recommended Use Cases
Selecting the correct bit width depends on your hardware constraints and accuracy requirements:
2-bit (Maximum Compression) Use this for memory-constrained deployments such as edge devices or massive corpora hosted on modest RAM. The ~16× compression ratio minimizes footprint, and the recall loss is often acceptable when you can afford larger candidate sets or when a downstream reranker corrects approximate results.
3-bit (Balanced Trade-off) Choose this when you have moderate RAM availability but still need decent recall. It provides more granularity than 2-bit quantization while keeping the memory footprint well below the original float32 size, making it suitable for mid-tier servers.
4-bit (High Recall & Speed) Select this for latency-sensitive queries and RAG pipelines where you require the best possible nearest-neighbor results. Beyond superior accuracy, the 4-bit configuration leverages hand-optimized AVX-512 SIMD intrinsics that beat FAISS performance on modern x86 hardware. On ARM architectures, both 2-bit and 4-bit configurations maintain a 10–19% speed advantage over FAISS, but 4-bit remains the optimal choice when recall is critical.
Implementation Examples
Python
import numpy as np
from turbovec import TurboQuantIndex
# 2-bit index for high compression on memory-constrained hosts
idx2 = TurboQuantIndex(dim=1536, bit_width=2)
idx2.add(np.random.rand(10_000, 1536).astype(np.float32))
scores2, ids2 = idx2.search(np.random.rand(5, 1536).astype(np.float32), k=10)
# 4-bit index for maximum recall and SIMD-accelerated search
idx4 = TurboQuantIndex(dim=1536, bit_width=4)
idx4.add(np.random.rand(10_000, 1536).astype(np.float32))
scores4, ids4 = idx4.search(np.random.rand(5, 1536).astype(np.float32), k=10)
Rust
use turbovec::TurboQuantIndex;
// 2-bit index – maximum compression for edge deployments
let mut idx2 = TurboQuantIndex::new(1536, 2).unwrap();
let vectors2 = vec![0.0_f32; 1536 * 8_000]; // 8k vectors
idx2.add(&vectors2);
let results2 = idx2.search(&vectors2[..1536 * 5], 10);
// 4-bit index – best recall with AVX-512 optimization
let mut idx4 = TurboQuantIndex::new(1536, 4).unwrap();
let vectors4 = vec![0.0_f32; 1536 * 8_000];
idx4.add(&vectors4);
let results4 = idx4.search(&vectors4[..1536 * 5], 10);
In both languages, the bit_width parameter is passed directly to the constructor, and invalid values trigger immediate errors propagated from the underlying Rust validation logic in turbovec/src/lib.rs.
Summary
- Validation: turbovec accepts only bit widths of 2, 3, or 4, enforced by
TurboQuantIndex::newandnew_lazyinturbovec/src/lib.rs. - Compression: 2-bit delivers ~16× size reduction versus float32; 4-bit delivers ~4× reduction but with superior fidelity.
- Performance: 4-bit quantization on x86 with AVX-512 outperforms FAISS, while 2-bit is optimized for minimal memory consumption.
- API Consistency: Both the Rust core and Python bindings (
turbovec-python/src/lib.rs) apply identical validation rules.
Frequently Asked Questions
What happens if I supply an unsupported bit width to turbovec?
The constructor will return a BitWidthOutOfRange error immediately. According to the implementation in turbovec/src/lib.rs (lines 22–26 and 52–56), only values of 2, 3, or 4 are permitted; any other integer causes an early exit before memory allocation occurs.
Which turbovec bit width option provides the fastest search performance?
4-bit quantization provides the fastest search on x86 CPUs supporting AVX-512, where the SIMD kernel outperforms FAISS. On ARM64 architectures, both 2-bit and 4-bit configurations run 10–19% faster than FAISS, but 4-bit maintains higher recall.
How much memory does each bit width option save?
Relative to 32-bit float vectors, 2-bit quantization achieves approximately 16× compression, 3-bit achieves roughly 10.6×, and 4-bit achieves approximately 8× compression. The exact savings depend on vector dimension alignment and padding.
Can I change the bit width of an existing TurboQuantIndex?
No. The bit width is an immutable property set at construction time. If you need to migrate between bit widths, you must create a new TurboQuantIndex instance with the desired setting and re-add your vectors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →