Performance Difference Between 2-Bit and 4-Bit Compression in TurboVec
TurboVec's 2-bit compression roughly halves single-threaded query latency compared to 4-bit (~2.66 ms vs ~5.34 ms), trading a small 1-2% recall reduction for significantly faster memory-bound vector search.
TurboVec stores vectors in a quantized packed-code format where the bit_width parameter controls the compression ratio and directly influences query latency and recall. In the RyanCodrai/turbovec repository, the core TurboQuantIndex struct validates bit widths between 2 and 4 in turbovec/src/lib.rs, and the choice between these settings creates measurable performance differences. Understanding the performance difference between 2-bit and 4-bit compression in TurboVec is essential for balancing search throughput against result accuracy.
Benchmarking the 2-Bit vs 4-Bit Performance Difference in TurboVec
The benchmarks/results/ directory contains controlled x86 timing tests that isolate the impact of bit width. On single-threaded runs with 3072-dimensional vectors, 2-bit queries execute in roughly half the time of 4-bit queries.
| Benchmark File | Dimension | Threads | Bit Width | Avg. TurboVec Time | Avg. FAISS Time |
|---|---|---|---|---|---|
speed_d3072_2bit_x86_st.json |
3072 | Single | 2 | 2.66 ms | 2.58 ms |
speed_d3072_4bit_x86_st.json |
3072 | Single | 4 | 5.34 ms | 5.47 ms |
speed_d3072_4bit_x86_mt.json |
3072 | Multi | 4 | 1.18 ms | 1.18 ms |
These results confirm that 2-bit compression is the faster format under single-threaded conditions. When multi-threading is enabled, the 4-bit implementation drops to approximately 1.18 ms, though the repository does not provide a corresponding 2-bit multi-threaded baseline for direct comparison.
Why 2-Bit Compression Outperforms 4-Bit in TurboVec
The performance gap stems from how tightly the packed representation couples memory bandwidth to compute. Three factors in the TurboVec source code drive the latency difference.
Memory Traffic
In turbovec/src/pack.rs, the packing logic computes bytes_per_vec = dim * self.bit_width / 8. A 2-bit code consumes exactly half the storage of a 4-bit code for the same dimensionality. During query execution in TurboQuantIndex::search, this halves the volume of data read from RAM.
Cache Friendliness
Smaller packed representations improve CPU cache utilization. When TurboQuantIndex iterates over the vector collection, the 2-bit format allows more vectors to reside in L1 and L2 cache simultaneously, reducing cache-miss penalties compared to the 4-bit layout.
Computational Overhead
The unpack-and-dot-product step scales linearly with bit_width. Halving the bit width from 4 to 2 roughly halves the number of bitwise operations required per vector comparison, lowering overall CPU cycles per query.
Accuracy Trade-Off: Recall Impact of 2-Bit Quantization
Speed comes at a minor accuracy cost. The recall benchmark scripts benchmarks/suite/recall_d3072_2bit.py and benchmarks/suite/recall_d3072_4bit.py demonstrate that moving from 4-bit to 2-bit compression lowers recall by approximately 1-2%, depending on the dataset. This trend is consistent across runs: higher compression yields faster queries with a small reduction in nearest-neighbor accuracy.
Code Example: Configuring 2-Bit and 4-Bit Indexes
You configure the compression level when constructing TurboQuantIndex through the Python bindings. The following snippet creates both index types, adds identical data, and runs search queries that exhibit the performance difference shown in the benchmarks.
import numpy as np
import turbovec
# 2-bit index (faster, slightly lower recall)
idx2 = turbovec.TurboQuantIndex(dim=3072, bit_width=2)
# 4-bit index (slower, higher recall)
idx4 = turbovec.TurboQuantIndex(dim=3072, bit_width=4)
# Add vectors (same data for both indexes)
vectors = np.random.randn(1000, 3072).astype(np.float32)
idx2.add(vectors)
idx4.add(vectors)
# Query – note the ~2× speed difference on the same hardware
query = np.random.randn(1, 3072).astype(np.float32)
neighbors2 = idx2.search(query, k=5) # ~2.6 ms per query
neighbors4 = idx4.search(query, k=5) # ~5.3 ms per query
Summary
- 2-bit compression in TurboVec delivers approximately 2× faster single-threaded query latency than 4-bit (~2.66 ms vs ~5.34 ms for 3072-dimensional vectors).
- Memory bandwidth is the primary driver:
bytes_per_vec = dim * bit_width / 8means 2-bit indexes read half the data per query. - Recall drops by roughly 1-2% when switching from 4-bit to 2-bit, as shown in
benchmarks/suite/recall_d3072_2bit.py. - Multi-threading improves 4-bit performance dramatically (down to ~1.18 ms), but the single-threaded results confirm that 2-bit remains the more lightweight representation.
- Key files governing this behavior are
turbovec/src/lib.rsfor index configuration andturbovec/src/pack.rsfor the packing math.
Frequently Asked Questions
Is 2-bit compression always faster than 4-bit in TurboVec?
Under single-threaded execution, the 2-bit format is consistently faster because it moves less data through the memory hierarchy and executes fewer bitwise operations per vector. The repository does not publish a 2-bit multi-threaded benchmark, but the underlying math in turbovec/src/pack.rs ensures the 2-bit representation remains more cache-friendly regardless of thread count.
How much recall do I lose with 2-bit vs 4-bit compression?
The benchmark scripts in benchmarks/suite/recall_d3072_2bit.py and recall_d3072_4bit.py report a drop of approximately 1-2% when moving from 4-bit to 2-bit compression. The exact figure depends on the dataset and distribution, but the trade-off is consistently small across runs.
Where does TurboVec enforce the bit width setting?
The constructor TurboQuantIndex::new(dim, bit_width) in turbovec/src/lib.rs validates that bit_width falls in the range 2..=4 and stores the value in the index struct. The query routine TurboQuantIndex::search and the packing logic in turbovec/src/pack.rs then use this value to compute bytes_per_vec, directly controlling the storage footprint and speed of each query.
When should I choose 4-bit compression over 2-bit?
Choose 4-bit compression when maximizing recall is more important than raw query throughput, or when you can offset the higher latency with multi-threading. If your workload is single-threaded or memory-bandwidth constrained, 2-bit compression in TurboVec provides substantially faster search with only a minor accuracy penalty.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →