Expected Memory Compression Ratios with TurboVec: 2‑Bit vs 4‑Bit Quantization
TurboVec achieves up to 16× memory compression with 2‑bit quantization and 8× compression with 4‑bit quantization, with real‑world workloads typically seeing approximately 7.7× overall reduction after accounting for per‑vector metadata.
TurboVec is a high‑performance Rust vector index that implements Google Research’s TurboQuant algorithm to store high‑dimensional embeddings in a highly compact binary format. For engineers capacity‑planning semantic search or RAG infrastructure, understanding the expected memory compression ratios with TurboVec is essential for accurately provisioning RAM and estimating storage costs.
How TurboVec Compresses Vector Data
TurboVec reduces memory footprint by quantizing each coordinate to a small number of bits and tightly packing those bits into a SIMD‑friendly layout. The process begins in src/encode.rs, which implements per‑coordinate calibration (TQ+) and Lloyd‑Max scalar quantization, and concludes in src/pack.rs, where quantized values are rearranged into a bit‑plane → SIMD‑blocked layout that yields the final compressed byte stream.
This quantization pipeline replaces the standard 32‑bit float per dimension with either 2‑bit or 4‑bit representations, resulting in theoretical compression factors of 16× and 8× respectively.
Compression Ratios by Bit Width
TurboVec supports configurable bit widths that directly determine the compression factor. For a standard 1536‑dimensional embedding (common in OpenAI models), the memory savings are calculated as follows:
2‑Bit Quantization (16× Theoretical Reduction)
Configuring bit_width=2 reduces each dimension from 4 bytes to 0.5 bytes (2 bits). A 1536‑dimensional vector shrinks from 6,144 bytes to 384 bytes, yielding a 16× compression ratio.
According to the repository’s README.md, this represents the maximum theoretical compression available in the current implementation.
4‑Bit Quantization (8× Theoretical Reduction)
Configuring bit_width=4 stores each coordinate in 4 bits (0.5 bytes). The same 1536‑dimensional vector compresses from 6,144 bytes to 768 bytes, achieving an 8× compression ratio.
Real‑World Corpus‑Level Compression
While theoretical ratios apply to pure vector data, production deployments must account for auxiliary metadata. TurboVec stores a single float per vector for length renormalization and additional bookkeeping structures.
For a realistic workload of 10 million documents with 1536‑dimensional embeddings, the raw float32 data consumes approximately 31 GB. After compression, the repository notes this fits into roughly 4 GB, representing an overall ≈ 7.7× reduction. This demonstrates that while individual vectors compress at 16× or 8×, corpus‑level efficiency typically falls slightly below the theoretical maximum due to fixed per‑vector overhead.
Measuring Compression Programmatically
You can verify these ratios using the public API exposed in src/lib.rs. The TurboQuantIndex class accepts a bit_width parameter that controls the quantization level.
Python Example
The following script creates a 2‑bit index and estimates memory usage:
import numpy as np
from turbovec import TurboQuantIndex
# 1 M vectors, 1536 dim, random float32 data
vectors = np.random.rand(1_000_000, 1536).astype(np.float32)
# 2‑bit quantisation (bit_width=2) → ~16× compression
idx = TurboQuantIndex(dim=1536, bit_width=2)
idx.add(vectors)
# Rough memory usage (indices are stored in‑memory)
print("Num vectors:", idx.len())
print("Estimated RAM (bytes):", idx.len() * 384) # 384 B per vector at 2‑bit
Changing bit_width=4 will allocate 768 bytes per vector, reflecting the 8× compression mode.
Rust Example
In Rust, instantiate TurboQuantIndex with the desired bit width:
use turbovec::TurboQuantIndex;
fn main() -> Result<(), turbovec::error::Error> {
// Assume `vectors` is a `Vec<Vec<f32>>` with shape (n, 1536)
let mut index = TurboQuantIndex::new(1536, 4); // 4‑bit → 8× compression
index.add(&vectors)?;
println!("Vectors stored: {}", index.len());
Ok(())
}
The internal quantization pipeline invoked by these examples is implemented across src/encode.rs (calibration and quantization) and src/pack.rs (bit packing).
Summary
- 2‑bit quantization yields approximately 16× compression (384 bytes per 1536‑dimensional vector).
- 4‑bit quantization yields approximately 8× compression (768 bytes per 1536‑dimensional vector).
- Production workloads typically observe ≈ 7.7× overall compression after accounting for metadata and renormalization scalars.
- Core implementation resides in
src/encode.rs(quantization logic) andsrc/pack.rs(SIMD packing). - Benchmark your own datasets using
benchmarks/suite/compression.py.
Frequently Asked Questions
What is the maximum memory compression TurboVec can achieve?
The theoretical maximum is 16× when using 2‑bit quantization on 32‑bit float vectors. This reduces storage from 4 bytes per dimension to 0.5 bytes per dimension. However, actual savings in a production index will be slightly lower due to the per‑vector length renormalization scalar and index metadata.
How does metadata affect the overall compression ratio?
While the raw vector data compresses at the full theoretical factor (16× or 8×), TurboVec stores an additional float per vector for length renormalization plus bookkeeping structures. For a 10‑million‑document corpus, this reduces the effective compression from 16× to approximately 7.7× overall.
Where is the quantization logic implemented in the source code?
The quantization pipeline is split across two key files. src/encode.rs handles per‑coordinate calibration (TQ+) and Lloyd‑Max scalar quantization, while src/pack.rs manages the final bit‑plane to SIMD‑blocked layout conversion that produces the compressed byte representation.
Can I benchmark compression ratios on my own dataset?
Yes. The repository includes benchmarks/suite/compression.py, which calculates the exact compression ratio for a given dataset. Additionally, you can instantiate TurboQuantIndex with bit_width=2 or bit_width=4 in Python or Rust and compare the in‑memory size against the original float32 array size.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →