TQ+ Per-Coordinate Calibration in turbovec: How It Improves Recall
TQ+ per-coordinate calibration learns a tiny shift and scale for every dimension to align rotated embeddings with a theoretical Beta distribution, eliminating quantizer mis-fit and boosting recall by up to 1.4 percentage points at R@1.
TurboQuant, the core algorithm behind the turbovec vector-search library by RyanCodrai, relies on a single random rotation to make every coordinate follow the same theoretical Beta distribution. Real-world embeddings are anisotropic, meaning some dimensions deviate from this ideal shape and cause the global Lloyd-Max codebook to mis-fit. TQ+ per-coordinate calibration solves this by learning per-dimension corrections once during the first batch, then freezing them for all future indexing and querying.
The Problem: Anisotropic Embeddings and Quantizer Mis-Fit
After normalization and random rotation, turbovec expects all coordinates to match a canonical Beta marginal. In turbovec/src/encode.rs, the rotation step (lines 68-84) projects vectors into this theoretical space. Real-world data, however, contains dimensions that are systematically wider or narrower than the Beta ideal. When the global scalar codebook—pre-computed for the perfect Beta shape—is applied to these shifted coordinates, quantization error increases and inner-product estimation suffers, directly lowering recall.
How TQ+ Per-Coordinate Calibration Works
TurboQuant+ (TQ+) solves the mis-match with two scalars per dimension: a shift and a scale. These are derived from empirical quantiles and applied to every vector exactly once.
Fitting Shift and Scale from Empirical Quantiles
The calibration is computed in compute_tqplus_calibration (turbovec/src/encode.rs, lines 36-82). For each rotated dimension, the algorithm measures the empirical 5% and 95% quantiles. It then solves for a shift and scale that map these empirical quantiles onto the theoretical 5% and 95% quantiles of the Beta marginal. This fit happens automatically during the first add call, provided the batch contains at least 1,000 vectors. Once computed, the parameters are frozen and reused forever.
Applying Calibration During Encoding
With parameters in hand, the encoder transforms each rotated coordinate:
u_calibrated[d] = (u_rot[d] + shift[d]) * scale_tq[d]
In turbovec/src/encode.rs (lines 98-108), this operation produces rotated_calib. The calibrated values then flow into fused_quantize_scale_pack (lines 86-114), where they are quantized using the global Lloyd-Max boundaries. Because the distribution now matches the codebook design, quantization distortion drops.
Inverting Calibration at Query Time
Search preserves inner-product accuracy by applying the inverse transform on the query side. The query vector is rotated, divided by scale_tq, and adjusted by a bias term equal to -⟨q_rot, shift⟩ (see comments in encode.rs, lines 24-27). Because both index and query paths use the same calibration, the SIMD scoring kernel requires no structural changes.
Recall Improvement and Benchmark Results
The README (README.md, lines 71-73) documents the net effect: a recall boost of up to +1.4 percentage points at R@1 on the hardest datasets, such as GloVe with 2-bit quantization. The gain is most pronounced at low bit-widths, where quantization error otherwise dominates the distortion budget. Benchmark data in benchmarks/results/recall_d1536_2bit.json concretely measures these improvements for 1,536-dimensional vectors.
Using TQ+ Calibration in Python and Rust
No extra API calls are required to enable TQ+.
Python
from turbovec import TurboQuantIndex
import numpy as np
# 1536-dimensional embeddings, 4-bit quantization
index = TurboQuantIndex(dim=1536, bit_width=4)
# First add triggers TQ+ calibration (requires >=1000 vectors for full gain)
vectors = np.random.randn(2000, 1536).astype(np.float32)
index.add(vectors)
# Subsequent adds reuse the same frozen calibration
more_vectors = np.random.randn(500, 1536).astype(np.float32)
index.add(more_vectors)
# Search applies the inverse calibration automatically
query = np.random.randn(1536).astype(np.float32)
scores, ids = index.search(query, k=10)
Rust
use turbovec::{TurboQuantIndex, encode};
let dim = 1536;
let bit_width = 4;
let mut index = TurboQuantIndex::new(dim, bit_width);
let vectors = vec![/* flatten d-dim f32 data */];
// None for existing calibration → triggers TQ+ fitting
let (packed, scales, shift, scale_tq) = encode::encode(
&vectors,
2000,
dim,
&rotation_matrix,
&boundaries,
¢roids,
bit_width,
None,
);
index.load_encoded(packed, scales, shift, scale_tq);
In both interfaces, the calibration vectors (shift and scale_tq) are persisted inside the index. The high-level TurboQuantIndex API handles this transparently, while the low-level encode function exposes the parameters for custom storage.
Summary
- TQ+ per-coordinate calibration learns a shift and scale per dimension by aligning empirical 5% and 95% quantiles to the theoretical Beta marginal.
- The fit occurs once during the first
add(batch ≥ 1000 vectors) and is frozen thereafter, adding no rebuild cost. - Calibrated vectors are quantized in
fused_quantize_scale_packwhile queries apply the inverse transform before the SIMD kernel. - The result is up to +1.4 pp higher R@1 recall, especially valuable at aggressive 2-bit and 4-bit settings.
- Memory overhead is negligible: only two
f32values per dimension.
Frequently Asked Questions
What triggers TQ+ calibration in turbovec?
The first call to add automatically triggers compute_tqplus_calibration in turbovec/src/encode.rs whenever the batch contains at least 1,000 vectors. If fewer vectors are supplied, calibration is deferred until enough data is available.
How much memory does per-coordinate calibration add?
Each dimension stores exactly two f32 values: shift[d] and scale_tq[d]. For a 1,536-dimensional index, this adds roughly 12 KB—negligible compared to the quantized payload.
Does TQ+ calibration require rebuilding the index?
No. The calibration parameters are computed once and then frozen. All subsequent add calls reuse the same shift and scale_tq, so there is no extra training phase or index rebuild.
Which bit-widths benefit most from TQ+ calibration?
Low bit-widths see the largest gains because quantization error is the dominant distortion source. The README highlights improvements on 2-bit GloVe embeddings, where matching the Beta marginal to the Lloyd-Max codebook is most critical.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →