TQ+ Per-Coordinate Calibration in TurboVec: Implementation and Usage
TQ+ per-coordinate calibration in TurboVec corrects anisotropic data distributions by fitting shift and scale parameters to each rotated coordinate, mapping empirical 5th and 95th percentiles to the theoretical Beta distribution expected after random rotation.
TurboVec is a high-performance vector quantization library that uses random rotation and Lloyd-Max quantization to compress high-dimensional vectors. Because real-world data is often anisotropic, the empirical distribution of rotated coordinates deviates from the theoretical expectation. TQ+ per-coordinate calibration addresses this by independently adjusting each dimension to match the canonical Beta distribution, improving quantization accuracy without modifying the core SIMD scoring kernel.
How TQ+ Calibration Works
Theoretical Foundation
After random rotation, high-dimensional vectors should theoretically follow a Beta((d-1)/2, (d-1)/2) distribution in each coordinate, where d is the dimensionality. Real-world datasets often diverge from this ideal due to anisotropic structure. TQ+ (per-coordinate shift + scale calibration) corrects this mismatch by transforming each coordinate to match the canonical distribution.
Quantile Matching Parameter Estimation
The calibration fits parameters by comparing empirical quantiles to theoretical targets:
-
Extract empirical quantiles: After rotating a batch, TurboVec calculates the 5% and 95% empirical quantiles for every coordinate using
statrs::Betafor the target distribution values. -
Solve for scale and shift: For each dimension
d, the system computes:scale_d = (q_95^β - q_5^β) / (x_95^emp - x_5^emp) shift_d = (q_5^β / scale_d) - x_5^empThis maps the empirical quantile pair onto the canonical Beta distribution.
-
Minimum batch size constraint: If the batch contains fewer than 1,000 vectors, or if a coordinate is effectively constant, calibration defaults to the identity transformation (
shift = 0,scale = 1) to avoid noisy quantile estimates.
Implementation in the TurboVec Codebase
Encoding and Calibration Fitting (encode.rs)
The compute_tqplus_calibration function in turbovec/src/encode.rs (lines 36-83) implements the quantile fitting logic. During encoding, the rotated vector u_rot is transformed before quantization:
u_calib[d] = (u_rot[d] + shift_d) * scale_d
This transformation occurs in the encode function (lines 86-95). The per-vector scale factor stored in the index remains the length-renormalization factor from the original paper, not the TQ+ scale.
Search-Time Inverse Transformation (search.rs)
At query time, the calibrate_queries function in turbovec/src/search.rs (lines 21-30) applies the inverse transformation. Queries are rotated, then each coordinate is divided by the stored scale. A bias correction term is computed as:
bias_q = -sum_d(u_rot[d] * shift_d)
This bias is folded into the per-query bias field of the lookup table, allowing the SIMD scoring kernel to remain unchanged while accounting for the calibration.
Storage and Serialization (lib.rs and io.rs)
Calibration vectors are stored in the index structure defined in turbovec/src/lib.rs (lines 114-119):
TurboIndex.tqplus_shift: Per-coordinate shift valuesTurboIndex.tqplus_scale: Per-coordinate scale values
The turbovec/src/io.rs module handles serialization of these vectors (lines 201-210), ensuring calibration persists across index saves and loads.
Handling Incremental Index Updates
When building an index incrementally, the first batch that calls add determines the calibration parameters. Subsequent batches must reuse the same (shift, scale) pair to prevent cross-batch drift. This is achieved via the existing_calibration argument of the encode function (lines 49-55 in encode.rs), which passes the previously computed calibration to new encoding operations.
Code Examples
Encoding with Automatic TQ+ Calibration
use turbovec::encode;
// `vectors` is a flat slice of `n * dim` f32 values.
let (packed, scales, shift, scale_tq) = encode(
&vectors,
n, // number of vectors
dim, // dimensionality
&rotation, // random rotation matrix (dim×dim)
&boundaries, // quantiser boundaries
¢roids, // Lloyd‑Max centroids
bit_width, // e.g. 4 for 4‑bit codes
None, // no prior calibration → fit new one
);
The shift and scale_tq variables contain the per-coordinate calibration parameters returned by compute_tqplus_calibration.
Reusing Calibration for Incremental Adds
let (packed2, scales2, _, _) = encode(
&new_vectors,
new_n,
dim,
&rotation,
&boundaries,
¢roids,
bit_width,
Some((&shift, &scale_tq)), // lock to the calibration of the first batch
);
Searching with Calibrated Indexes
use turbovec::search;
// All calibration vectors are stored inside the index; they are passed to `search`.
let (scores, ids) = search(
&queries,
nq,
&rotation,
&blocked_codes,
¢roids,
&vec_scales,
&index.tqplus_shift, // may be empty for legacy v2 indexes
&index.tqplus_scale,
bits,
dim,
n_vectors,
n_blocks,
k,
None, // optional mask
);
Inside search, the calibrate_queries function performs the inverse transformation and bias correction before the SIMD kernel executes.
Summary
- TQ+ calibration corrects anisotropic data by mapping empirical quantiles to the theoretical Beta((d-1)/2, (d-1)/2) distribution expected after random rotation.
- Per-coordinate parameters (
shiftandscale) are computed by matching 5th and 95th percentiles, with a minimum batch size of 1,000 vectors required for reliable estimation. - Implementation spans
encode.rs(fitting and application),search.rs(inverse transformation), andlib.rs/io.rs(storage and serialization). - Incremental indexing requires passing existing calibration parameters via
existing_calibrationto maintain consistency across batches. - Zero kernel changes: The inverse transformation is folded into query bias terms, leaving the SIMD scoring kernel unchanged from the standard implementation.
Frequently Asked Questions
What is the minimum batch size required for TQ+ calibration?
TurboVec requires at least 1,000 vectors in a batch to compute reliable quantile estimates for TQ+ calibration. If the batch is smaller, or if any coordinate is effectively constant, the system defaults to identity calibration (shift = 0, scale = 1) to prevent noise from distorting the quantization.
How does TQ+ calibration affect the SIMD scoring kernel?
The SIMD scoring kernel remains completely unchanged. Instead of modifying the kernel, TurboVec applies the inverse calibration to queries at search time (dividing by scale and adding a bias correction), then folds this correction into the per-query bias field of the lookup table. This maintains performance while achieving the accuracy benefits of calibration.
Can I disable TQ+ calibration for legacy compatibility?
Yes. TQ+ calibration is optional. If you pass empty vectors for tqplus_shift and tqplus_scale (or None for the calibration parameter in older API versions), the system operates in legacy mode without per-coordinate calibration. Legacy v2 indexes store empty calibration vectors by default for backward compatibility.
How is calibration handled when adding vectors incrementally?
The first batch processed during incremental indexing determines the calibration parameters. All subsequent batches must reuse these exact parameters by passing them via the existing_calibration argument to encode(). This ensures all vectors in the index share a single consistent calibration, preventing drift between different add operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →