How Per-Coordinate Calibration (TQ+) Works in Turbovec: A Complete Guide
TLDR: Per-Coordinate Calibration (TQ+) in Turbovec is a two-vector shift-and-scale transformation (tqplus_shift and tqplus_scale) derived from a user-supplied sample, applied to rotated coordinates before quantization during both encoding and search to dramatically improve recall over plain TurboQuant.
Turbovec (also called TurboQuantum) is a Rust library for approximate nearest neighbor search that stores vectors in a compressed, data-oblivious format. But the standard compression has a weakness: its fixed codebook boundaries, derived from bit-width alone, rarely match the actual distribution of your data. To fix this, the repository introduces an optional per-coordinate calibration step, called TQ+, that lets you adapt the quantization to your data with just a few lines of code. This article walks through the complete TQ+ workflow — what it is, how the calibration vectors are computed, and how they're applied during encoding and search.
What Is Per-Coordinate Calibration (TQ+)?
TQ+ is a data-driven adjustment layer on top of TurboQuant's base quantization. It consists of two per-coordinate vectors:
tqplus_shift: a length-dimvector of float values acting as an offset for each coordinate.tqplus_scale: a length-dimvector of float values acting as a multiplicative scaling factor.
When both vectors are non-empty, the index is in the CalibrationState::Calibrated state. An empty pair indicates CalibrationState::Uncalibrated, meaning the index uses plain TurboQuant without calibration.
The key insight is that these calibration vectors are derived entirely from a sample you supply to the TurboQuantIndex::calibrate method. No other data — not the stored rows, not the query set, not any external metadata — influences them. Once fitted, the calibration vectors remain immutable for the life of the index.
The Calibration Fitting Process
The calibration process itself is straightforward. Let's break down the steps as implemented in Turbovec's source.
Step 1: Sampling
You — the caller — provide a sample set of vectors, typically a few thousand rows from your actual dataset. This sample gives the algorithm a statistical picture of what the real data looks like.
Step 2: Computing Order Statistics
For each coordinate d in 0..dim, the algorithm computes the empirical low and high quantiles of the rotated coordinate values. These are not fixed percentiles — they are the actual rank-based order statistics of the sampled coordinate.
Step 3: Mapping to Codebook Boundaries
The quantiles are then mapped onto the codebook boundaries that the encoder expects:
tqplus_shift[d]= the low quantile value, acting as an offset.tqplus_scale[d]= the ratio of the high‑quantile width (distance between high and low quantiles) to the codebook's numeric width, acting as a multiplicative scale.
This is the crucial step — you're remapping the data distribution so that it fills the codebook's numeric range as evenly as possible.
Step 4: Storing the Pair
The resulting (shift, scale) pair is stored in the index fields tqplus_shift and tqplus_scale.
When and Where TQ+ is Applied
During Encoding
When vectors are added to the index, the encoder in turbovec/src/encode.rs receives the current calibration pair (or None if uncalibrated) and applies the transformation before quantizing. Here is a simplified version:
// inside encode::encode – simplified
for d in 0..dim {
let rotated = rotated_coords[d];
let calibrated = (rotated - shift[d]) * scale[d];
// then quantise `calibrated` against the codebook
}
During Search
The query vectors go through exactly the same shift-and-scale transformation before being compared against the stored codes. The low-level search kernel in turbovec/src/search.rs includes a dedicated function:
// turbovec/src/search.rs – signature excerpt
fn calibrate_queries(
q_rot: &mut [f32],
tqplus_shift: &[f32],
tqplus_scale: &[f32],
nq: usize,
dim: usize,
)
This kernel uses the identical formulas as the encoder, guaranteeing that a query encoded with the current calibration will be directly comparable to rowthat were encoded under the same calibration.
The Calibration Lifecycle and State Transitions
The calibration vectors have a strict, well-documented lifecycle in Turbovec. Knowing when they change — and when they don't — is essential to working with TQ+.
| Operation | Effect on calibration vectors |
|---|---|
TurboQuantIndex::new / new_lazy |
Both vectors start empty (Uncalibrated). |
TurboQuantIndex::calibrate(sample) |
Computes a new (shift, scale) pair and re-encodes every stored row with it, moving to Calibrated. |
add, swap_remove, load, save |
Never change the calibration vectors; they remain exactly what the last calibrate call produced. |
remove_calibration (via API) |
Not provided – clearing the vectors is not a public operation; an index stays calibrated once fitted. |
The critical takeaway: the stored codes are a pure function of (rows, calibration). Because the calibration vectors are immutable after they're set, adding or removing rows does not alter the calibration, and loading a pre-TQ+ file yields empty vectors, placing the index in the Uncalibrated state.
A Complete Code Example
Here's a fully runnable example that demonstrates the entire per-coordinate calibration workflow, from index creation to search:
use turbovec::{TurboQuantIndex, CalibrationState};
fn main() -> Result<(), Box<dyn std::error::Error>> {
// 1. Create an index (e.g. 1536-dim, 4-bit per coordinate)
let mut idx = TurboQuantIndex::new(1536, 4)?;
// 2. Add some vectors (un-calibrated at this point)
idx.add(&vec![0.0_f32; 1536 * 10]);
// 3. Inspect calibration state – should be Uncalibrated
assert_eq!(idx.calibration_state(), CalibrationState::Uncalibrated);
// 4. Provide a sample for per-coordinate calibration
let sample = vec![0.0_f32; 1536 * 5]; // normally a real sample
idx.calibrate(&sample)?;
// 5. Now the index is calibrated
assert_eq!(idx.calibration_state(), CalibrationState::Calibrated);
assert!(!idx.tqplus_shift().is_empty()); // non-empty shift vector
assert!(!idx.tqplus_scale().is_empty()); // non-empty scale vector
// 6. Search – queries are automatically calibrated with the same pair
let queries = vec![0.0_f32; 1536 * 2];
let results = idx.search(&queries, 10);
println!("top-10 for each query: {:?}", results.indices);
Ok(())
}
Key Source File Locations
To dig deeper into the implementation yourself, these are the files you should look at:
turbovec/src/lib.rs— Defines theCalibrationStateenum, storestqplus_shift/tqplus_scale, and documents the immutable calibration lifecycle. linkturbovec/src/encode.rs— Implements the encoder; receives the optional(shift, scale)pair and applies the per-coordinate transformation before quantization. linkturbovec/src/search.rs— Low-level query kernel;calibrate_queriesapplies the same shift/scale to rotated query coordinates. linkturbovec/tests/tqplus_calibration.rs— Test suite for calibration fitting, re-encoding, and persistence. linkturbovec/tests/explicit_calibration.rs— Additional tests for calibration-state transitions and correctness after reload. link
Summary
- TQ+ (Per-Coordinate Calibration) uses two vectors,
tqplus_shiftandtqplus_scale, to remap data into the codebook's optimal range. - The calibration vectors are derived solely from a user-supplied sample via
TurboQuantIndex::calibrate()— no other data affects them. - The same shift-and-scale transformation is applied identically during encoding and search, ensuring query comparability.
- Calibration-state transitions are strictly controlled: once you calibrate, the vectors are immutable until you reload the entire index from disk.
- This approach is implemented in cleanly in
lib.rs,encode.rs, andsearch.rs, with thorough test coverage inturbovec/tests/.
Frequently Asked Questions
When should I call calibrate in Turbovec?
You should call calibrate after adding your data but before intensive search workloads — the calibration must reflect the actual distribution of your data to be useful. Passing a sample of a few thousand rows that are representative of your entire dataset is typically sufficient.
Does calibrating an index re-encode all existing rows?
Yes. When you call TurboQuantIndex::calibrate(sample), the library computes the new (shift, scale) pair and re-encodes every stored row under the new calibration. This is a synchronous operation with the index, so plan for the added CPU cost across current run the first time you calibrate.
Can you remove calibration once it's applied?
No — Turbovec does not expose a public API to clear the calibration vectors. Once an index is calibrated, it stays calibrated until you create a new index or load a file that was saved without calibration (which leaves the vectors empty, putting the index in CalibrationState::Uncalibrated).
How does TQ+ differ from plain TurboQuant?
Plain TurboQuant uses a fixed, data-oblivious codebook — the quantization boundaries are the same regardless of what your actual data looks like. TQ+ adds a simple per-coordinate shift-and-scale transformation that spreads your data across the full codebook width, measurable improving recall in real-world scenarios, as validated in the turbovec test suite.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →