How Per-Coordinate Calibration (TQ+) Works in Turbovec: A Complete Guide

TLDR: Per-Coordinate Calibration (TQ+) in Turbovec is a two-vector shift-and-scale transformation (tqplus_shift and tqplus_scale) derived from a user-supplied sample, applied to rotated coordinates before quantization during both encoding and search to dramatically improve recall over plain TurboQuant.

Turbovec (also called TurboQuantum) is a Rust library for approximate nearest neighbor search that stores vectors in a compressed, data-oblivious format. But the standard compression has a weakness: its fixed codebook boundaries, derived from bit-width alone, rarely match the actual distribution of your data. To fix this, the repository introduces an optional per-coordinate calibration step, called TQ+, that lets you adapt the quantization to your data with just a few lines of code. This article walks through the complete TQ+ workflow — what it is, how the calibration vectors are computed, and how they're applied during encoding and search.

What Is Per-Coordinate Calibration (TQ+)?

TQ+ is a data-driven adjustment layer on top of TurboQuant's base quantization. It consists of two per-coordinate vectors:

  • tqplus_shift: a length-dim vector of float values acting as an offset for each coordinate.
  • tqplus_scale: a length-dim vector of float values acting as a multiplicative scaling factor.

When both vectors are non-empty, the index is in the CalibrationState::Calibrated state. An empty pair indicates CalibrationState::Uncalibrated, meaning the index uses plain TurboQuant without calibration.

The key insight is that these calibration vectors are derived entirely from a sample you supply to the TurboQuantIndex::calibrate method. No other data — not the stored rows, not the query set, not any external metadata — influences them. Once fitted, the calibration vectors remain immutable for the life of the index.

The Calibration Fitting Process

The calibration process itself is straightforward. Let's break down the steps as implemented in Turbovec's source.

Step 1: Sampling

You — the caller — provide a sample set of vectors, typically a few thousand rows from your actual dataset. This sample gives the algorithm a statistical picture of what the real data looks like.

Step 2: Computing Order Statistics

For each coordinate d in 0..dim, the algorithm computes the empirical low and high quantiles of the rotated coordinate values. These are not fixed percentiles — they are the actual rank-based order statistics of the sampled coordinate.

Step 3: Mapping to Codebook Boundaries

The quantiles are then mapped onto the codebook boundaries that the encoder expects:

  • tqplus_shift[d] = the low quantile value, acting as an offset.
  • tqplus_scale[d] = the ratio of the high‑quantile width (distance between high and low quantiles) to the codebook's numeric width, acting as a multiplicative scale.

This is the crucial step — you're remapping the data distribution so that it fills the codebook's numeric range as evenly as possible.

Step 4: Storing the Pair

The resulting (shift, scale) pair is stored in the index fields tqplus_shift and tqplus_scale.

When and Where TQ+ is Applied

During Encoding

When vectors are added to the index, the encoder in turbovec/src/encode.rs receives the current calibration pair (or None if uncalibrated) and applies the transformation before quantizing. Here is a simplified version:

// inside encode::encode – simplified
for d in 0..dim {
    let rotated = rotated_coords[d];
    let calibrated = (rotated - shift[d]) * scale[d];
    // then quantise `calibrated` against the codebook
}

The query vectors go through exactly the same shift-and-scale transformation before being compared against the stored codes. The low-level search kernel in turbovec/src/search.rs includes a dedicated function:

// turbovec/src/search.rs – signature excerpt
fn calibrate_queries(
    q_rot: &mut [f32],
    tqplus_shift: &[f32],
    tqplus_scale: &[f32],
    nq: usize,
    dim: usize,
)

This kernel uses the identical formulas as the encoder, guaranteeing that a query encoded with the current calibration will be directly comparable to rowthat were encoded under the same calibration.

The Calibration Lifecycle and State Transitions

The calibration vectors have a strict, well-documented lifecycle in Turbovec. Knowing when they change — and when they don't — is essential to working with TQ+.

Operation Effect on calibration vectors
TurboQuantIndex::new / new_lazy Both vectors start empty (Uncalibrated).
TurboQuantIndex::calibrate(sample) Computes a new (shift, scale) pair and re-encodes every stored row with it, moving to Calibrated.
add, swap_remove, load, save Never change the calibration vectors; they remain exactly what the last calibrate call produced.
remove_calibration (via API) Not provided – clearing the vectors is not a public operation; an index stays calibrated once fitted.

The critical takeaway: the stored codes are a pure function of (rows, calibration). Because the calibration vectors are immutable after they're set, adding or removing rows does not alter the calibration, and loading a pre-TQ+ file yields empty vectors, placing the index in the Uncalibrated state.

A Complete Code Example

Here's a fully runnable example that demonstrates the entire per-coordinate calibration workflow, from index creation to search:

use turbovec::{TurboQuantIndex, CalibrationState};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // 1. Create an index (e.g. 1536-dim, 4-bit per coordinate)
    let mut idx = TurboQuantIndex::new(1536, 4)?;

    // 2. Add some vectors (un-calibrated at this point)
    idx.add(&vec![0.0_f32; 1536 * 10]);

    // 3. Inspect calibration state – should be Uncalibrated
    assert_eq!(idx.calibration_state(), CalibrationState::Uncalibrated);

    // 4. Provide a sample for per-coordinate calibration
    let sample = vec![0.0_f32; 1536 * 5]; // normally a real sample
    idx.calibrate(&sample)?;

    // 5. Now the index is calibrated
    assert_eq!(idx.calibration_state(), CalibrationState::Calibrated);
    assert!(!idx.tqplus_shift().is_empty()); // non-empty shift vector
    assert!(!idx.tqplus_scale().is_empty()); // non-empty scale vector

    // 6. Search – queries are automatically calibrated with the same pair
    let queries = vec![0.0_f32; 1536 * 2];
    let results = idx.search(&queries, 10);
    println!("top-10 for each query: {:?}", results.indices);

    Ok(())
}

Key Source File Locations

To dig deeper into the implementation yourself, these are the files you should look at:

Summary

  • TQ+ (Per-Coordinate Calibration) uses two vectors, tqplus_shift and tqplus_scale, to remap data into the codebook's optimal range.
  • The calibration vectors are derived solely from a user-supplied sample via TurboQuantIndex::calibrate() — no other data affects them.
  • The same shift-and-scale transformation is applied identically during encoding and search, ensuring query comparability.
  • Calibration-state transitions are strictly controlled: once you calibrate, the vectors are immutable until you reload the entire index from disk.
  • This approach is implemented in cleanly in lib.rs, encode.rs, and search.rs, with thorough test coverage in turbovec/tests/.

Frequently Asked Questions

When should I call calibrate in Turbovec?

You should call calibrate after adding your data but before intensive search workloads — the calibration must reflect the actual distribution of your data to be useful. Passing a sample of a few thousand rows that are representative of your entire dataset is typically sufficient.

Does calibrating an index re-encode all existing rows?

Yes. When you call TurboQuantIndex::calibrate(sample), the library computes the new (shift, scale) pair and re-encodes every stored row under the new calibration. This is a synchronous operation with the index, so plan for the added CPU cost across current run the first time you calibrate.

Can you remove calibration once it's applied?

No — Turbovec does not expose a public API to clear the calibration vectors. Once an index is calibrated, it stays calibrated until you create a new index or load a file that was saved without calibration (which leaves the vectors empty, putting the index in CalibrationState::Uncalibrated).

How does TQ+ differ from plain TurboQuant?

Plain TurboQuant uses a fixed, data-oblivious codebook — the quantization boundaries are the same regardless of what your actual data looks like. TQ+ adds a simple per-coordinate shift-and-scale transformation that spreads your data across the full codebook width, measurable improving recall in real-world scenarios, as validated in the turbovec test suite.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →