# How Per-Coordinate Calibration (TQ+) Works in Turbovec: A Complete Guide

> Learn how Per-Coordinate Calibration TQ+ works in Turbovec. Understand the shift-and-scale transformation for improved recall in TurboQuant. A complete guide.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: deep-dive
- Published: 2026-08-22

---

**TLDR: Per-Coordinate Calibration (TQ+) in Turbovec is a two-vector shift-and-scale transformation (`tqplus_shift` and `tqplus_scale`) derived from a user-supplied sample, applied to rotated coordinates before quantization during both encoding and search to dramatically improve recall over plain TurboQuant.**

Turbovec (also called TurboQuantum) is a Rust library for approximate nearest neighbor search that stores vectors in a compressed, data-oblivious format. But the standard compression has a weakness: its fixed codebook boundaries, derived from bit-width alone, rarely match the actual distribution of your data. To fix this, the repository introduces an optional **per-coordinate calibration** step, called **TQ+**, that lets you adapt the quantization to your data with just a few lines of code. This article walks through the complete TQ+ workflow — what it is, how the calibration vectors are computed, and how they're applied during encoding and search.

## What Is Per-Coordinate Calibration (TQ+)?

**TQ+ is a data-driven adjustment layer** on top of TurboQuant's base quantization. It consists of two per-coordinate vectors:

- **`tqplus_shift`**: a length-`dim` vector of float values acting as an offset for each coordinate.
- **`tqplus_scale`**: a length-`dim` vector of float values acting as a multiplicative scaling factor.

When both vectors are non-empty, the index is in the `CalibrationState::Calibrated` state. An empty pair indicates `CalibrationState::Uncalibrated`, meaning the index uses plain TurboQuant without calibration.

The key insight is that these calibration vectors are **derived entirely from a sample** you supply to the [`TurboQuantIndex::calibrate`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs#L58-L86) method. No other data — not the stored rows, not the query set, not any external metadata — influences them. Once fitted, the calibration vectors remain **immutable** for the life of the index.

## The Calibration Fitting Process

The calibration process itself is straightforward. Let's break down the steps as implemented in Turbovec's source.

### Step 1: Sampling

You — the caller — provide a sample set of vectors, typically a few thousand rows from your actual dataset. This sample gives the algorithm a statistical picture of what the real data looks like.

### Step 2: Computing Order Statistics

For each coordinate `d` in `0..dim`, the algorithm computes the **empirical low and high quantiles** of the rotated coordinate values. These are not fixed percentiles — they are the actual rank-based order statistics of the sampled coordinate.

### Step 3: Mapping to Codebook Boundaries

The quantiles are then mapped onto the codebook boundaries that the encoder expects:

- `tqplus_shift[d]` = the low quantile value, acting as an offset.
- `tqplus_scale[d]` = the ratio of the high‑quantile width (distance between high and low quantiles) to the codebook's numeric width, acting as a multiplicative scale.

This is the crucial step — you're remapping the data distribution so that it fills the codebook's numeric range as evenly as possible.

### Step 4: Storing the Pair

The resulting `(shift, scale)` pair is stored in the index fields `tqplus_shift` and `tqplus_scale`.

## When and Where TQ+ is Applied

### During Encoding

When vectors are added to the index, the encoder in [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs) receives the current calibration pair (or `None` if uncalibrated) and applies the transformation *before* quantizing. Here is a simplified version:

```rust
// inside encode::encode – simplified
for d in 0..dim {
    let rotated = rotated_coords[d];
    let calibrated = (rotated - shift[d]) * scale[d];
    // then quantise `calibrated` against the codebook
}

```

### During Search

The query vectors go through exactly the same shift-and-scale transformation before being compared against the stored codes. The low-level search kernel in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) includes a dedicated function:

```rust
// turbovec/src/search.rs – signature excerpt
fn calibrate_queries(
    q_rot: &mut [f32],
    tqplus_shift: &[f32],
    tqplus_scale: &[f32],
    nq: usize,
    dim: usize,
)

```

This kernel uses the identical formulas as the encoder, guaranteeing that a query encoded with the current calibration will be directly comparable to rowthat were encoded under the same calibration.

## The Calibration Lifecycle and State Transitions

The calibration vectors have a strict, well-documented lifecycle in Turbovec. Knowing when they change — and when they *don't* — is essential to working with TQ+.

| Operation | Effect on calibration vectors |
|---|---|
| `TurboQuantIndex::new` / `new_lazy` | Both vectors start empty (`Uncalibrated`). |
| `TurboQuantIndex::calibrate(sample)` | Computes a new `(shift, scale)` pair and **re-encodes every stored row** with it, moving to `Calibrated`. |
| `add`, `swap_remove`, `load`, `save` | **Never change** the calibration vectors; they remain exactly what the last `calibrate` call produced. |
| `remove_calibration` (via API) | Not provided – clearing the vectors is not a public operation; an index stays calibrated once fitted. |

The critical takeaway: **the stored codes are a pure function of `(rows, calibration)`**. Because the calibration vectors are immutable after they're set, adding or removing rows does not alter the calibration, and loading a pre-TQ+ file yields empty vectors, placing the index in the `Uncalibrated` state.

## A Complete Code Example

Here's a fully runnable example that demonstrates the entire per-coordinate calibration workflow, from index creation to search:

```rust
use turbovec::{TurboQuantIndex, CalibrationState};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // 1. Create an index (e.g. 1536-dim, 4-bit per coordinate)
    let mut idx = TurboQuantIndex::new(1536, 4)?;

    // 2. Add some vectors (un-calibrated at this point)
    idx.add(&vec![0.0_f32; 1536 * 10]);

    // 3. Inspect calibration state – should be Uncalibrated
    assert_eq!(idx.calibration_state(), CalibrationState::Uncalibrated);

    // 4. Provide a sample for per-coordinate calibration
    let sample = vec![0.0_f32; 1536 * 5]; // normally a real sample
    idx.calibrate(&sample)?;

    // 5. Now the index is calibrated
    assert_eq!(idx.calibration_state(), CalibrationState::Calibrated);
    assert!(!idx.tqplus_shift().is_empty()); // non-empty shift vector
    assert!(!idx.tqplus_scale().is_empty()); // non-empty scale vector

    // 6. Search – queries are automatically calibrated with the same pair
    let queries = vec![0.0_f32; 1536 * 2];
    let results = idx.search(&queries, 10);
    println!("top-10 for each query: {:?}", results.indices);

    Ok(())
}

```

## Key Source File Locations

To dig deeper into the implementation yourself, these are the files you should look at:

- **[`turbovec/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs)** — Defines the `CalibrationState` enum, stores `tqplus_shift` / `tqplus_scale`, and documents the immutable calibration lifecycle. [link](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs)
- **[`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs)** — Implements the encoder; receives the optional `(shift, scale)` pair and applies the per-coordinate transformation before quantization. [link](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs)
- **[`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs)** — Low-level query kernel; `calibrate_queries` applies the same shift/scale to rotated query coordinates. [link](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs)
- **[`turbovec/tests/tqplus_calibration.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/tqplus_calibration.rs)** — Test suite for calibration fitting, re-encoding, and persistence. [link](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/tqplus_calibration.rs)
- **[`turbovec/tests/explicit_calibration.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/explicit_calibration.rs)** — Additional tests for calibration-state transitions and correctness after reload. [link](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/explicit_calibration.rs)

## Summary

- **TQ+ (Per-Coordinate Calibration)** uses two vectors, `tqplus_shift` and `tqplus_scale`, to remap data into the codebook's optimal range.
- The calibration vectors are derived **solely from a user-supplied sample** via `TurboQuantIndex::calibrate()` — no other data affects them.
- The same shift-and-scale transformation is applied **identically** during encoding and search, ensuring query comparability.
- Calibration-state transitions are strictly controlled: **once you calibrate, the vectors are immutable** until you reload the entire index from disk.
- This approach is implemented in cleanly in [`lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/lib.rs), [`encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/encode.rs), and [`search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/search.rs), with thorough test coverage in `turbovec/tests/`.

## Frequently Asked Questions

### When should I call `calibrate` in Turbovec?

You should call `calibrate` after adding your data but *before* intensive search workloads — the calibration must reflect the actual distribution of your data to be useful. Passing a sample of a few thousand rows that are representative of your entire dataset is typically sufficient.

### Does calibrating an index re-encode all existing rows?

Yes. When you call `TurboQuantIndex::calibrate(sample)`, the library computes the new `(shift, scale)` pair and **re-encodes every stored row** under the new calibration. This is a synchronous operation with the index, so plan for the added CPU cost across current run the first time you calibrate.

### Can you remove calibration once it's applied?

No — Turbovec does not expose a public API to clear the calibration vectors. Once an index is calibrated, it stays calibrated until you create a new index or load a file that was saved without calibration (which leaves the vectors empty, putting the index in `CalibrationState::Uncalibrated`).

### How does TQ+ differ from plain TurboQuant?

Plain TurboQuant uses a fixed, data-oblivious codebook — the quantization boundaries are the same regardless of what your actual data looks like. TQ+ adds a simple per-coordinate shift-and-scale transformation that spreads your data across the full codebook width, measurable improving recall in real-world scenarios, as validated in the turbovec test suite.