# TQ+ Calibration in TurboVec: When and How It Improves Recall

> Discover TQ+ calibration in turbovec. Learn how this vector transformation aligns distributions to significantly boost recall, especially with large, non-isotropic datasets.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: deep-dive
- Published: 2026-06-16

---

**TQ+ calibration is a per-coordinate shift-and-scale transformation that aligns empirical vector distributions with the Beta distribution assumed by the Lloyd-Max quantizer, significantly boosting recall when fitting on 1,000 or more vectors from non-isotropic datasets.**

**TQ+ calibration** (TurboQuant Plus) is an adaptive preprocessing step in the `RyanCodrai/turbovec` vector search library that fixes distribution mismatches between real-world data and the theoretical assumptions of the scalar quantization codebook. While TurboVec's base pipeline normalizes, randomly rotates, and then quantizes vectors using a Lloyd-Max codebook trained on a Beta distribution, real datasets often deviate from this ideal. TQ+ measures these deviations and corrects them before quantization, reducing distortion and improving search recall under specific conditions.

## What Is TQ+ Calibration in TurboVec?

TurboVec stores vectors after three sequential steps: **normalize → random-rotate → quantize**. The Lloyd-Max quantizer assumes each rotated coordinate follows a theoretical *Beta* distribution `Beta((d-1)/2, (d-1)/2)`. In practice, data anisotropy means the empirical marginal distributions rarely match this assumption.

As implemented in [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs) (lines 10-27), TQ+ calibration computes a **shift** and **scale** for each dimension such that the empirical 5% and 95% quantiles align with the canonical Beta quantiles:

```text
u_calibrated[d] = (u_rot[d] + shift[d]) * scale_tq[d]

```

The calibrated values are then fed to the quantizer. At query time, the inverse transformation is applied on the fly to maintain consistency between the stored codes and the query vector.

## When TQ+ Calibration Activates in TurboVec

The calibration is **fit automatically on the first `add` call** of a batch. However, the system guards against noisy statistics with a minimum sample threshold.

According to the source in [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs) (lines 39-45 and 49-54), TurboVec checks `TQPLUS_MIN_SAMPLES` (currently set to **1000**):

- If the first batch contains **≥ 1,000 vectors**, TurboVec computes fresh `shift` and `scale_tq` arrays via `compute_tqplus_calibration`.
- If the batch is smaller, it falls back to **identity calibration** (shift = 0, scale = 1), incurring no overhead but providing no recall benefit.

Subsequent batches can reuse existing calibration parameters via `add_with_calibration` to maintain index consistency.

## Why TQ+ Calibration Improves Recall

By aligning the empirical coordinate distributions with the codebook's assumed Beta marginal, TQ+ reduces quantization error. This error reduction directly translates to higher recall, particularly under three conditions:

1. **Sufficient sample size** – The 1,000-vector threshold provides stable estimates of the 5%/95% quantiles.
2. **Non-isotropic data** – Datasets where per-coordinate variances differ significantly from the ideal isotropic case (e.g., GloVe embeddings) show the largest gains.
3. **Higher-bit encodings** – The effect is most pronounced at 4-bit precision, where quantization error dominates and the calibrated shift/scale can recover lost precision.

Benchmarks on the GLUE/GloVe and D3072 datasets (visualized in `docs/recall_glove.svg` and `docs/recall_d3072.svg`) demonstrate clear recall uplift when TQ+ is active compared to identity calibration.

## Implementing TQ+ Calibration in Rust

The following example demonstrates fitting calibration on the first batch and reusing it for subsequent additions:

```rust
use turbovec::TurboQuantIndex;

// Create a new index (dimension, bit-width, etc.)
let mut idx = TurboQuantIndex::new(dim, bit_width);

// First add – calibration is fitted automatically if ≥1000 vectors.
let (codes, scales, shift, scale_tq) = idx.add(&vectors)?;
// `shift` and `scale_tq` contain the TQ+ parameters for this batch.

// Subsequent adds – reuse the same calibration to keep the index consistent.
let existing = Some((&shift[..], &scale_tq[..]));
let (codes2, scales2, _, _) = idx.add_with_calibration(&more_vectors, existing)?;

```

Internally, the `encode` function in [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs) handles the logic:

```rust
// Inside encode.rs (excerpt)
let (shift, scale_tq) = match existing_calibration {
    Some((s, sc)) => (s.to_vec(), sc.to_vec()),
    None => compute_tqplus_calibration(rotated, n, dim),
};

```

At search time, the inverse transformation is applied to query vectors in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) (`q_calib[d] = q_rot[d] / scale_tq[d]`), ensuring the query and database vectors occupy the same calibrated space.

## Summary

- **TQ+ calibration** applies per-coordinate shift and scale factors to align empirical distributions with the Beta distribution assumed by the Lloyd-Max quantizer.
- It activates automatically on the first batch but requires **at least 1,000 vectors** (`TQPLUS_MIN_SAMPLES`) to compute reliable statistics; otherwise, it falls back to identity calibration.
- Recall improvements are most significant for **non-isotropic datasets** and **higher-bit encodings** (e.g., 4-bit), where quantization error reduction has the highest impact.
- Calibration parameters can be extracted after the first `add` and reused via `add_with_calibration` to ensure consistent encoding across multiple batches.

## Frequently Asked Questions

### What is the minimum number of vectors required for TQ+ calibration?

TurboVec requires **1,000 vectors** in the first added batch to activate TQ+ calibration. This threshold, defined as `TQPLUS_MIN_SAMPLES` in [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs), ensures the empirical 5% and 95% quantiles are stable enough to compute reliable shift and scale parameters. Smaller batches automatically use identity calibration (no shift, scale = 1).

### How does TQ+ calibration differ from standard TurboVec quantization?

Standard TurboVec quantization assumes each rotated coordinate follows a theoretical Beta distribution and quantizes directly. **TQ+ calibration** first measures the actual empirical distribution of your specific dataset and applies a linear transformation (shift and scale) to map those quantiles onto the expected Beta quantiles before quantization. This compensates for anisotropy in real-world data that the base quantizer cannot account for.

### Can I reuse TQ+ calibration parameters across different batches?

Yes. After the initial `add` call returns the `shift` and `scale_tq` arrays, you can pass them to `add_with_calibration` for subsequent batches. This is critical for maintaining consistency across an index built from multiple data chunks, ensuring all vectors are encoded in the same calibrated coordinate system.

### Does TQ+ calibration affect query-time performance?

No. The calibration parameters are **inverted and applied on the fly** to query vectors during the search phase (as seen in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs)). This adds negligible computational overhead because it consists of a single division per dimension (`q_rot[d] / scale_tq[d]`), which is dwarfed by the distance computation costs. The recall gains from reduced quantization error far outweigh this minimal latency increase.