TQ+ Per-Coordinate Calibration in turbovec: How It Improves Recall

TQ+ per-coordinate calibration learns a tiny shift and scale for every dimension to align rotated embeddings with a theoretical Beta distribution, eliminating quantizer mis-fit and boosting recall by up to 1.4 percentage points at R@1.

TurboQuant, the core algorithm behind the turbovec vector-search library by RyanCodrai, relies on a single random rotation to make every coordinate follow the same theoretical Beta distribution. Real-world embeddings are anisotropic, meaning some dimensions deviate from this ideal shape and cause the global Lloyd-Max codebook to mis-fit. TQ+ per-coordinate calibration solves this by learning per-dimension corrections once during the first batch, then freezing them for all future indexing and querying.

The Problem: Anisotropic Embeddings and Quantizer Mis-Fit

After normalization and random rotation, turbovec expects all coordinates to match a canonical Beta marginal. In turbovec/src/encode.rs, the rotation step (lines 68-84) projects vectors into this theoretical space. Real-world data, however, contains dimensions that are systematically wider or narrower than the Beta ideal. When the global scalar codebook—pre-computed for the perfect Beta shape—is applied to these shifted coordinates, quantization error increases and inner-product estimation suffers, directly lowering recall.

How TQ+ Per-Coordinate Calibration Works

TurboQuant+ (TQ+) solves the mis-match with two scalars per dimension: a shift and a scale. These are derived from empirical quantiles and applied to every vector exactly once.

Fitting Shift and Scale from Empirical Quantiles

The calibration is computed in compute_tqplus_calibration (turbovec/src/encode.rs, lines 36-82). For each rotated dimension, the algorithm measures the empirical 5% and 95% quantiles. It then solves for a shift and scale that map these empirical quantiles onto the theoretical 5% and 95% quantiles of the Beta marginal. This fit happens automatically during the first add call, provided the batch contains at least 1,000 vectors. Once computed, the parameters are frozen and reused forever.

Applying Calibration During Encoding

With parameters in hand, the encoder transforms each rotated coordinate:

u_calibrated[d] = (u_rot[d] + shift[d]) * scale_tq[d]

In turbovec/src/encode.rs (lines 98-108), this operation produces rotated_calib. The calibrated values then flow into fused_quantize_scale_pack (lines 86-114), where they are quantized using the global Lloyd-Max boundaries. Because the distribution now matches the codebook design, quantization distortion drops.

Inverting Calibration at Query Time

Search preserves inner-product accuracy by applying the inverse transform on the query side. The query vector is rotated, divided by scale_tq, and adjusted by a bias term equal to -⟨q_rot, shift⟩ (see comments in encode.rs, lines 24-27). Because both index and query paths use the same calibration, the SIMD scoring kernel requires no structural changes.

Recall Improvement and Benchmark Results

The README (README.md, lines 71-73) documents the net effect: a recall boost of up to +1.4 percentage points at R@1 on the hardest datasets, such as GloVe with 2-bit quantization. The gain is most pronounced at low bit-widths, where quantization error otherwise dominates the distortion budget. Benchmark data in benchmarks/results/recall_d1536_2bit.json concretely measures these improvements for 1,536-dimensional vectors.

Using TQ+ Calibration in Python and Rust

No extra API calls are required to enable TQ+.

Python

from turbovec import TurboQuantIndex
import numpy as np

# 1536-dimensional embeddings, 4-bit quantization

index = TurboQuantIndex(dim=1536, bit_width=4)

# First add triggers TQ+ calibration (requires >=1000 vectors for full gain)

vectors = np.random.randn(2000, 1536).astype(np.float32)
index.add(vectors)

# Subsequent adds reuse the same frozen calibration

more_vectors = np.random.randn(500, 1536).astype(np.float32)
index.add(more_vectors)

# Search applies the inverse calibration automatically

query = np.random.randn(1536).astype(np.float32)
scores, ids = index.search(query, k=10)

Rust

use turbovec::{TurboQuantIndex, encode};

let dim = 1536;
let bit_width = 4;
let mut index = TurboQuantIndex::new(dim, bit_width);

let vectors = vec![/* flatten d-dim f32 data */];

// None for existing calibration → triggers TQ+ fitting
let (packed, scales, shift, scale_tq) = encode::encode(
    &vectors,
    2000,
    dim,
    &rotation_matrix,
    &boundaries,
    &centroids,
    bit_width,
    None,
);

index.load_encoded(packed, scales, shift, scale_tq);

In both interfaces, the calibration vectors (shift and scale_tq) are persisted inside the index. The high-level TurboQuantIndex API handles this transparently, while the low-level encode function exposes the parameters for custom storage.

Summary

  • TQ+ per-coordinate calibration learns a shift and scale per dimension by aligning empirical 5% and 95% quantiles to the theoretical Beta marginal.
  • The fit occurs once during the first add (batch ≥ 1000 vectors) and is frozen thereafter, adding no rebuild cost.
  • Calibrated vectors are quantized in fused_quantize_scale_pack while queries apply the inverse transform before the SIMD kernel.
  • The result is up to +1.4 pp higher R@1 recall, especially valuable at aggressive 2-bit and 4-bit settings.
  • Memory overhead is negligible: only two f32 values per dimension.

Frequently Asked Questions

What triggers TQ+ calibration in turbovec?

The first call to add automatically triggers compute_tqplus_calibration in turbovec/src/encode.rs whenever the batch contains at least 1,000 vectors. If fewer vectors are supplied, calibration is deferred until enough data is available.

How much memory does per-coordinate calibration add?

Each dimension stores exactly two f32 values: shift[d] and scale_tq[d]. For a 1,536-dimensional index, this adds roughly 12 KB—negligible compared to the quantized payload.

Does TQ+ calibration require rebuilding the index?

No. The calibration parameters are computed once and then frozen. All subsequent add calls reuse the same shift and scale_tq, so there is no extra training phase or index rebuild.

Which bit-widths benefit most from TQ+ calibration?

Low bit-widths see the largest gains because quantization error is the dominant distortion source. The README highlights improvements on 2-bit GloVe embeddings, where matching the Beta marginal to the Lloyd-Max codebook is most critical.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →