When and How to Apply TQ+ Calibration in Turbovec for Maximum Recall

TurboVect's TQ+ calibration applies per-coordinate shift and scale transformations that boost recall by roughly +2.5 points on R@10 on average, reaching up to +8.7 points on highly anisotropic data when applied before adding vectors.

The TQ+ calibration feature in RyanCodrai/turbovec is the mechanism that delivers the recall improvements promised by TurboQuant's quantization strategy. By fitting empirical quantiles to a representative sample, TQ+ calibration determines per-coordinate shifts and scales that counteract anisotropic data distributions. However, timing matters—the operation yields maximum benefit only when invoked before vector ingestion, as documented in docs/api.md and verified in turbovec/tests/tqplus_calibration.rs.

Understanding TQ+ Calibration States

The calibration lifecycle defined in the API reference consists of two distinct states that determine how vectors are encoded:

  • "uncalibrated" — No calibration has been committed; the index functions but without the TQ+ recall gain because quantization uses default parameters.
  • "calibrated" — A calibration is committed; every stored vector, including those added before the calibration call, is re-encoded using the fitted per-coordinate shift and scale.

You can inspect the current state at any time via the idx.calibration_state property, as validated in turbovec/tests/explicit_calibration.rs.

When to Apply TQ+ Calibration for Maximum Recall

The timing of the calibrate() call relative to add() operations determines the magnitude of recall improvement. The source code and documentation identify three distinct scenarios:

Before Adding Vectors (Optimal)

Call idx.calibrate(sample) before any add operations. This ensures every vector is encoded using the fitted shift and scale from the start, avoiding quantization error accumulation. This yields the maximum possible recall because vectors undergo only a single quantization step.

After Adding a Small Set (Acceptable)

If the index contains fewer than roughly 1,000 vectors, you can still invoke calibration. The call will re-encode every stored row using the new calibration parameters. Recall loss remains minimal provided the sample is representative, as the documentation notes that the codebook's empirical quantiles stabilize around the TQPLUS_MIN_SAMPLES threshold (≈ 1,000 rows).

After Large Uncalibrated Ingest (Sub-optimal)

Performing calibration after many uncalibrated vectors have been added incurs a noticeable recall penalty. The re-encode becomes a second quantization step that compounds error. According to the API documentation, calibrating after a large uncalibrated ingest "costs several points of recall versus calibrating first."

Step-by-Step Calibration Process

Follow these steps to implement TQ+ calibration correctly in Python:

1. Prepare a Representative Sample

Select approximately 1,024 random rows to estimate the data distribution. Using 1,024 rows provides roughly half a point of R@10 improvement compared to the full corpus, while 2,048 rows are considered safe for all corpora according to the API reference.

import numpy as np

# vectors is your full dataset (numpy array)

rng = np.random.default_rng(seed=42)
sample_idx = rng.choice(len(vectors), size=1024, replace=False)
sample = vectors[sample_idx]

2. Fit and Commit the Calibration

The calibrate() method fits a (shift, scale) pair for each coordinate using empirical quantiles of the sample. This commits the calibration and transitions calibration_state to "calibrated".

from turbovec import TurboQuantIndex

idx = TurboQuantIndex(bit_width=4)  # dimension inferred on first add

idx.calibrate(sample)               # commits TQ+ calibration

3. Verify the Calibration State

Confirm the transition to ensure your vectors will be encoded correctly.

assert idx.calibration_state == "calibrated"

4. Add Your Vectors

With calibration committed, all subsequent additions use the optimized parameters.

idx.add(vectors)  # all vectors encoded with calibrated shift/scale

Handling Existing Indexes

The appropriate action depends on the current state of your index:

Situation Recommended Action
Building a new index Sample → calibrate() → add().
Existing index with < 1,000 vectors Calibrate immediately; re-encoding cost is negligible.
Existing index with ≥ 1,000 vectors Rebuild from source vectors with calibration applied before the first add. If rebuilding is impossible, call calibrate() now and accept a modest recall drop.
Repeated calibrations Safe, but ensure each new sample remains representative to avoid over-biasing.

Programmatically Verifying Calibration Parameters

For debugging or logging purposes, you can inspect the fitted parameters after calibration:


# Only available when idx.calibration_state == "calibrated"

print("Shift vector length:", len(idx.tqplus_shift()))   # equals dimension

print("Scale vector length:", len(idx.tqplus_scale()))   # equals dimension

If you need to refine calibration later, simply call calibrate() again with a new sample. Provided the new calibration is close to the previous one, the operation is essentially free because the codes reach a fixed point. However, a badly biased calibration cannot be repaired—you must rebuild the index from original vectors.

Summary

  • TQ+ calibration in Turbovec delivers +2.5 to +8.7 R@10 improvement by applying per-coordinate shift and scale transformations.
  • For maximum recall, always call calibrate() before adding vectors, using a random sample of approximately 1,024 representative rows.
  • The index maintains a calibration_state property ("uncalibrated" or "calibrated") that you can query at any time.
  • Calibration re-encodes all existing vectors, making early calibration critical to avoid double-quantization penalties on large datasets.
  • Access fitted parameters via tqplus_shift() and tqplus_scale() methods when the index is calibrated.

Frequently Asked Questions

What is the minimum sample size for TQ+ calibration?

While the TQPLUS_MIN_SAMPLES constant gates automatic fitting at approximately 1,000 vectors, the documentation recommends using 1,024 rows for robust estimation. Using 2,048 rows provides additional safety margins for highly anisotropic distributions without significant computational overhead.

Can I recalibrate an index after adding vectors?

Yes. Calling calibrate() again re-encodes every stored row using the new parameters. This is safe when the new sample is representative and similar to the previous calibration. However, if the initial calibration was severely biased, subsequent recalibration cannot fully recover the lost recall—you must rebuild the index from the original vectors.

How much recall improvement does TQ+ calibration provide?

According to the README.md performance summaries and API documentation, TQ+ calibration typically improves R@10 by approximately 2.5 points on average, with gains reaching up to 8.7 points on highly anisotropic datasets where coordinate distributions vary significantly.

What happens if I use a biased sample for calibration?

A biased sample produces incorrect empirical quantiles, leading to suboptimal shift and scale parameters. The calibrate() method commits these parameters immediately, and all vectors (existing and new) are encoded using them. The turbovec/tests/explicit_calibration.rs test suite confirms that once committed, a bad calibration persists until you rebuild the index or recalibrate with a representative sample.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →