Effective TQ+ Calibration Rules in Turbovec: A Complete Guide

TLDR: TQ+ calibration in Turbovec is an explicit, two-step process that requires a random representative sample of about 2,048 vectors, must be performed before bulk ingest for best results, and cannot repair a biased earlier calibration — so you must calibrate early and correctly the first time.

TurboVect's TQ+ (per-coordinate shift/scale) calibration is a deterministic procedure that dramatically improves recall on vector search — averaging roughly +2.5 recall points and reaching up to +8.7 points on highly anisotropic data. This guide walks through the exact rules for effective TQ+ calibration as implemented in the RyanCodrai/turbovec repository, with code examples, validation constraints, and best practices drawn directly from the source.

What Is TQ+ Calibration in Turbovec?

TQ+ is a calibration mechanism that applies a per-coordinate shift and scale transformation to stored vectors before quantization. It is a two-step deterministic process: you supply a sample of vectors, and Turbovec computes an optimal shift/scale pair that reduces quantization error for your specific data distribution.

The calibration lifecycle is enforced in the core library and documented in docs/api.md. Understanding the rules is essential — using TQ+ incorrectly can degrade recall, and a bad calibration is irreversible without a full rebuild.

Rule 1: Calibration Is Explicit Only

TQ+ is never fitted automatically. The index will only calibrate when you explicitly call:

  • Python: idx.calibrate(sample)
  • Rust: calibrate or calibrate_2d

According to the API docs, this explicit-only design guarantees that stored vectors are always encoded with the same calibrated coordinate system — there's no hidden state that could diverge between your code and the index's internal representation.

Rule 2: Understand the Calibration State

The index reports one of two states:

  • "uncalibrated" — no TQ+ gain is applied.
  • "calibrated" — a TQ+ shift/scale pair is committed.

When you call calibrate(), every stored row — including rows added before the calibration call — is re-encoded under the new shift/scale pair. This means calibration retroactively transforms your entire dataset, not just future inserts.

Rule 3: Sample Size and Quality Matter

The caller supplies a representative set of vectors. Empirical results from the Turbovec documentation show:

  • ≈1,024 random rows yield roughly half of the maximal recall improvement.
  • ≈2,048 rows achieve peak recall on most corpora.

More important than size is sample quality. The sample must be:

  • Random — never sorted or sequentially ordered.
  • Representative — drawn from the actual distribution of vectors you will search.

A biased, clustered, or sorted subset leads to a calibration that degrades recall and, as noted below, cannot be repaired later.

Rule 4: When to Calibrate — Timing Rules

Calibrate Before Bulk Ingest

Best practice: calibrate before bulk ingestion. Calibrating after a large uncalibrated ingest forces a second-quantization cost — you lose a few recall points that you cannot recover.

The recommendation in docs/api.md is clear:

Calibrate at index creation time, before adding the bulk of your vectors.

Re-Calibration Is Allowed but Risky

Re-calibration is permitted at any time; it simply re-encodes existing rows with the new pair. However, there is a critical caveat:

If you perform a poorly biased early calibration, later refits cannot fix the problem. The narrow fit clamped coordinates irreversibly.

In such cases the only remedy is a full rebuild from the original vectors. There is no in-place repair path — so the first calibration decision is essentially permanent.

Rule 5: Validation Guarantees

The TQ+ calibration pair (shift & scale) must satisfy strict constraints, enforced in the low-level constructor (from_parts) and loader:

  • Each shift must be finite.
  • Each scale must be finite and > 0.
  • The pair length must equal dim (or be empty, which means identity).

These constraints are implemented in turbovec/src/lib.rs, around lines 315–324. They guarantee that any persisted index round-trips correctly through:

  • write / load
  • to_bytes / from_bytes
  • Python pickling

Rule 6: Persistence of Calibration

The TQ+ pair is stored in the file format (.tv / .tvim) and is automatically re-loaded with the index. Importantly, draining an index (removing all rows) preserves the calibration state — a later rebuild can reuse the same pair if desired.

Code Examples: TQ+ Calibration in Python and Rust

Python — Typical Workflow


# -------------------------------------------------

# Python usage – typical workflow

# -------------------------------------------------

from turbovec import TurboQuantIndex
import numpy as np

# 1️⃣ Create an index (dim optional, inferred on first add)

idx = TurboQuantIndex(bit_width=4)

# 2️⃣ Add some vectors (any number, can be lazy)

vectors = np.random.randn(5000, 1536).astype(np.float32)
idx.add(vectors)

# 3️⃣ Sample 2,048 random rows for calibration (must be representative)

sample = vectors[np.random.choice(len(vectors), 2048, replace=False)]

# 4️⃣ Calibrate – this re-encodes all existing rows

idx.calibrate(sample)          # -> idx.calibration_state == "calibrated"

# 5️⃣ Continue adding more vectors; they inherit the TQ+ pair

more_vectors = np.random.randn(3000, 1536).astype(np.float32)
idx.add(more_vectors)

# 6️⃣ Search as usual – you now get the recall boost from TQ+

queries = np.random.randn(10, 1536).astype(np.float32)
scores, ids = idx.search(queries, k=10)

Rust — Lower-Level API

// -------------------------------------------------
// Rust usage – low-level API
// -------------------------------------------------
use turbovec::{TurboQuantIndex, CalibrateError};

fn main() -> Result<(), CalibrateError> {
    // Build a lazy index
    let mut idx = TurboQuantIndex::new(None, 4)?;

    // Add some vectors (flattened 2-D slice)
    let vectors: Vec<f32> = (0..5_000 * 1_536)
        .map(|_| rand::random::<f32>())
        .collect();
    idx.add_2d(&vectors, 1_536)?;

    // Pick 2,048 random rows as a sample
    let sample: Vec<f32> = idx
        .as_slice()
        .chunks(1_536)
        .choose_multiple(&mut rand::thread_rng(), 2_048)
        .flatten()
        .copied()
        .collect();

    // Calibrate – this rewrites all stored rows
    idx.calibrate(&sample)?;

    // Continue adding more vectors; they inherit the calibration
    let more: Vec<f32> = (0..3_000 * 1_536)
        .map(|_| rand::random::<f32>())
        .collect();
    idx.add_2d(&more, 1_536)?;

    Ok(())
}

Key Files in the Turbovec Repository

File Purpose Link
turbovec/src/lib.rs Core implementation; defines the calibration state, validation rules, and the public calibrate / calibrate_2d methods. https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs
turbovec/src/id_map.rs Wrapper around TurboQuantIndex; forwards calibration calls to the inner index. https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/id_map.rs
docs/api.md The documentation that describes the TQ+ calibration lifecycle, usage guidance, and sample-size recommendations. https://github.com/RyanCodrai/turbovec/blob/main/docs/api.md#L134-L152
turbovec/tests/explicit_calibration.rs Test suite confirming calibration behavior, re-encoding, and error handling. https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/explicit_calibration.rs
turbovec/tests/calibration_bounds.rs Validation tests ensuring shift/scale values respect finiteness and positivity constraints. https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/calibration_bounds.rs

Summary

  • TQ+ calibration is explicit: You must call calibrate() — the index never fits calibration automatically.

  • Sample quality matters: Use 2,048 random, representative rows for the best recall gains. A mis-selected sample degrades recall and cannot be fixed.

  • Calibrate before bulk ingest for optimal performance — a post-ingest calibration loses a few recall points via re-quantization.

  • Affects all rows: Calibration re-encodes every existing vector, not just new ones.

  • Persists across storage: The calibration pair is stored in the file and reloaded automatically.

  • Recalibration is permanent: A biased first calibration cannot be undone; it requires a full rebuild from the original vectors.

Frequent Asked Questions

How much recall improvement can I expect from TQ+ calibration?

Based on the Turbovec documentation, you can expect an average recall gain of approximately +2.5 points, and up to +8.7 points on highly anisotropic datasets.

Can I calibrate after adding all my vectors?

Technically yes, TurboVect supports you — it re-encodes every stored at the time of the call. However, it is not recommended because you incur a second quantization cost that costs you a few recall points. The best practice is to calibrate before bulk insertion.

What happens if I calibrate with a bad sample?

If the sample is biased (e.g., sorted or clustered), the resulting shift/scale pair will degrade recall. Publicly, a later re-calibrate cannot fix the damage because the coordinates are already clamped. you must rebuild the index from the original uncompressed vectors.

Is the calibration state persisted with the index?

Yes. The TQ+ pair is stored in the file format (.tv / .tvim) and automatically reloaded with the index. Even if you remove all rows (draining the index), the calibration state is preserved — a later rebuild can reuse the same mapping.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →