How the TQ+ Calibration Trailer Enables Online Calibration in Turbovec

The TQ+ calibration trailer stores per-coordinate shift and scale parameters at the end of the index file, allowing Turbovec to fit calibration on the first batch of vectors and automatically reuse it across process restarts without offline retraining.

TurboQuant+ (TQ+) is a quantization method implemented in the RyanCodrai/turbovec repository that learns per-coordinate shift and scale parameters to map rotated coordinates onto a canonical Beta distribution. Unlike traditional quantization approaches requiring offline calibration datasets, Turbovec's TQ+ calibration trailer enables online calibration by persisting these parameters directly in the index file format. This design allows the system to fit calibration on the first batch of added vectors and seamlessly reuse those parameters for all future additions and queries.

What Is TQ+ Calibration?

TQ+ improves upon standard quantization by learning per-coordinate shift and scale parameters that map the empirical distribution of each rotated coordinate onto the canonical Beta-marginal used to train the Lloyd-Max codebook. This coordinate-wise transformation ensures that the quantized values better match the expected distribution, reducing quantization error. In turbovec/src/encode.rs, the compute_tqplus_calibration function (lines 36-45) fits these parameters by analyzing the first batch of vectors added to the index.

How the TQ+ Calibration Trailer Works

The online calibration capability relies on three complementary mechanisms baked into the indexing pipeline and file format:

Fitting Calibration on First Add

When a batch of vectors is added for the first time, TurboQuantIndex::add calls encode::encode with existing_calibration = None. Inside encode.rs (lines 36-45), the compute_tqplus_calibration function fits a fresh (shift, scale) pair per dimension. The returned vectors are stored in the index fields tqplus_shift and tqplus_scale (defined in turbovec/src/lib.rs, lines 107-112).

Persisting Parameters in the Trailer

The fitted calibration survives process restarts through a dedicated trailer appended to the index file. In turbovec/src/io.rs, the write_core function (lines 97-112) appends the TQ+ trailer after the per-vector scales, writing the length (n_calib) followed by the two float arrays containing the shift and scale values. This binary format ensures the calibration travels with the index data on disk.

Reusing Calibration on Subsequent Adds

When a file is later loaded, io::read_core_v3 (lines 44-58) reads the trailer and restores the same calibration vectors into memory. On any subsequent call to add, the index detects that tqplus_shift is non-empty and constructs existing = Some((shift, scale)) (as seen in lib.rs, lines 66-73). The encoder then locks the calibration (lines 89-95 in encode.rs), ensuring that all vectors—both old and new—share the same coordinate system without refitting.

The Binary Trailer Format

The TQ+ calibration trailer occupies the end of the Turbovec index file. According to the source code in io.rs, the trailer consists of:

  • A length field (n_calib) indicating the number of calibrated dimensions
  • A float array for the shift parameters
  • A float array for the scale parameters

This append-only design allows the calibration metadata to be written atomically with the index data and read back efficiently when the index is loaded via TurboQuantIndex::load.

Practical Implementation

The following Rust code demonstrates the complete workflow, from initial creation through persistence and reuse:

// 1️⃣ Create a new index (lazy dim) and add the first batch
let mut idx = TurboQuantIndex::new_lazy(4).unwrap(); // 4‑bit quantization
idx.add_2d(&first_batch, dim).unwrap(); // fits TQ+ calibration internally

// 2️⃣ Save the index – the calibration trailer is written automatically
idx.write("my_index.tv").unwrap();

// 3️⃣ Load the index later – calibration is restored from the trailer
let mut loaded = TurboQuantIndex::load("my_index.tv").unwrap();

// 4️⃣ Add more vectors – the same calibration is reused (no re‑fit)
loaded.add_2d(&second_batch, dim).unwrap();

// 5️⃣ Search – query vectors are un‑calibrated on the fly, using the stored shift/scale
let results = loaded.search(&queries, 10);

Because the calibration vectors travel with the index on disk, as implemented in RyanCodrai/turbovec, a freshly loaded index automatically knows the exact (shift, scale) that was used when the vectors were originally encoded. This eliminates any need for an offline "re-train" step and lets the system calibrate online: the moment the first vectors are added, the calibration is created, saved, and later reused without recomputation.

Summary

  • Online fitting: The compute_tqplus_calibration function in encode.rs fits shift and scale parameters during the first call to add when no existing calibration is present.
  • Persistent storage: The TQ+ calibration trailer written by io::write_core (lines 97-112) stores these parameters at the end of the index file.
  • Automatic restoration: io::read_core_v3 (lines 44-58) reads the trailer when loading an index, restoring calibration into tqplus_shift and tqplus_scale.
  • Locked reuse: Subsequent calls to add detect existing calibration in lib.rs (lines 66-73) and lock it during encoding (lines 89-95), ensuring all vectors share the same coordinate system.
  • Process independence: The trailer enables online calibration across process restarts without requiring access to the original calibration dataset.

Frequently Asked Questions

Where is the TQ+ calibration trailer stored in the index file?

The trailer is appended at the very end of the file after the per-vector scales. In turbovec/src/io.rs, the write_core function writes the length field followed by the shift and scale float arrays (lines 97-112), making it the final component of the binary format.

What happens if I add vectors to an existing index that already has calibration?

The index checks if tqplus_shift is non-empty in lib.rs (lines 66-73). If calibration exists, it constructs existing = Some((shift, scale)) and passes it to the encoder. The encoder then locks these parameters (lines 89-95 in encode.rs) and applies them to new vectors without refitting, ensuring consistency with previously stored vectors.

Does Turbovec require an offline calibration dataset?

No. Turbovec performs online calibration by fitting the shift and scale parameters on the first batch of vectors added to the index. The compute_tqplus_calibration function in encode.rs analyzes the empirical distribution of the first add operation and stores the results in the TQ+ trailer for all future use.

How does the trailer ensure calibration survives process restarts?

When TurboQuantIndex::load is called, io::read_core_v3 (lines 44-58) reads the TQ+ trailer from the end of the file and restores the shift and scale vectors into the index instance. This allows the loaded index to immediately use the same calibration parameters that were fitted during the initial creation, without requiring access to the original training data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →