How Length-Renormalized Scoring Removes Quantization Bias in TurboVec Vector Search
Length-renormalized scoring removes quantization bias by storing a per-vector scale factor equal to ‖v‖ / ⟨u, x̂⟩ at encode time and multiplying each raw SIMD inner-product estimate by that factor during search, undoing the systematic shrinkage introduced by Lloyd-Max scalar quantization.
TurboVec is an open-source vector search engine that uses length-renormalized scoring to recover unbiased similarity scores without sacrificing query latency. As implemented in RyanCodrai/turbovec, the correction is computed once per vector during indexing and fused directly into the low-level SIMD kernels that score quantized codes. Because the scale factor is embedded in the existing search-time register operations, the technique removes bias while adding no measurable latency.
The Source of Quantization Bias
Why Lloyd-Max Quantization Shrinks Reconstructed Vectors
TurboVec compresses high-dimensional vectors with a Lloyd-Max codebook. During this process, scalar quantization shortens the reconstructed directions: after normalizing a vector to unit length and applying random rotation, the resulting quantized centroid x̂ is slightly smaller than the original direction v / ‖v‖. This shrinkage is inherent to the quantization process and becomes more severe as bit-width decreases.
The Systematic Inner-Product Under-Estimate
When a query q is scored against a compressed vector using the raw inner product ⟨q, x̂⟩, the reconstruction error produces a systematic under-estimate of the true similarity. Because x̂ is shorter than the original unit direction, the raw score is biased downward regardless of the query vector. TurboVec eliminates this bias with a per-vector multiplier derived from the original magnitude.
Encode-Time Computation in turbovec/src/encode.rs
The length-renormalization factor is computed during indexing inside turbovec/src/encode.rs. Specifically, the function fused_quantize_scale_pack returns the ratio of the original vector norm to the inner product between the true unit direction u and the quantized reconstruction x̂:
s = ‖v‖ / ⟨u, x̂⟩
To avoid division by zero, the implementation clamps the denominator to a small epsilon before computing the scale:
// inside fused_quantize_scale_pack (turbovec/src/encode.rs)
let inner = inner.max(1e-10) as f32;
norm / inner // per-vector scale stored in `scales`
The resulting Vec<f32> of scales is written alongside the packed quantization codes and later loaded into the search index.
Search-Time Application in turbovec/src/search.rs
At query time, the scale is applied inside the SIMD kernels defined in turbovec/src/search.rs. After accumulating the raw inner-product estimate for a block of vectors, each lane is multiplied by the corresponding per-vector vec_scales entry.
In the ARM NEON kernel, the correction looks like this:
// score_4bit_block_neon
let vec_scales_ptr = vec_scales.as_ptr().add(base_vec);
// ...
*out_ptr.add(lane) = float_accum[lane] * *vec_scales_ptr.add(lane);
The x86 kernels follow the identical pattern. For AVX2, the multiplication is expressed with wide SIMD loads:
// avx2_post_flush_heap_update
let s0 = _mm256_mul_ps(fa[0], _mm256_loadu_ps(vec_scales_ptr));
Because the kernels already stream per-vector constants into registers, applying vec_scales adds zero effective overhead beyond a single fused multiply.
Why Length-Renormalized Scoring Works
The stored factor vec_scales[i] is exactly the ratio ‖v_i‖ / ⟨u_i, x̂_i⟩. Multiplying the raw quantized score by this value performs two corrections simultaneously:
- Magnitude restoration — The numerator
‖v‖recovers the original vector length that was lost during normalization. - Shrinkage compensation — The denominator
⟨u, x̂⟩accounts for the reduced length of the reconstructed unit direction.
According to the TurboVec source code, this yields an unbiased estimator of the true inner product while remaining computationally free at search time. The recall improvement is most pronounced at aggressive compression settings such as 2-bit and 4-bit quantization, where centroid reconstruction shrinkage is largest.
Practical Code Examples
Encoding a Batch in Rust
use turbovec::encode::encode;
// vectors: flat f32 slice, n vectors, dim dimensions
let (packed_codes, scales, shift, scale_tq) = encode(
&vectors,
n,
dim,
&rotation,
&boundaries,
¢roids,
bit_width,
None, // no previous calibration
);
// `scales` now contains the length-renormalization factor for each vector
Searching with Python
import turbovec
# Load an index that was built with the Rust encoder
index = turbovec.load_index("my_index.tvec")
# query is a raw f32 vector (same dim as indexed vectors)
scores, ids = index.search(query, k=10)
# Internally the index multiplies each raw inner-product estimate by
# the stored per-vector scale, so the returned `scores` are unbiased.
Inspecting Stored Scales
// After loading an index
let vec_scales = index.vec_scales(); // &[f32] with one entry per stored vector
println!("first vector scale = {}", vec_scales[0]);
Summary
- TurboVec uses Lloyd-Max scalar quantization, which unavoidably shortens reconstructed vectors and biases inner-product scores downward.
- The
fused_quantize_scale_packfunction inturbovec/src/encode.rspre-computes a per-vector scale equal tonorm / inner, i.e.,‖v‖ / ⟨u, x̂⟩. - During search, the SIMD kernels in
turbovec/src/search.rsmultiply each raw accumulated score by the corresponding entry invec_scales, removing the bias. - Because the multiplier is stored with each compressed vector and applied inside existing SIMD pipelines, the correction adds no measurable search-time overhead.
- Length-renormalized scoring delivers the largest recall gains at low bit-widths such as 2-bit and 4-bit.
Frequently Asked Questions
What is quantization bias in vector search?
Quantization bias is a systematic error introduced when compressed vector reconstructions are shorter than the original directions. In TurboVec, Lloyd-Max scalar quantization shrinks the centroid-reconstructed vector x̂, causing raw inner-product scores ⟨q, x̂⟩ to under-estimate the true similarity ⟨q, v⟩ for every query.
How does TurboVec compute the length-renormalization factor?
The factor is computed once per vector during encoding in turbovec/src/encode.rs by the fused_quantize_scale_pack function. It calculates s = ‖v‖ / ⟨u, x̂⟩, where ‖v‖ is the original vector norm and ⟨u, x̂⟩ is the inner product between the true unit direction and its quantized reconstruction.
Does length-renormalized scoring slow down vector search?
No. The per-vector scale is applied inside the same SIMD kernels that compute the raw scores, as seen in turbovec/src/search.rs. The correction requires only a single extra floating-point multiplication per vector, fused into the existing NEON, AVX2, and AVX-512 instruction streams.
Which bit-widths benefit most from length-renormalized scoring?
The improvement is most noticeable at aggressive compression levels, specifically 2-bit and 4-bit quantization. At these low bit-widths, the Lloyd-Max reconstruction error is largest, so the length shrinkage—and therefore the bias removed by renormalization—is greatest.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →