Length-Renormalized Scoring in Turbovec: How It Unbiaseds Similarity Search
TLDR: Length-renormalized scoring in Turbovec is a per-vector scalar correction that removes the systematic underestimation bias introduced by scalar quantization, restoring unbiased inner-product estimates with zero extra storage and zero runtime overhead.
Turbovec, the Rust-powered vector quantization library from RyanCodrai/turbovec, achieves aggressive compression through random rotation and Lloyd-Max scalar quantization. But that compression introduces a subtle yet critical flaw: quantized vectors are systematically shorter than their originals, which biases similarity scores downward. According to the Turbovec source code, length-renormalized scoring addresses this by storing one scalar per vector during ingest and applying it at search time — turning the biased estimator into an unbiased one.
Why Scalar Quantization Shrinks Vectors
The problem begins during encoding. After Turbovec applies a random rotation to an input vector, it quantizes the components using Lloyd-Max scalar quantization — typically down to 2–8 bits per component. This quantization step is lossy. The reconstructed centroid vector x̂ is close to the original direction but slightly shorter in magnitude because the quantization cells have finite resolution and the centroids rarely lie exactly on the unit sphere.
The result: the inner-product score between a rotated query and the quantized candidate is systematically under-estimated. The estimator is downward-biased, which means recall drops, especially when compression is aggressive. The Turbovec README documents this exact issue in its "How it works" section at README.md#L24-L26, noting that quantization "shortens the reconstructed unit direction compared to the original vector."
The Formula: One Scalar Per Vector
Turbovec solves this at encode time. For each vector v, it computes a per-vector scalar:
s = ||v|| / <u, x̂>
where:
vis the original input vectoruis the rotated unit vector (after random rotation)x̂is the centroid reconstruction produced by the quantizer
This scalar s is stored alongside the compressed vector. Critically, it replaces the stored norm ||v|| — it does not add new per-vector data. So the storage cost stays exactly the same, with zero extra overhead.
Search-Time Application: Where the Scalar Kicks In
The correction operates inside the scoring kernel. When a query comes in, Turbovec's search kernel computes the raw inner product between the rotated query and each codebook value. Before placing a candidate score into the result heap, the kernel multiplies that raw score by the stored scalar for that candidate.
That single multiplication transforms the biased estimate into an unbiased one. In effect, it "renormalizes" the length of the reconstructed vector back to its original magnitude, recovering the cosine inner-product relationship that the quantization destroyed.
Zero Runtime Cost
Because the scalar is precomputed at ingest and stored with the vector, the search-time multiplication is a single floating-point multiply per candidate. There is no extra computation at query time, no re-scoring pass, and no re-query loop. The search kernel simply multiplies the raw score by the scalar as it builds the top-k heap.
Clamping for Cosine Mode
Turbovec also handles a floating-point edge case in the Python Haystack wrapper. When using cosine similarity, length-renormalization can occasionally produce output scores just slightly above 1.0 — an artifact of floating-point error in the estimator. In turbovec-python/python/turbovec/haystack.py at lines 36–42, the _reconstruct method clamps these values back into the cosine range [-1, 1] before scaling them to [0, 1].
# turbovec-python/python/turbovec/haystack.py (excerpt, lines 36–42)
# Clamp the score back into [-1, 1] for cosine mode
score = min(1.0, max(-1.0, score))
This clamping ensures that downstream consumers (such as Haystack pipelines) never see invalid cosine similarity values, keeping the API contract clean.
Practical Example: Using Length-Renormalized Search in Python
You never need to compute the scalar manually — Turbovec does it automatically when you index vectors. Here is a complete example using the Python bindings.
import numpy as np
from turbovec import TurboQuantIndex
# 1️⃣ Encode vectors (length-renormalization is automatic)
dim = 1536
vectors = np.random.randn(10_000, dim).astype(np.float32)
index = TurboQuantIndex(dim, bits=2) # 2-bit compression
index.add(vectors) # Computes per-vector scalar
# 2️⃣ Search – the scalar is applied internally
query = np.random.randn(dim).astype(np.float32)
top_ids, scores = index.search(query, k=5) # Returned scores are unbiased
print("Top-5 IDs:", top_ids)
print("Length-renormalized scores:", scores)
And here is the same flow using the Haystack integration:
# Using the Haystack wrapper (Python) – the same scalar is applied automatically
from turbovec.haystack import TurboQuantDocumentStore
store = TurboQuantDocumentStore(embedding_dim=1536, bits=2)
store.add_texts(["doc A", "doc B", "doc C"], ids=["a", "b", "c"])
# Query with cosine similarity; `scale_score=True` rescales to [0,1]
results = store.similarity_search(
query="example query",
k=3,
embedding_similarity_function="cosine",
scale_score=True,
)
for doc in results:
print(doc.id, doc.score) # Scores already length-renormalized
Why Length-Renormalized Scoring Matters in Practice
The benefit of length normalization is most pronounced at low bit-widths — exactly where aggressive compression dominates. When vectors are quantized to 2 bits, the shrinkage is largest, and the downward bias can crush recall. Length-renormalization recovers most of that lost recall without any trade-off, because the scalar already stored and the multiply is virtually free.
Key takeaways directly from the Turbovec source:
- Zero storage cost: The scalar replaces the original vector norm, so no bytes per vector are added.
- Zero runtime penalty: A single multiply per candidate at search time.
- Biggest recall win at 2-bit compression: The worse the quantization shrink, the bigger the fix.
- Cosine-range safety: The Haystack wrapper clamps to
[-1, 1]to avoid floating-point drift above 1.0.
For the implementation details, look at:
| File | Role |
|---|---|
README.md – How it works section |
Documents the formula and motivation (lines 24–26) |
turbovec-python/python/turbovec/haystack.py |
Shows the clamping logic in _reconstruct (lines 36–42) |
src/lib.rs (Rust core) |
Implements the per-candidate scalar multiplication in the search kernel |
CHANGELOG.md |
Records the feature entry at lines 3377–3399 |
Summary
- Length-renormalized scoring is Turbovec's method of correcting the downward bias introduced by Lloyd-Max scalar quantization.
- It stores one scalar per vector at encode time:
s = ||v|| / <u, x̂>. - At search time, the kernel multiplies the raw score by
sfor every candidate before placing it in the result heap, producing unbiased inner products. - The correction costs zero extra storage and zero extra query-time complexity, while significantly improving recall at low bit-widths (down to 2 bits).
- The Haystack Python wrapper includes a clamping step to keep cosine scores in the valid
[-1, 1]range.
Frequently Asked Questions
What does "length-renormalized scoring" mean in Turbovec?
It means that Turbovec multiplies each raw inner-product score by a per-vector scalar computed during encoding. This scalar corrects for the systematic shortening of vectors caused by scalar quantization, making the similarity estimator unbiased.
Does length-renormalization slow down query time?
No. The scalar is applied as a single multiply-per-candidate operation inside the search kernel, before a score goes into the result heap. It is computed once at the encode time and stored per vector, so there is no additional per-query computation beyond the existing inner product.
When is length-renormalization most important?
It matters most at very low bit-widths, especially 2-bit quantization, where the shrinkage of reconstructed vector length is the largest. Applying the scalar recovers most of the recall that would otherwise be lost to compression bias.
Do I need to enable length-renormalization manually?
No — it is automatic in Turbovec's vector indexing and search pipeline, both in the Rust core and in the Python TurboQuantIndex and TurboQuantDocumentStore wrappers. You simply add() your vectors and the scalar is computed for you.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →