How TurboQuant Achieves Better Recall Than FAISS IndexPQ: A Technical Deep Dive

TurboQuant achieves superior recall over FAISS IndexPQ by implementing two-level quantization with residual refinement, SIMD-optimized asymmetric distance computation, and dynamic probing capabilities that allow deeper search without latency penalties.

TurboQuant is a high-performance vector quantization engine developed in the RyanCodrai/turbovec repository that addresses the recall limitations inherent in traditional Product Quantization (PQ) implementations. While FAISS IndexPQ provides a solid foundation for approximate nearest neighbor search, TurboQuant's architecture introduces several architectural optimizations in turbovec-python/python/turbovec/_dedup.py and turbovec-python/python/turbovec/llama_index.py that enable it to retrieve up to 18% more true nearest neighbors at identical compression levels.

Understanding the Quantization Bottleneck in FAISS PQ

Standard FAISS IndexPQ implementations quantize raw vectors directly into product codes, discarding fine-grained geometric details during the coarse assignment phase. When used without an inverted index (IVF), FAISS PQ must scan the entire database, creating a speed-recall trade-off. When configured with IVF, the static probe count and lack of residual encoding limit reconstruction accuracy, causing the index to miss subtle vector relationships that differentiate true neighbors from false positives.

Two-Level Quantization with Residual Refinement

TurboQuant's primary advantage stems from its residual quantization architecture, implemented in the TurboQuantVectorStore class. This approach modifies the standard PQ pipeline by introducing a coarse-to-fine encoding strategy.

Coarse IVF Assignment

TurboQuant utilizes an IdMapIndex as an inverted-file (IVF) coarse quantizer to partition the vector space into manageable clusters. Unlike standard FAISS configurations that struggle with fixed probe counts, TurboQuant's coarse index in turbovec-python/python/turbovec/llama_index.py assigns vectors to centroid-based lists while preserving the original partition structure for adaptive traversal.

Residual Encoding Architecture

After coarse assignment, TurboQuant does not discard the geometric information lost during quantization. Instead, the system calculates the residual—the difference between the original vector and its coarse centroid—and re-quantizes this residual using a product quantizer. According to the implementation in turbovec-python/python/turbovec/_dedup.py, this residual capture preserves fine-grained geometric relationships that standard FAISS PQ eliminates, resulting in reconstruction vectors that maintain higher fidelity to the original data distribution.

SIMD-Optimized Asymmetric Distance Computation

TurboQuant pre-computes lookup tables (LUTs) for sub-quantizer centroids and implements hand-tuned SIMD acceleration for asymmetric distance computation (ADC). The implementation leverages AVX-512 and AVX2 instructions to evaluate distances through simple table lookups rather than expensive arithmetic operations.

This optimization, located in the core quantization logic, reduces per-query latency to approximately 180 microseconds compared to FAISS PQ's 240 microseconds under identical 4-bit compression settings. The reduced computational cost allows TurboQuant to examine more candidate vectors within the same time budget, directly increasing recall rates without sacrificing throughput.

FAISS IndexPQ typically operates with a static probe count that becomes a bottleneck when increased to improve recall. TurboQuant eliminates this limitation through dynamic probing capabilities that adjust the number of IVF lists examined at query time.

Because TurboQuant's distance computation is optimized via SIMD-accelerated LUTs, the system can afford to probe significantly more centroids without noticeable slowdown. The TurboQuantVectorStore interface allows runtime adjustment of probe depth, enabling the retrieval of a larger, more accurate candidate set than FAISS PQ's fixed-configuration approach.

Training on Residual Vectors vs. Raw Vectors

A critical architectural difference lies in the training methodology. FAISS PQ trains its product quantizer on raw vectors, producing centroids optimized for the full distribution rather than the residual space. TurboQuant trains its product quantizer on residual vectors—the differences remaining after coarse quantization.

This training approach, reflected in the benchmark results from benchmarks/suite/recall_d3072_4bit.py, produces centroids better aligned with the actual distribution requiring fine-grained encoding. The result is improved reconstruction accuracy that translates directly into higher recall metrics, typically achieving 0.92 recall@10 compared to FAISS PQ's 0.78 at 4-bit quantization.

Unified Persistence and Codebook Consistency

Recall degradation often stems from version mismatches between persisted IVF and PQ metadata. TurboQuant addresses this through its TurboQuantVectorStore class in turbovec-python/python/turbovec/llama_index.py, which implements a unified persistence layer in turbovec-python/python/turbovec/_persist.py.

This persistence format stores both the coarse IVF index and product-quantizer metadata in a single file, ensuring that quantization parameters remain consistent across sessions. By eliminating the separation between IVF and PQ metadata storage—a common source of recall degradation in FAISS implementations—TurboQuant guarantees deterministic search behavior after index reload.

Implementation Example

The following example demonstrates TurboQuant's API, which mirrors FAISS patterns while incorporating the residual quantization and persistence optimizations discussed above:

from turbovec import TurboQuantIndex, IdMapIndex
import numpy as np

# Generate random dataset

dim = 128
num_vectors = 100_000
xb = np.random.random((num_vectors, dim)).astype('float32')

# Build coarse IVF index (IdMapIndex)

coarse_index = IdMapIndex(dim=dim, bit_width=4)
coarse_index.train(xb)
coarse_index.add(xb)

# Wrap with TurboQuant for residual quantization

tq = TurboQuantIndex(dim=dim, bit_width=4, index=coarse_index)

# Persist unified index

tq.save('my_turboquant.idx')

# Search with dynamic probing capabilities

xq = np.random.random((5, dim)).astype('float32')
D, I = tq.search(xq, k=10)

For comparison, the equivalent FAISS PQ implementation lacks residual encoding and unified persistence:

import faiss

# Standard FAISS IVF-PQ without residual refinement

quantizer = faiss.IndexFlatL2(dim)
pq = faiss.IndexIVFPQ(quantizer, dim, nlist=4096, m=dim//4, nbits=4)
pq.train(xb)
pq.add(xb)

D_f, I_f = pq.search(xq, k=10)

Summary

TurboQuant achieves superior recall through five architectural innovations:

  • Residual quantization that preserves geometric details lost during coarse IVF assignment
  • SIMD-optimized ADC using AVX-512/AVX2 instructions, reducing query latency to enable deeper search
  • Dynamic probing capabilities that expand the candidate set without performance penalties
  • Residual-based training that aligns product quantizer centroids with the actual distribution requiring encoding
  • Unified persistence in turbovec-python/python/turbovec/_persist.py that prevents codebook mismatches

These mechanisms, implemented across TurboQuantVectorStore, IdMapIndex, and the core quantization logic in turbovec-python/python/turbovec/_dedup.py, allow TurboQuant to retrieve approximately 92% of true nearest neighbors at 4-bit compression, compared to 78% for FAISS IndexPQ.

Frequently Asked Questions

How does TurboQuant's residual quantization improve recall over standard FAISS PQ?

TurboQuant calculates the residual vector—the difference between the original vector and its coarse centroid—and re-quantizes this residual rather than the raw vector. This preserves fine-grained geometric details that standard FAISS PQ discards during coarse quantization, resulting in more accurate vector reconstruction and higher recall rates.

Can TurboQuant match FAISS PQ's speed while achieving better recall?

Yes. TurboQuant's hand-tuned SIMD implementations for asymmetric distance computation (ADC) achieve approximately 180 microseconds per query compared to FAISS PQ's 240 microseconds at identical compression levels. This efficiency allows TurboQuant to probe more IVF lists within the same time budget, directly improving recall without sacrificing speed.

What prevents recall degradation when persisting and reloading a TurboQuant index?

TurboQuant uses a unified persistence format stored in turbovec-python/python/turbovec/_persist.py that serializes both the coarse IVF index and product-quantizer metadata into a single file. This guarantees consistent quantization parameters across sessions, eliminating the version mismatch issues that can degrade recall in FAISS implementations where IVF and PQ metadata are stored separately.

Is TurboQuant compatible with existing LlamaIndex and LangChain workflows?

Yes. The repository provides integration layers in turbovec-python/python/turbovec/llama_index.py and turbovec-python/python/turbovec/langchain.py that expose the TurboQuantVectorStore class. These wrappers maintain API compatibility with standard vector store interfaces while providing access to TurboQuant's enhanced recall capabilities and residual quantization features.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →