PLAID vs Voyager Indexes in PyLate: Key Differences and When to Use Each
PLAID uses product-quantization with IVF for memory-efficient search with optional subset filtering, while Voyager implements HNSW-style approximate nearest neighbor search with different tuning parameters and requires separate installation.
When working with late-interaction retrieval models in PyLate (lightonai/pylate), choosing the right index implementation is critical for performance and functionality. The library provides two distinct indexing strategies: PLAID (Product-Lookups for AI and Dense retrieval) and Voyager (an HNSW-based approximate nearest neighbor library). Each follows different algorithmic approaches, offers unique configuration options, and imposes different requirements on your deployment environment.
Core Architectural Differences
PLAID: Product Quantization with IVF
The PLAID index in PyLate implements a product-quantization (PQ) approach combined with an inverted file (IVF) structure. This architecture compresses embeddings into compact codes using configurable nbits parameters, significantly reducing memory footprint while maintaining retrieval speed. The implementation supports fine-tuning through parameters like kmeans_niters, max_points_per_centroid, and n_ivf_probe, which control clustering quality and search depth.
Voyager: HNSW-Style Approximate Nearest Neighbors
Voyager utilizes a graph-based HNSW (Hierarchical Navigable Small World) algorithm for approximate nearest neighbor search. Unlike PLAID's quantization approach, Voyager builds a navigable graph structure using parameters M (number of sub-quantizers/neighbors), ef_construction (build-time search depth), and ef_search (query-time search depth). This approach typically offers different latency-recall tradeoffs compared to IVF-based methods.
Backend Implementations and Installation
PLAID Backend Options
The pylate/indexes/plaid.py file implements a unified interface that selects between two backends via the use_fast flag:
- Fast-Plaid: A high-performance Rust implementation (default) that provides optimized product-quantization and IVF search. This backend is bundled with PyLate and requires no additional dependencies.
- Stanford PLAID: The original Python/NLP implementation (deprecated). This backend is maintained for compatibility but lacks performance optimizations and subset filtering capabilities.
The backend selection occurs in PLAID.__init__, with the use_fast=True default ensuring optimal performance for new projects.
Voyager External Dependency
The pylate/indexes/voyager.py implementation wraps the external Voyager library, which is not included in the base PyLate installation. To use this index, you must install the optional dependency:
pip install "pylate[voyager]"
# or directly
pip install voyager
The Voyager class initializes a voyager.Index object with Space.Cosine and the specified HNSW parameters (M, ef_construction, ef_search).
Key Functional Differences
Embedding Shape Handling
PLAID expects token-level embeddings directly without reshaping, operating on the native output shape from ColBERT-style encoders.
Voyager requires embeddings shaped (batch, n_tokens, dim). The pylate/indexes/voyager.py file includes a reshape_embeddings helper function that automatically adds a batch dimension when needed, ensuring compatibility with single-document inputs.
Document-ID Mapping
PLAID manages document-to-embedding mappings internally within the Fast-Plaid backend, storing IDs alongside the quantized vectors.
Voyager maintains explicit mapping files: document_ids_to_embeddings.pkl and embeddings_to_documents_ids.pkl. These pickle files persist the bidirectional mapping between your string/integer document IDs and Voyager's internal vector indices, as implemented in the mapping I/O methods of pylate/indexes/voyager.py.
Subset Filtering Capabilities
Subset filtering—restricting search to a specific set of document IDs—is supported in PLAID only when use_fast=True. The pylate/indexes/plaid.py file explicitly raises a ValueError if you attempt to use the subset argument with the deprecated Stanford backend.
Voyager does not support subset filtering. The index always searches the entire collection, returning results from all indexed documents regardless of the k parameter.
Query Return Types
PLAID returns a list of RerankResult objects, maintaining consistency with other indexes in the PyLate ecosystem and providing structured access to document IDs and scores.
Voyager returns a dictionary with two keys: "documents_ids" and "distances", each containing nested arrays corresponding to query batches. This raw array format requires different post-processing compared to PLAID's object-oriented results.
Code Examples
Using PLAID with Fast-Plaid Backend
from pylate import indexes, models
# Initialize the PLAID index with Fast-Plaid backend (default)
plaid_idx = indexes.PLAID(
index_folder="plaid_index",
index_name="colbert",
override=True, # Recreate if exists
use_fast=True, # Use Rust implementation (default)
nbits=4, # Product quantization bits
n_ivf_probe=8, # IVF clusters to probe
)
# Load a ColBERT model and encode documents
model = models.ColBERT("lightonai/GTE-ModernColBERT-v1")
doc_emb = model.encode(["Doc 1", "Doc 2"], is_query=False)
# Add documents to index
plaid_idx = plaid_idx.add_documents(
documents_ids=[0, 1],
documents_embeddings=doc_emb,
)
# Query with optional subset filtering
query_emb = model.encode(["search query"], is_query=True)
results = plaid_idx(query_emb, k=5, subset=[0, 1]) # Subset only works with use_fast=True
print(results) # List of RerankResult objects
Using Voyager Index
from pylate import indexes, models
# Initialize Voyager (requires: pip install "pylate[voyager]")
voyager_idx = indexes.Voyager(
index_folder="voyager_index",
index_name="colbert",
override=True,
embedding_size=128,
M=64, # Number of sub-quantizers
ef_construction=200, # Build-time search depth
ef_search=200, # Query-time search depth
)
# Encode documents with batch dimension handling
model = models.ColBERT("sentence-transformers/all-MiniLM-L6-v2")
doc_emb = model.encode(
["Doc A", "Doc B", "Doc C"],
is_query=False,
)
# Add documents (mappings stored in pickle files)
voyager_idx = voyager_idx.add_documents(
documents_ids=["A", "B", "C"],
documents_embeddings=doc_emb,
)
# Query (no subset filtering available)
query_emb = model.encode(["search"], is_query=True)
matches = voyager_idx(query_emb, k=3)
print(matches["documents_ids"]) # Nested list of document IDs
print(matches["distances"]) # Corresponding distances
Summary
- PLAID implements product-quantization with IVF, offering configurable compression via
nbitsand optional subset filtering when using the default Fast-Plaid Rust backend. - Voyager provides HNSW-style graph search with parameters
M,ef_construction, andef_search, requiring separate installation viapip install "pylate[voyager]". - PLAID manages document IDs internally and returns
RerankResultobjects, while Voyager uses explicit pickle files for ID mapping and returns raw dictionary arrays. - Subset filtering is only available in PLAID with
use_fast=True; Voyager always searches the entire index. - Both indexes are implemented in
pylate/indexes/plaid.pyandpylate/indexes/voyager.pyrespectively, following the common interface defined inpylate/indexes/base.py.
Frequently Asked Questions
Which index is faster for large-scale retrieval?
Performance depends on your specific latency-recall requirements and hardware. PLAID's Fast-Plaid backend uses highly optimized Rust implementations of product-quantization, which typically excel in memory-constrained environments with high compression ratios. Voyager's HNSW implementation generally provides faster query times at higher recall levels but consumes more memory. For ColBERT-style late interaction retrieval, PLAID is often preferred due to its native handling of token-level embeddings without reshaping overhead.
Can I use subset filtering with Voyager?
No, subset filtering is not supported in the Voyager index. According to the implementation in pylate/indexes/voyager.py, the index always searches the entire collection and returns results from all indexed documents. If your application requires searching within specific document subsets (such as user-specific corpora or time-bounded collections), you must use the PLAID index with use_fast=True, which is the only configuration that supports the subset argument as implemented in pylate/indexes/plaid.py.
Do I need to install extra dependencies for PLAID?
No, PLAID works out-of-the-box with the base PyLate installation. The default Fast-Plaid backend bundles pre-compiled Rust wheels that handle product-quantization and IVF search without additional dependencies. The alternative Stanford PLAID backend (deprecated) also requires no extra packages as it uses pure Python. This contrasts with Voyager, which explicitly requires pip install "pylate[voyager]" or pip install voyager to access the external C++ library.
Which index should I choose for ColBERT embeddings?
For most ColBERT workflows, PLAID is the recommended choice because it natively handles token-level embeddings without requiring batch dimension reshaping and supports subset filtering for targeted retrieval. The Fast-Plaid backend provides optimized product-quantization specifically designed for late-interaction models. However, choose Voyager if you specifically require HNSW-style graph search characteristics, already depend on the Voyager library in your infrastructure, or need the specific latency-recall tradeoffs that HNSW provides. Note that Voyager requires explicit reshaping of embeddings to (batch, n_tokens, dim) format and maintains document mappings through external pickle files rather than internal storage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →