PLAID vs Voyager Indexes in PyLate: Key Differences and When to Use Each

PLAID uses product-quantization with IVF for memory-efficient search with optional subset filtering, while Voyager implements HNSW-style approximate nearest neighbor search with different tuning parameters and requires separate installation.

When working with late-interaction retrieval models in PyLate (lightonai/pylate), choosing the right index implementation is critical for performance and functionality. The library provides two distinct indexing strategies: PLAID (Product-Lookups for AI and Dense retrieval) and Voyager (an HNSW-based approximate nearest neighbor library). Each follows different algorithmic approaches, offers unique configuration options, and imposes different requirements on your deployment environment.

Core Architectural Differences

PLAID: Product Quantization with IVF

The PLAID index in PyLate implements a product-quantization (PQ) approach combined with an inverted file (IVF) structure. This architecture compresses embeddings into compact codes using configurable nbits parameters, significantly reducing memory footprint while maintaining retrieval speed. The implementation supports fine-tuning through parameters like kmeans_niters, max_points_per_centroid, and n_ivf_probe, which control clustering quality and search depth.

Voyager: HNSW-Style Approximate Nearest Neighbors

Voyager utilizes a graph-based HNSW (Hierarchical Navigable Small World) algorithm for approximate nearest neighbor search. Unlike PLAID's quantization approach, Voyager builds a navigable graph structure using parameters M (number of sub-quantizers/neighbors), ef_construction (build-time search depth), and ef_search (query-time search depth). This approach typically offers different latency-recall tradeoffs compared to IVF-based methods.

Backend Implementations and Installation

PLAID Backend Options

The pylate/indexes/plaid.py file implements a unified interface that selects between two backends via the use_fast flag:

  • Fast-Plaid: A high-performance Rust implementation (default) that provides optimized product-quantization and IVF search. This backend is bundled with PyLate and requires no additional dependencies.
  • Stanford PLAID: The original Python/NLP implementation (deprecated). This backend is maintained for compatibility but lacks performance optimizations and subset filtering capabilities.

The backend selection occurs in PLAID.__init__, with the use_fast=True default ensuring optimal performance for new projects.

Voyager External Dependency

The pylate/indexes/voyager.py implementation wraps the external Voyager library, which is not included in the base PyLate installation. To use this index, you must install the optional dependency:

pip install "pylate[voyager]"

# or directly

pip install voyager

The Voyager class initializes a voyager.Index object with Space.Cosine and the specified HNSW parameters (M, ef_construction, ef_search).

Key Functional Differences

Embedding Shape Handling

PLAID expects token-level embeddings directly without reshaping, operating on the native output shape from ColBERT-style encoders.

Voyager requires embeddings shaped (batch, n_tokens, dim). The pylate/indexes/voyager.py file includes a reshape_embeddings helper function that automatically adds a batch dimension when needed, ensuring compatibility with single-document inputs.

Document-ID Mapping

PLAID manages document-to-embedding mappings internally within the Fast-Plaid backend, storing IDs alongside the quantized vectors.

Voyager maintains explicit mapping files: document_ids_to_embeddings.pkl and embeddings_to_documents_ids.pkl. These pickle files persist the bidirectional mapping between your string/integer document IDs and Voyager's internal vector indices, as implemented in the mapping I/O methods of pylate/indexes/voyager.py.

Subset Filtering Capabilities

Subset filtering—restricting search to a specific set of document IDs—is supported in PLAID only when use_fast=True. The pylate/indexes/plaid.py file explicitly raises a ValueError if you attempt to use the subset argument with the deprecated Stanford backend.

Voyager does not support subset filtering. The index always searches the entire collection, returning results from all indexed documents regardless of the k parameter.

Query Return Types

PLAID returns a list of RerankResult objects, maintaining consistency with other indexes in the PyLate ecosystem and providing structured access to document IDs and scores.

Voyager returns a dictionary with two keys: "documents_ids" and "distances", each containing nested arrays corresponding to query batches. This raw array format requires different post-processing compared to PLAID's object-oriented results.

Code Examples

Using PLAID with Fast-Plaid Backend

from pylate import indexes, models

# Initialize the PLAID index with Fast-Plaid backend (default)

plaid_idx = indexes.PLAID(
    index_folder="plaid_index",
    index_name="colbert",
    override=True,          # Recreate if exists

    use_fast=True,         # Use Rust implementation (default)

    nbits=4,               # Product quantization bits

    n_ivf_probe=8,         # IVF clusters to probe

)

# Load a ColBERT model and encode documents

model = models.ColBERT("lightonai/GTE-ModernColBERT-v1")
doc_emb = model.encode(["Doc 1", "Doc 2"], is_query=False)

# Add documents to index

plaid_idx = plaid_idx.add_documents(
    documents_ids=[0, 1],
    documents_embeddings=doc_emb,
)

# Query with optional subset filtering

query_emb = model.encode(["search query"], is_query=True)
results = plaid_idx(query_emb, k=5, subset=[0, 1])  # Subset only works with use_fast=True

print(results)  # List of RerankResult objects

Using Voyager Index

from pylate import indexes, models

# Initialize Voyager (requires: pip install "pylate[voyager]")

voyager_idx = indexes.Voyager(
    index_folder="voyager_index",
    index_name="colbert",
    override=True,
    embedding_size=128,
    M=64,                  # Number of sub-quantizers

    ef_construction=200,   # Build-time search depth

    ef_search=200,         # Query-time search depth

)

# Encode documents with batch dimension handling

model = models.ColBERT("sentence-transformers/all-MiniLM-L6-v2")
doc_emb = model.encode(
    ["Doc A", "Doc B", "Doc C"],
    is_query=False,
)

# Add documents (mappings stored in pickle files)

voyager_idx = voyager_idx.add_documents(
    documents_ids=["A", "B", "C"],
    documents_embeddings=doc_emb,
)

# Query (no subset filtering available)

query_emb = model.encode(["search"], is_query=True)
matches = voyager_idx(query_emb, k=3)
print(matches["documents_ids"])   # Nested list of document IDs

print(matches["distances"])       # Corresponding distances

Summary

  • PLAID implements product-quantization with IVF, offering configurable compression via nbits and optional subset filtering when using the default Fast-Plaid Rust backend.
  • Voyager provides HNSW-style graph search with parameters M, ef_construction, and ef_search, requiring separate installation via pip install "pylate[voyager]".
  • PLAID manages document IDs internally and returns RerankResult objects, while Voyager uses explicit pickle files for ID mapping and returns raw dictionary arrays.
  • Subset filtering is only available in PLAID with use_fast=True; Voyager always searches the entire index.
  • Both indexes are implemented in pylate/indexes/plaid.py and pylate/indexes/voyager.py respectively, following the common interface defined in pylate/indexes/base.py.

Frequently Asked Questions

Which index is faster for large-scale retrieval?

Performance depends on your specific latency-recall requirements and hardware. PLAID's Fast-Plaid backend uses highly optimized Rust implementations of product-quantization, which typically excel in memory-constrained environments with high compression ratios. Voyager's HNSW implementation generally provides faster query times at higher recall levels but consumes more memory. For ColBERT-style late interaction retrieval, PLAID is often preferred due to its native handling of token-level embeddings without reshaping overhead.

Can I use subset filtering with Voyager?

No, subset filtering is not supported in the Voyager index. According to the implementation in pylate/indexes/voyager.py, the index always searches the entire collection and returns results from all indexed documents. If your application requires searching within specific document subsets (such as user-specific corpora or time-bounded collections), you must use the PLAID index with use_fast=True, which is the only configuration that supports the subset argument as implemented in pylate/indexes/plaid.py.

Do I need to install extra dependencies for PLAID?

No, PLAID works out-of-the-box with the base PyLate installation. The default Fast-Plaid backend bundles pre-compiled Rust wheels that handle product-quantization and IVF search without additional dependencies. The alternative Stanford PLAID backend (deprecated) also requires no extra packages as it uses pure Python. This contrasts with Voyager, which explicitly requires pip install "pylate[voyager]" or pip install voyager to access the external C++ library.

Which index should I choose for ColBERT embeddings?

For most ColBERT workflows, PLAID is the recommended choice because it natively handles token-level embeddings without requiring batch dimension reshaping and supports subset filtering for targeted retrieval. The Fast-Plaid backend provides optimized product-quantization specifically designed for late-interaction models. However, choose Voyager if you specifically require HNSW-style graph search characteristics, already depend on the Voyager library in your infrastructure, or need the specific latency-recall tradeoffs that HNSW provides. Note that Voyager requires explicit reshaping of embeddings to (batch, n_tokens, dim) format and maintains document mappings through external pickle files rather than internal storage.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →