# PLAID vs Voyager Indexes in PyLate: Key Differences and When to Use Each

> Discover the key differences between PLAID and Voyager indexes in PyLate. Learn when to use each for efficient similarity search in your AI applications.

- Repository: [LightOn/pylate](https://github.com/lightonai/pylate)
- Tags: deep-dive
- Published: 2026-03-06

---

**PLAID uses product-quantization with IVF for memory-efficient search with optional subset filtering, while Voyager implements HNSW-style approximate nearest neighbor search with different tuning parameters and requires separate installation.**

When working with late-interaction retrieval models in PyLate (`lightonai/pylate`), choosing the right index implementation is critical for performance and functionality. The library provides two distinct indexing strategies: **PLAID** (Product-Lookups for AI and Dense retrieval) and **Voyager** (an HNSW-based approximate nearest neighbor library). Each follows different algorithmic approaches, offers unique configuration options, and imposes different requirements on your deployment environment.

## Core Architectural Differences

### PLAID: Product Quantization with IVF

The PLAID index in PyLate implements a product-quantization (PQ) approach combined with an inverted file (IVF) structure. This architecture compresses embeddings into compact codes using configurable `nbits` parameters, significantly reducing memory footprint while maintaining retrieval speed. The implementation supports fine-tuning through parameters like `kmeans_niters`, `max_points_per_centroid`, and `n_ivf_probe`, which control clustering quality and search depth.

### Voyager: HNSW-Style Approximate Nearest Neighbors

Voyager utilizes a graph-based HNSW (Hierarchical Navigable Small World) algorithm for approximate nearest neighbor search. Unlike PLAID's quantization approach, Voyager builds a navigable graph structure using parameters `M` (number of sub-quantizers/neighbors), `ef_construction` (build-time search depth), and `ef_search` (query-time search depth). This approach typically offers different latency-recall tradeoffs compared to IVF-based methods.

## Backend Implementations and Installation

### PLAID Backend Options

The [`pylate/indexes/plaid.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/plaid.py) file implements a unified interface that selects between two backends via the `use_fast` flag:

- **Fast-Plaid**: A high-performance Rust implementation (default) that provides optimized product-quantization and IVF search. This backend is bundled with PyLate and requires no additional dependencies.
- **Stanford PLAID**: The original Python/NLP implementation (deprecated). This backend is maintained for compatibility but lacks performance optimizations and subset filtering capabilities.

The backend selection occurs in `PLAID.__init__`, with the `use_fast=True` default ensuring optimal performance for new projects.

### Voyager External Dependency

The [`pylate/indexes/voyager.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/voyager.py) implementation wraps the external Voyager library, which is not included in the base PyLate installation. To use this index, you must install the optional dependency:

```bash
pip install "pylate[voyager]"

# or directly

pip install voyager

```

The Voyager class initializes a `voyager.Index` object with `Space.Cosine` and the specified HNSW parameters (`M`, `ef_construction`, `ef_search`).

## Key Functional Differences

### Embedding Shape Handling

PLAID expects token-level embeddings directly without reshaping, operating on the native output shape from ColBERT-style encoders.

Voyager requires embeddings shaped `(batch, n_tokens, dim)`. The [`pylate/indexes/voyager.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/voyager.py) file includes a `reshape_embeddings` helper function that automatically adds a batch dimension when needed, ensuring compatibility with single-document inputs.

### Document-ID Mapping

PLAID manages document-to-embedding mappings internally within the Fast-Plaid backend, storing IDs alongside the quantized vectors.

Voyager maintains explicit mapping files: `document_ids_to_embeddings.pkl` and `embeddings_to_documents_ids.pkl`. These pickle files persist the bidirectional mapping between your string/integer document IDs and Voyager's internal vector indices, as implemented in the mapping I/O methods of [`pylate/indexes/voyager.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/voyager.py).

### Subset Filtering Capabilities

Subset filtering—restricting search to a specific set of document IDs—is supported in PLAID **only** when `use_fast=True`. The [`pylate/indexes/plaid.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/plaid.py) file explicitly raises a `ValueError` if you attempt to use the `subset` argument with the deprecated Stanford backend.

Voyager does **not** support subset filtering. The index always searches the entire collection, returning results from all indexed documents regardless of the `k` parameter.

### Query Return Types

PLAID returns a list of `RerankResult` objects, maintaining consistency with other indexes in the PyLate ecosystem and providing structured access to document IDs and scores.

Voyager returns a dictionary with two keys: `"documents_ids"` and `"distances"`, each containing nested arrays corresponding to query batches. This raw array format requires different post-processing compared to PLAID's object-oriented results.

## Code Examples

### Using PLAID with Fast-Plaid Backend

```python
from pylate import indexes, models

# Initialize the PLAID index with Fast-Plaid backend (default)

plaid_idx = indexes.PLAID(
    index_folder="plaid_index",
    index_name="colbert",
    override=True,          # Recreate if exists

    use_fast=True,         # Use Rust implementation (default)

    nbits=4,               # Product quantization bits

    n_ivf_probe=8,         # IVF clusters to probe

)

# Load a ColBERT model and encode documents

model = models.ColBERT("lightonai/GTE-ModernColBERT-v1")
doc_emb = model.encode(["Doc 1", "Doc 2"], is_query=False)

# Add documents to index

plaid_idx = plaid_idx.add_documents(
    documents_ids=[0, 1],
    documents_embeddings=doc_emb,
)

# Query with optional subset filtering

query_emb = model.encode(["search query"], is_query=True)
results = plaid_idx(query_emb, k=5, subset=[0, 1])  # Subset only works with use_fast=True

print(results)  # List of RerankResult objects

```

### Using Voyager Index

```python
from pylate import indexes, models

# Initialize Voyager (requires: pip install "pylate[voyager]")

voyager_idx = indexes.Voyager(
    index_folder="voyager_index",
    index_name="colbert",
    override=True,
    embedding_size=128,
    M=64,                  # Number of sub-quantizers

    ef_construction=200,   # Build-time search depth

    ef_search=200,         # Query-time search depth

)

# Encode documents with batch dimension handling

model = models.ColBERT("sentence-transformers/all-MiniLM-L6-v2")
doc_emb = model.encode(
    ["Doc A", "Doc B", "Doc C"],
    is_query=False,
)

# Add documents (mappings stored in pickle files)

voyager_idx = voyager_idx.add_documents(
    documents_ids=["A", "B", "C"],
    documents_embeddings=doc_emb,
)

# Query (no subset filtering available)

query_emb = model.encode(["search"], is_query=True)
matches = voyager_idx(query_emb, k=3)
print(matches["documents_ids"])   # Nested list of document IDs

print(matches["distances"])       # Corresponding distances

```

## Summary

- **PLAID** implements product-quantization with IVF, offering configurable compression via `nbits` and optional subset filtering when using the default Fast-Plaid Rust backend.
- **Voyager** provides HNSW-style graph search with parameters `M`, `ef_construction`, and `ef_search`, requiring separate installation via `pip install "pylate[voyager]"`.
- PLAID manages document IDs internally and returns `RerankResult` objects, while Voyager uses explicit pickle files for ID mapping and returns raw dictionary arrays.
- Subset filtering is only available in PLAID with `use_fast=True`; Voyager always searches the entire index.
- Both indexes are implemented in [`pylate/indexes/plaid.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/plaid.py) and [`pylate/indexes/voyager.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/voyager.py) respectively, following the common interface defined in [`pylate/indexes/base.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/base.py).

## Frequently Asked Questions

### Which index is faster for large-scale retrieval?

Performance depends on your specific latency-recall requirements and hardware. PLAID's Fast-Plaid backend uses highly optimized Rust implementations of product-quantization, which typically excel in memory-constrained environments with high compression ratios. Voyager's HNSW implementation generally provides faster query times at higher recall levels but consumes more memory. For ColBERT-style late interaction retrieval, PLAID is often preferred due to its native handling of token-level embeddings without reshaping overhead.

### Can I use subset filtering with Voyager?

No, subset filtering is not supported in the Voyager index. According to the implementation in [`pylate/indexes/voyager.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/voyager.py), the index always searches the entire collection and returns results from all indexed documents. If your application requires searching within specific document subsets (such as user-specific corpora or time-bounded collections), you must use the PLAID index with `use_fast=True`, which is the only configuration that supports the `subset` argument as implemented in [`pylate/indexes/plaid.py`](https://github.com/lightonai/pylate/blob/main/pylate/indexes/plaid.py).

### Do I need to install extra dependencies for PLAID?

No, PLAID works out-of-the-box with the base PyLate installation. The default Fast-Plaid backend bundles pre-compiled Rust wheels that handle product-quantization and IVF search without additional dependencies. The alternative Stanford PLAID backend (deprecated) also requires no extra packages as it uses pure Python. This contrasts with Voyager, which explicitly requires `pip install "pylate[voyager]"` or `pip install voyager` to access the external C++ library.

### Which index should I choose for ColBERT embeddings?

For most ColBERT workflows, **PLAID** is the recommended choice because it natively handles token-level embeddings without requiring batch dimension reshaping and supports subset filtering for targeted retrieval. The Fast-Plaid backend provides optimized product-quantization specifically designed for late-interaction models. However, choose **Voyager** if you specifically require HNSW-style graph search characteristics, already depend on the Voyager library in your infrastructure, or need the specific latency-recall tradeoffs that HNSW provides. Note that Voyager requires explicit reshaping of embeddings to `(batch, n_tokens, dim)` format and maintains document mappings through external pickle files rather than internal storage.