Strategies for Optimizing Retrieval Performance with Large Document Collections in PyLate

Use the FastPlaid backend with tuned product quantization (nbits) and inverted file probes (n_ivf_probe) to achieve sub-second retrieval across millions of token-level embeddings while balancing index size and recall.

PyLate is a flexible retrieval library from LightOnAI that separates indexing from query-time scoring to handle large-scale document collections efficiently. When working with millions of token-level embeddings, performance depends on selecting the right index backend, tuning quantization parameters, and optimizing batch processing. This guide covers concrete strategies implemented in the lightonai/pylate repository to maximize retrieval speed without sacrificing accuracy.

Choose the FastPlaid Backend for Production Workloads

PyLate offers two backends in pylate/indexes/plaid.py: FastPlaid (Rust-based, SIMD-optimized) and the older Stanford PLAID (Python implementation). For large document collections, FastPlaid provides dramatically better throughput and memory efficiency.

FastPlaid implements IVF-PQ (Inverted File with Product Quantization) search in Rust, enabling subset filtering and parallel batch processing that the Python backend cannot match. When initializing your index, explicitly set use_fast=True to ensure you leverage the optimized implementation:

from pylate import indexes

index = indexes.PLAID(
    index_folder="my_index",
    index_name="colbert_fast",
    use_fast=True,  # Critical for performance

    nbits=4,
    n_ivf_probe=16,
)

Tune Product Quantization Bits (nbits) for Size vs. Accuracy

The nbits parameter in pylate/indexes/fast_plaid.py controls the bit depth of product quantization, directly impacting index size, search latency, and recall. Lower bit values create smaller indexes with faster distance calculations but may reduce retrieval accuracy.

Consider these trade-offs when configuring your index:

  • nbits=8: Larger index footprint, fastest search, approximately 70% recall
  • nbits=4: Balanced size and speed, approximately 85% recall (recommended default)
  • nbits=2: Smallest index, slowest search due to increased quantization error, approximately 92% recall but requires more probes to achieve it

Adjust nbits during index construction in PLAID.__init__ and validate against your specific corpus to find the optimal balance.

Optimize Inverted File Probes (n_ivf_probe) for Recall

The n_ivf_probe parameter determines how many inverted file lists FastPlaid examines during the initial candidate retrieval phase. More probes increase recall by searching deeper into the index but consume additional CPU or GPU resources.

In pylate/indexes/fast_plaid.py, the IVF search step uses this parameter to limit the candidate pool before exact scoring. For high-recall scenarios, increase n_ivf_probe to 32 or 64, while latency-sensitive applications may use 4-8 probes at the cost of lower recall.

Align k and k_token for Accurate Top-K Results

PyLate distinguishes between k (final results returned) and k_token (token-level candidates retrieved before re-ranking). The ColBERT.retrieve method in pylate/retrieve/colbert.py automatically raises k_token to match k when the former is smaller, but manual tuning provides better control.

For extreme recall requirements, set k_token substantially larger than k (e.g., k=10, k_token=500) and let the re-ranker in pylate/rank/rank.py prune to the final set using exact dot-product calculations. This two-stage approach balances efficiency with accuracy.

Adjust Batch Size and Device Placement for Hardware

Memory usage and throughput depend heavily on batch_size and device selection. The iter_batch utility in pylate/utils/iter_batch.py splits large query lists into manageable chunks, while FastPlaid accepts device arguments for GPU acceleration.

On GPUs with 16GB VRAM, use batch_size=128 for encoding and retrieval. For CPU-only machines, reduce to batch_size=32 to prevent memory pressure. Always specify device="cuda" when available, as FastPlaid defaults to CUDA if present but falls back gracefully to CPU.

Leverage Subset Filtering for Targeted Searches

When searching specific document collections (e.g., a single tenant’s data), use subset filtering available exclusively in FastPlaid. The FastPlaid.__call__ method in pylate/indexes/fast_plaid.py accepts a subset parameter that restricts the search to pre-selected document IDs.

This avoids scanning the entire corpus, reducing latency from seconds to milliseconds for targeted queries. Note that subset filtering is not supported by the Stanford PLAID backend.

Persist ID Mappings for Incremental Updates

FastPlaid maintains bidirectional mappings between document IDs and internal PLAID IDs in pickle files (documents_ids_to_plaid_ids.pkl, plaid_ids_to_documents_ids.pkl) stored under <index_folder>/<index_name>/. These mappings enable incremental document addition via FastPlaid.add_documents without rebuilding the entire index.

Avoid manually deleting these pickle files unless performing a full rebuild. The incremental update capability is crucial for production systems where documents arrive continuously.

End-to-End Optimization Example

import torch
from pylate import indexes, models, retrieve

# Load model on GPU if available

device = "cuda" if torch.cuda.is_available() else "cpu"
model = models.ColBERT(
    model_name_or_path="sentence-transformers/all-MiniLM-L6-v2",
    device=device,
)

# Encode 1M documents in optimized batches

doc_ids = [f"doc_{i}" for i in range(1_000_000)]
texts = ["document content here..."] * 1_000_000  # Your actual corpus

doc_embeddings = model.encode(
    sentences=texts,
    batch_size=256,
    is_query=False,
)

# Build FastPlaid index with performance tuning

index = indexes.PLAID(
    index_folder="optimized_index",
    index_name="fast_plaid_4bit",
    use_fast=True,
    nbits=4,
    n_ivf_probe=16,
    n_full_scores=4096,
    override=True,
)

index.add_documents(
    documents_ids=doc_ids,
    documents_embeddings=doc_embeddings,
)

# Retrieve with hardware-optimized settings

retriever = retrieve.ColBERT(index=index)
queries = ["quantum computing applications", "machine learning optimization"]
query_emb = model.encode(queries, batch_size=1, is_query=True)

results = retriever.retrieve(
    queries_embeddings=query_emb,
    k=10,
    k_token=200,
    batch_size=64,
    device=device,
)

# Display results

for q, hits in zip(queries, results):
    print(f"\nQuery: {q}")
    for r in hits:
        print(f"  • {r.id} (score={r.score:.4f})")

This example demonstrates the complete workflow: GPU-accelerated encoding, FastPlaid index construction with 4-bit quantization, and batched retrieval with tuned candidate pools.

Summary

  • Select FastPlaid (use_fast=True) for production workloads handling millions of embeddings, as implemented in pylate/indexes/plaid.py and pylate/indexes/fast_plaid.py.
  • Tune nbits (product quantization) and n_ivf_probe (inverted file probes) to balance index size, latency, and recall according to your validation metrics.
  • Align k and k_token in pylate/retrieve/colbert.py to ensure sufficient candidates enter the re-ranking stage in pylate/rank/rank.py.
  • Optimize hardware utilization by adjusting batch_size and device parameters based on available GPU memory, leveraging the iter_batch utility in pylate/utils/iter_batch.py.
  • Use subset filtering (FastPlaid only) to restrict searches to specific document collections without scanning the entire index.
  • Preserve ID mapping pickles to enable incremental updates via add_documents without rebuilding the entire index structure.

Frequently Asked Questions

What is the difference between FastPlaid and the standard Stanford PLAID backend in PyLate?

FastPlaid is a Rust-based implementation that uses SIMD-accelerated IVF-PQ search and supports advanced features like subset filtering and parallel batch processing. The Stanford PLAID backend is a pure Python implementation that is slower and lacks subset filtering capabilities. For production systems handling large document collections, FastPlaid (use_fast=True) is the recommended choice as implemented in pylate/indexes/plaid.py.

How does the nbits parameter affect retrieval performance and accuracy?

The nbits parameter controls the bit depth of product quantization in FastPlaid, directly impacting the trade-off between index size and recall. Lower values (e.g., nbits=2) create smaller indexes but require more computational overhead during search and may reduce accuracy. Higher values (e.g., nbits=8) improve recall and speed but increase disk usage. For most applications, nbits=4 provides an optimal balance between storage efficiency and retrieval accuracy.

Can I add new documents to an existing PyLate index without rebuilding it?

Yes, FastPlaid supports incremental updates through the add_documents method in pylate/indexes/fast_plaid.py. The index maintains bidirectional ID mappings in pickle files (documents_ids_to_plaid_ids.pkl and plaid_ids_to_documents_ids.pkl) that enable the system to integrate new embeddings without reconstructing the entire IVF-PQ structure. Avoid deleting these mapping files unless you intend to perform a full index rebuild from scratch.

What is the purpose of the k_token parameter in PyLate retrieval?

The k_token parameter determines the number of token-level candidates retrieved from the index before the exact re-ranking stage in pylate/rank/rank.py. While k specifies the final number of results returned, k_token controls the size of the candidate pool used for computing exact dot-product scores. Setting k_token larger than k (e.g., k=10, k_token=200) improves recall by ensuring high-quality candidates survive the initial approximate search phase, while the re-ranker prunes to the final k results.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →