Comparing fast_plaid and stanford_plaid Implementations in PyLate: Architecture, Performance, and Use Cases

PyLate offers two distinct PLAID index backends—FastPlaid for GPU-accelerated C++/CUDA retrieval and StanfordPlaid for research-friendly ColBERT-style indexing—each with unique embedding conversion logic, update mechanisms, and configuration parameters.

PyLate is an open-source library that implements multi-vector retrieval for late interaction models like ColBERT. When building a PLAID index for scalable semantic search, developers must choose between two backends: fast_plaid and stanford_plaid. This comparison examines their architectural differences, API implementations, and performance characteristics based on the actual source code in the lightonai/pylate repository.

Backend Architecture and Design Philosophy

FastPlaid: High-Performance C++/CUDA Engine

FastPlaid wraps the fast-plaid library—a high-performance C++/CUDA implementation that quantizes vectors using product quantization (PQ-IVF) and optionally leverages Triton kernels for accelerated k-means clustering. This backend is optimized for production deployments requiring low-latency batch searches across massive corpora.

In pylate/indexes/fast_plaid.py, the FastPlaid class initializes the native index with explicit GPU-oriented parameters like n_ivf_probe and n_full_scores (lines 98-107 and 122-129).

StanfordPlaid: Research-Friendly ColBERT Implementation

StanfordPlaid (exposed as indexes.PLAID) wraps Stanford's original ColBERT-style PLAID implementation (stanford-nlp). It builds on the original PLAID C++ code but exposes functionality through Python wrappers including Indexer and Searcher objects. This backend prioritizes flexibility for research prototypes and compatibility with the original ColBERT ecosystem.

The implementation in pylate/indexes/stanford_plaid.py uses the Indexer class for initial builds and IndexUpdater for incremental modifications (lines 99-112).

Data Handling and Embedding Conversion

The two backends expect different embedding formats and use distinct conversion utilities.

FastPlaid: Torch-Centric Conversion

FastPlaid accepts NumPy arrays, PyTorch tensors, or plain Python lists, converting them to a list of torch.Tensor objects via convert_embeddings_to_torch. This matches the fast-plaid API expectations for GPU batch processing.

from pylate.indexes.fast_plaid import convert_embeddings_to_torch
import numpy as np

# FastPlaid handles 2D or 3D embeddings internally

embeddings = np.random.randn(10, 128).astype(np.float32)
torch_embeddings = convert_embeddings_to_torch(embeddings)

# Returns list of torch.Tensor objects

See pylate/indexes/fast_plaid.py lines 19-44 for the implementation details.

StanfordPlaid: NumPy Reshaping

StanfordPlaid expects embeddings shaped (batch, tokens, dim)—the standard ColBERT multi-vector format. The helper reshape_embeddings expands 2-D arrays to 3-D and converts PyTorch tensors to NumPy for the underlying C++ engine.

from pylate.indexes.stanford_plaid import reshape_embeddings
import torch

# StanfordPlaid requires 3D format (batch, tokens, dim)

embeddings_2d = torch.randn(10, 128)
embeddings_3d = reshape_embeddings(embeddings_2d)

# Returns numpy array with shape (10, 1, 128) if input was 2D

See pylate/indexes/stanford_plaid.py lines 18-33 for the implementation.

Index Management and CRUD Operations

Creating and Updating Indexes

FastPlaid creates indexes via self.fast_plaid.create() with explicit parameters controlling k-means iterations, centroid limits, and quantization bits. The index persists under fast_plaid_index.


# FastPlaid index creation parameters

fast_index = indexes.FastPlaid(
    index_folder="indexes",
    index_name="fast_index",
    nbits=4,                    # Quantization bits

    kmeans_niters=10,           # K-means iterations

    max_points_per_centroid=128,
    n_ivf_probe=8,              # IVF probe count for search

    use_triton=True,            # Enable Triton kernels

)

See pylate/indexes/fast_plaid.py lines 98-107 and 122-129.

StanfordPlaid uses the Indexer object for initial builds and IndexUpdater for incremental additions. It supports updating existing indexes without full rebuilds.


# StanfordPlaid uses Indexer/IndexUpdater pattern

stanford_index = indexes.PLAID(
    index_folder="indexes",
    index_name="stanford_index",
    embedding_size=128,
    nbits=2,
)

# First call creates index, subsequent calls use IndexUpdater

stanford_index.add_documents(
    documents_ids=["doc1", "doc2"],
    documents_embeddings=doc_embeddings,
)

See pylate/indexes/stanford_plaid.py lines 99-112.

Document Removal Strategies

FastPlaid does not support true vector removal. The remove_documents method deletes ID mappings, calls self.fast_plaid.delete(plaid_ids_to_remove), and re-indexes the remaining IDs locally.


# FastPlaid removal re-indexes remaining documents

fast_index.remove_documents(["doc1"])  # Marks for deletion and re-indexes

See pylate/indexes/fast_plaid.py lines 30-44 and 56-72.

StanfordPlaid uses IndexUpdater.remove(plaid_ids) for true removal and persists changes immediately.


# StanfordPlaid supports direct removal via IndexUpdater

stanford_index.remove_documents(["doc1"])  # Calls IndexUpdater.remove()

See pylate/indexes/stanford_plaid.py lines 26-44.

Query Execution and Retrieval

Both implementations expose a __call__ method for search, but handle query processing differently.

FastPlaid converts query embeddings, optionally translates document-ID subsets to PLAID internal IDs, then runs self.fast_plaid.search(). Results map back to user IDs via pickle mappings.


# FastPlaid search with optional document subset filtering

results = fast_index(
    query_embeddings, 
    k=10,
    documents_ids=["doc1", "doc2"]  # Optional subset filtering

)

See pylate/indexes/fast_plaid.py lines 76-84 and 115-130.

StanfordPlaid reshapes query embeddings and calls self.searcher.search() for each query, building RerankResult objects from the ID mapping.


# StanfordPlaid returns RerankResult objects

results = stanford_index(query_embeddings, k=10)
for hit in results[0]:
    print(f"{hit.id}: {hit.score}")

See pylate/indexes/stanford_plaid.py lines 52-66 and 70-82.

Configuration Parameters and Performance Tuning

Both backends expose different tuning knobs reflecting their underlying architectures.

Parameter FastPlaid StanfordPlaid
Quantization nbits (PQ bits) nbits (PQ bits)
Clustering kmeans_niters, max_points_per_centroid, n_samples_kmeans kmeans_niters, ndocs
Search n_ivf_probe, n_full_scores, batch_size ncells, centroid_score_threshold, search_batch_size
Hardware use_triton (Triton kernels) use_triton, nranks (GPU ranks)
Dimensions Inferred from data embedding_size (required)

FastPlaid optimizes for GPU-accelerated batch searches with large IVF-PQ tables. Configure n_ivf_probe (number of IVF clusters to probe) and n_full_scores (number of candidates for full scoring) to balance recall versus latency.

StanfordPlaid provides research-oriented parameters like centroid_score_threshold for filtering and nranks for multi-GPU indexing. It generally exhibits higher overhead due to Python wrapper layers but offers greater flexibility for experimental configurations.

Summary

  • FastPlaid leverages the fast-plaid C++/CUDA library with Triton kernel support, converting embeddings to torch.Tensor lists via convert_embeddings_to_torch and optimizing for high-throughput GPU batch searches with parameters like n_ivf_probe and n_full_scores.
  • StanfordPlaid wraps Stanford's ColBERT-style PLAID implementation using Indexer and IndexUpdater objects, reshaping embeddings to (batch, tokens, dim) via reshape_embeddings, and supporting true incremental updates and removal via IndexUpdater.remove.
  • Key architectural differences include FastPlaid's lack of true vector removal (requiring local re-indexing) versus StanfordPlaid's native removal support, and FastPlaid's GPU-optimized IVF-PQ quantization versus StanfordPlaid's research-flexible Python wrappers.
  • Both implementations share the same high-level API (add_documents, remove_documents, __call__) and maintain persistent ID-mapping pickles in pylate/indexes/fast_plaid.py and pylate/indexes/stanford_plaid.py to ensure stable user-defined IDs across index rebuilds.

Frequently Asked Questions

What is the main performance difference between FastPlaid and StanfordPlaid?

FastPlaid is optimized for production latency and throughput, utilizing C++/CUDA kernels with optional Triton acceleration and efficient IVF-PQ quantization. StanfordPlaid, while functionally equivalent, incurs Python wrapper overhead from the original Stanford ColBERT implementation, making it more suitable for research prototyping than high-throughput serving.

Can I switch between FastPlaid and StanfordPlaid without changing my embedding model?

Yes, both backends accept the same ColBERT-style multi-vector embeddings and share identical high-level APIs in PyLate. However, you must ensure embeddings conform to each backend's expected format: FastPlaid uses convert_embeddings_to_torch for list-of-tensors conversion, while StanfordPlaid uses reshape_embeddings to ensure (batch, tokens, dim) NumPy arrays.

Does FastPlaid support incremental document removal?

No, FastPlaid does not support true vector removal. The remove_documents method in pylate/indexes/fast_plaid.py (lines 30-44 and 56-72) deletes ID mappings and calls self.fast_plaid.delete(), but effectively requires re-indexing remaining documents locally. For use cases requiring frequent deletions, StanfordPlaid is preferable as it uses IndexUpdater.remove() for true incremental removal.

Which backend should I choose for large-scale production deployment?

Choose FastPlaid for large-scale production deployments where query latency and throughput are critical. Configure n_ivf_probe and n_full_scores to tune the recall-speed trade-off, and enable use_triton for optimized GPU kernels. Choose StanfordPlaid if you require the specific research features of the original Stanford implementation, need true incremental updates with removal support, or are prototyping with the ColBERT ecosystem.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →