Comparing fast_plaid and stanford_plaid Implementations in PyLate: Architecture, Performance, and Use Cases
PyLate offers two distinct PLAID index backends—FastPlaid for GPU-accelerated C++/CUDA retrieval and StanfordPlaid for research-friendly ColBERT-style indexing—each with unique embedding conversion logic, update mechanisms, and configuration parameters.
PyLate is an open-source library that implements multi-vector retrieval for late interaction models like ColBERT. When building a PLAID index for scalable semantic search, developers must choose between two backends: fast_plaid and stanford_plaid. This comparison examines their architectural differences, API implementations, and performance characteristics based on the actual source code in the lightonai/pylate repository.
Backend Architecture and Design Philosophy
FastPlaid: High-Performance C++/CUDA Engine
FastPlaid wraps the fast-plaid library—a high-performance C++/CUDA implementation that quantizes vectors using product quantization (PQ-IVF) and optionally leverages Triton kernels for accelerated k-means clustering. This backend is optimized for production deployments requiring low-latency batch searches across massive corpora.
In pylate/indexes/fast_plaid.py, the FastPlaid class initializes the native index with explicit GPU-oriented parameters like n_ivf_probe and n_full_scores (lines 98-107 and 122-129).
StanfordPlaid: Research-Friendly ColBERT Implementation
StanfordPlaid (exposed as indexes.PLAID) wraps Stanford's original ColBERT-style PLAID implementation (stanford-nlp). It builds on the original PLAID C++ code but exposes functionality through Python wrappers including Indexer and Searcher objects. This backend prioritizes flexibility for research prototypes and compatibility with the original ColBERT ecosystem.
The implementation in pylate/indexes/stanford_plaid.py uses the Indexer class for initial builds and IndexUpdater for incremental modifications (lines 99-112).
Data Handling and Embedding Conversion
The two backends expect different embedding formats and use distinct conversion utilities.
FastPlaid: Torch-Centric Conversion
FastPlaid accepts NumPy arrays, PyTorch tensors, or plain Python lists, converting them to a list of torch.Tensor objects via convert_embeddings_to_torch. This matches the fast-plaid API expectations for GPU batch processing.
from pylate.indexes.fast_plaid import convert_embeddings_to_torch
import numpy as np
# FastPlaid handles 2D or 3D embeddings internally
embeddings = np.random.randn(10, 128).astype(np.float32)
torch_embeddings = convert_embeddings_to_torch(embeddings)
# Returns list of torch.Tensor objects
See pylate/indexes/fast_plaid.py lines 19-44 for the implementation details.
StanfordPlaid: NumPy Reshaping
StanfordPlaid expects embeddings shaped (batch, tokens, dim)—the standard ColBERT multi-vector format. The helper reshape_embeddings expands 2-D arrays to 3-D and converts PyTorch tensors to NumPy for the underlying C++ engine.
from pylate.indexes.stanford_plaid import reshape_embeddings
import torch
# StanfordPlaid requires 3D format (batch, tokens, dim)
embeddings_2d = torch.randn(10, 128)
embeddings_3d = reshape_embeddings(embeddings_2d)
# Returns numpy array with shape (10, 1, 128) if input was 2D
See pylate/indexes/stanford_plaid.py lines 18-33 for the implementation.
Index Management and CRUD Operations
Creating and Updating Indexes
FastPlaid creates indexes via self.fast_plaid.create() with explicit parameters controlling k-means iterations, centroid limits, and quantization bits. The index persists under fast_plaid_index.
# FastPlaid index creation parameters
fast_index = indexes.FastPlaid(
index_folder="indexes",
index_name="fast_index",
nbits=4, # Quantization bits
kmeans_niters=10, # K-means iterations
max_points_per_centroid=128,
n_ivf_probe=8, # IVF probe count for search
use_triton=True, # Enable Triton kernels
)
See pylate/indexes/fast_plaid.py lines 98-107 and 122-129.
StanfordPlaid uses the Indexer object for initial builds and IndexUpdater for incremental additions. It supports updating existing indexes without full rebuilds.
# StanfordPlaid uses Indexer/IndexUpdater pattern
stanford_index = indexes.PLAID(
index_folder="indexes",
index_name="stanford_index",
embedding_size=128,
nbits=2,
)
# First call creates index, subsequent calls use IndexUpdater
stanford_index.add_documents(
documents_ids=["doc1", "doc2"],
documents_embeddings=doc_embeddings,
)
See pylate/indexes/stanford_plaid.py lines 99-112.
Document Removal Strategies
FastPlaid does not support true vector removal. The remove_documents method deletes ID mappings, calls self.fast_plaid.delete(plaid_ids_to_remove), and re-indexes the remaining IDs locally.
# FastPlaid removal re-indexes remaining documents
fast_index.remove_documents(["doc1"]) # Marks for deletion and re-indexes
See pylate/indexes/fast_plaid.py lines 30-44 and 56-72.
StanfordPlaid uses IndexUpdater.remove(plaid_ids) for true removal and persists changes immediately.
# StanfordPlaid supports direct removal via IndexUpdater
stanford_index.remove_documents(["doc1"]) # Calls IndexUpdater.remove()
See pylate/indexes/stanford_plaid.py lines 26-44.
Query Execution and Retrieval
Both implementations expose a __call__ method for search, but handle query processing differently.
FastPlaid converts query embeddings, optionally translates document-ID subsets to PLAID internal IDs, then runs self.fast_plaid.search(). Results map back to user IDs via pickle mappings.
# FastPlaid search with optional document subset filtering
results = fast_index(
query_embeddings,
k=10,
documents_ids=["doc1", "doc2"] # Optional subset filtering
)
See pylate/indexes/fast_plaid.py lines 76-84 and 115-130.
StanfordPlaid reshapes query embeddings and calls self.searcher.search() for each query, building RerankResult objects from the ID mapping.
# StanfordPlaid returns RerankResult objects
results = stanford_index(query_embeddings, k=10)
for hit in results[0]:
print(f"{hit.id}: {hit.score}")
See pylate/indexes/stanford_plaid.py lines 52-66 and 70-82.
Configuration Parameters and Performance Tuning
Both backends expose different tuning knobs reflecting their underlying architectures.
| Parameter | FastPlaid | StanfordPlaid |
|---|---|---|
| Quantization | nbits (PQ bits) |
nbits (PQ bits) |
| Clustering | kmeans_niters, max_points_per_centroid, n_samples_kmeans |
kmeans_niters, ndocs |
| Search | n_ivf_probe, n_full_scores, batch_size |
ncells, centroid_score_threshold, search_batch_size |
| Hardware | use_triton (Triton kernels) |
use_triton, nranks (GPU ranks) |
| Dimensions | Inferred from data | embedding_size (required) |
FastPlaid optimizes for GPU-accelerated batch searches with large IVF-PQ tables. Configure n_ivf_probe (number of IVF clusters to probe) and n_full_scores (number of candidates for full scoring) to balance recall versus latency.
StanfordPlaid provides research-oriented parameters like centroid_score_threshold for filtering and nranks for multi-GPU indexing. It generally exhibits higher overhead due to Python wrapper layers but offers greater flexibility for experimental configurations.
Summary
- FastPlaid leverages the
fast-plaidC++/CUDA library with Triton kernel support, converting embeddings totorch.Tensorlists viaconvert_embeddings_to_torchand optimizing for high-throughput GPU batch searches with parameters liken_ivf_probeandn_full_scores. - StanfordPlaid wraps Stanford's ColBERT-style PLAID implementation using
IndexerandIndexUpdaterobjects, reshaping embeddings to(batch, tokens, dim)viareshape_embeddings, and supporting true incremental updates and removal viaIndexUpdater.remove. - Key architectural differences include FastPlaid's lack of true vector removal (requiring local re-indexing) versus StanfordPlaid's native removal support, and FastPlaid's GPU-optimized IVF-PQ quantization versus StanfordPlaid's research-flexible Python wrappers.
- Both implementations share the same high-level API (
add_documents,remove_documents,__call__) and maintain persistent ID-mapping pickles inpylate/indexes/fast_plaid.pyandpylate/indexes/stanford_plaid.pyto ensure stable user-defined IDs across index rebuilds.
Frequently Asked Questions
What is the main performance difference between FastPlaid and StanfordPlaid?
FastPlaid is optimized for production latency and throughput, utilizing C++/CUDA kernels with optional Triton acceleration and efficient IVF-PQ quantization. StanfordPlaid, while functionally equivalent, incurs Python wrapper overhead from the original Stanford ColBERT implementation, making it more suitable for research prototyping than high-throughput serving.
Can I switch between FastPlaid and StanfordPlaid without changing my embedding model?
Yes, both backends accept the same ColBERT-style multi-vector embeddings and share identical high-level APIs in PyLate. However, you must ensure embeddings conform to each backend's expected format: FastPlaid uses convert_embeddings_to_torch for list-of-tensors conversion, while StanfordPlaid uses reshape_embeddings to ensure (batch, tokens, dim) NumPy arrays.
Does FastPlaid support incremental document removal?
No, FastPlaid does not support true vector removal. The remove_documents method in pylate/indexes/fast_plaid.py (lines 30-44 and 56-72) deletes ID mappings and calls self.fast_plaid.delete(), but effectively requires re-indexing remaining documents locally. For use cases requiring frequent deletions, StanfordPlaid is preferable as it uses IndexUpdater.remove() for true incremental removal.
Which backend should I choose for large-scale production deployment?
Choose FastPlaid for large-scale production deployments where query latency and throughput are critical. Configure n_ivf_probe and n_full_scores to tune the recall-speed trade-off, and enable use_triton for optimized GPU kernels. Choose StanfordPlaid if you require the specific research features of the original Stanford implementation, need true incremental updates with removal support, or are prototyping with the ColBERT ecosystem.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →