Understanding the Skiplist Feature in ColBERT Encoding with PyLate
The skiplist feature in PyLate's ColBERT implementation automatically masks punctuation and custom tokens from document embeddings, reducing noise and improving retrieval speed by excluding irrelevant tokens from similarity calculations.
The skiplist is a specialized masking mechanism in the PyLate library that filters designated tokens during ColBERT encoding. By default, it removes punctuation marks from document representations, ensuring that only semantically meaningful tokens contribute to similarity scores in late interaction retrieval systems.
What Is the Skiplist in PyLate ColBERT?
In late interaction retrieval, the skiplist serves as a pre-processing filter that identifies and masks specific token IDs before they reach the scoring function. Unlike traditional embedding models that produce a single vector per document, ColBERT generates a token-level embedding matrix. The skiplist ensures that certain tokens—typically punctuation or user-defined stop words—are treated as invisible during the MaxSim operation, preventing them from inflating similarity scores artificially.
How the Skiplist Is Configured
Default Punctuation Masking
If no custom configuration is provided, PyLate automatically masks all standard punctuation symbols. In pylate/models/colbert.py, the constructor initializes skiplist_words using Python's string.punctuation constant:
# pylate/models/colbert.py
self.skiplist_words = (
skiplist_words
if skiplist_words is not None
else self.skiplist_words
if self.skiplist_words is not None
else list(string.punctuation)
)
# Convert skiplist words to token IDs
self.skiplist = [
self.tokenizer.convert_tokens_to_ids(word) for word in self.skiplist_words
]
These token IDs are stored in self.skiplist and referenced throughout the encoding pipeline for efficient lookup.
Customizing Skiplist Words
Users can override the default behavior by passing a custom list to the skiplist_words parameter. This enables domain-specific filtering, such as removing frequent stop words or special formatting tokens:
from pylate import models
# Initialize with custom skiplist including stop words
model = models.ColBERT(
model_name_or_path="sentence-transformers/all-MiniLM-L6-v2",
device="cpu",
skiplist_words=[",", ".", "the", "and", "is"],
)
Technical Implementation in the ColBERT Model
The core masking logic resides in the static method skiplist_mask within pylate/models/colbert.py. This method generates a boolean tensor where False values indicate positions to be excluded:
# pylate/models/colbert.py
@staticmethod
def skiplist_mask(input_ids: torch.Tensor, skiplist: list[int]) -> torch.Tensor:
"""Create a mask for the set of input_ids that are in the skiplist."""
skiplist = torch.tensor(data=skiplist, dtype=torch.long, device=input_ids.device)
mask = torch.ones_like(input=input_ids, dtype=torch.bool) # start with all True
for token_id in skiplist:
mask = torch.where(
condition=input_ids == token_id,
input=torch.tensor(data=0, dtype=torch.bool, device=input_ids.device),
other=mask,
)
return mask
This mask is subsequently combined with the standard attention_mask using a logical AND operation, ensuring that both padding tokens and skiplist tokens are excluded from gradient computation and similarity scoring.
Applying Skiplist Masks During Training
During contrastive training, the extract_skiplist_mask function in pylate/losses/contrastive.py orchestrates the masking strategy. It preserves all tokens for the query (index 0) while applying the skiplist to all document features:
# pylate/losses/contrastive.py
skiplist_masks = [
torch.ones_like(sentence_features[0]["input_ids"], dtype=torch.bool)
]
skiplist_masks.extend(
[
ColBERT.skiplist_mask(
input_ids=sentence_feature["input_ids"], skiplist=skiplist
)
for sentence_feature in sentence_features[1:]
]
)
return [
torch.logical_and(skiplist_mask, attention_mask)
for skiplist_mask, attention_mask in zip(skiplist_masks, attention_masks)
]
The resulting masks are passed to colbert_scores, where they prevent skiplist tokens from participating in the late interaction similarity calculation. The same masking approach appears in pylate/losses/distillation.py and pylate/losses/cached_contrastive.py, ensuring consistent behavior across different training objectives.
Performance Benefits of Using Skiplist
Noise reduction – Punctuation and formatting characters rarely carry semantic information for passage retrieval. Masking these tokens prevents the model from learning spurious patterns based on sentence structure rather than content meaning.
Computational efficiency – Fewer active tokens directly translate to fewer dot-product operations during the MaxSim scoring phase. This acceleration applies to both training and inference without requiring architectural modifications.
Domain flexibility – The configurable skiplist_words parameter enables practitioners to adapt masking strategies to specific corpora, such as removing citation markers in academic text or boilerplate in legal documents.
Practical Code Examples
Initializing ColBERT with a Custom Skiplist
Configure the model to ignore specific punctuation marks or domain-specific tokens during encoding:
from pylate import models
# Initialize with custom skiplist (only commas and periods)
model = models.ColBERT(
model_name_or_path="sentence-transformers/all-MiniLM-L6-v2",
device="cpu",
skiplist_words=[",", "."],
)
Training with Skiplist Masking
Demonstrate how the contrastive loss automatically applies skiplist masks to documents while preserving query tokens:
from pylate import models, losses
# Use default punctuation skiplist
model = models.ColBERT("sentence-transformers/all-MiniLM-L6-v2")
contrastive = losses.Contrastive(model=model)
# Tokenize query and documents
query = model.tokenize(["What is the capital of France?"], is_query=True)
doc1 = model.tokenize(["Paris, the capital of France."], is_query=False)
doc2 = model.tokenize(["Berlin is the capital of Germany."], is_query=False)
# Compute loss - commas and periods are automatically masked
loss = contrastive(sentence_features=[query, doc1, doc2])
print("Loss value:", loss.item())
Multi-Process Encoding with Default Skiplist
Scale document encoding across multiple processes while maintaining punctuation masking:
from pylate import models
model = models.ColBERT(
"sentence-transformers/all-MiniLM-L6-v2",
device="cpu",
# Uses default list(string.punctuation) skiplist
)
pool = model.start_multi_process_pool()
embeddings = model.encode_multi_process(
sentences=[
"Hello, world!", # punctuation masked
"How does skiplist work?"
],
pool=pool,
is_query=False, # documents → skiplist applied
)
model.stop_multi_process_pool(pool)
print(embeddings.shape) # (2, num_tokens, embedding_dim)
Summary
- The skiplist in PyLate's ColBERT implementation filters unwanted tokens from document embeddings by masking them in the attention mechanism.
- By default,
list(string.punctuation)is masked, but users can customizeskiplist_wordsfor domain-specific requirements. - The masking logic is implemented in
pylate/models/colbert.pyvia theskiplist_maskstatic method, which generates boolean masks for token IDs. - During training,
pylate/losses/contrastive.pyapplies these masks to documents (not queries) throughextract_skiplist_mask, ensuring punctuation doesn't affect similarity scores. - This mechanism reduces semantic noise, decreases computational overhead during MaxSim scoring, and provides flexible domain adaptation.
Frequently Asked Questions
What tokens are included in the default PyLate skiplist?
By default, PyLate initializes the skiplist with all standard punctuation symbols using Python's string.punctuation constant. This includes characters such as periods, commas, exclamation marks, quotation marks, and other ASCII punctuation marks. You can inspect this default behavior in pylate/models/colbert.py where the constructor falls back to list(string.punctuation) when no custom skiplist_words are provided.
Can I add stop words to the skiplist for domain-specific retrieval?
Yes, the skiplist_words parameter accepts any list of strings, allowing you to mask domain-specific stop words or special formatting tokens. For example, when processing legal documents, you might add ["the", "and", "§", "¶"] to remove both common stop words and section symbols. Simply pass your custom list when initializing the ColBERT model, and the tokenizer will convert these words to the appropriate token IDs for masking.
Does the skiplist affect query embeddings or only document embeddings?
During standard contrastive training, the skiplist applies exclusively to document embeddings. In pylate/losses/contrastive.py, the extract_skiplist_mask function creates a mask of all ones for the query (the first sentence feature in the batch) while building actual skiplist masks for all subsequent document features. This preserves query tokens that might include punctuation for semantic clarification while ensuring that document representations remain clean of formatting artifacts.
How does skiplist masking impact retrieval accuracy and speed?
Skiplist masking typically improves both metrics simultaneously. By removing punctuation tokens from the MaxSim calculation, the model avoids false similarity matches based on formatting rather than content, leading to more precise semantic retrieval. Computationally, fewer active tokens directly reduce the number of dot-product operations required during the late interaction phase, accelerating both training convergence and inference latency without modifying the underlying ColBERT architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →