# Understanding the Skiplist Feature in ColBERT Encoding with PyLate

> Unlock faster retrieval with PyLate's skiplist feature. It automatically masks punctuation and custom tokens in ColBERT encoding, reducing noise for better embedding accuracy.

- Repository: [LightOn/pylate](https://github.com/lightonai/pylate)
- Tags: deep-dive
- Published: 2026-03-06

---

**The skiplist feature in PyLate's ColBERT implementation automatically masks punctuation and custom tokens from document embeddings, reducing noise and improving retrieval speed by excluding irrelevant tokens from similarity calculations.**

The **skiplist** is a specialized masking mechanism in the [PyLate](https://github.com/lightonai/pylate) library that filters designated tokens during **ColBERT encoding**. By default, it removes punctuation marks from document representations, ensuring that only semantically meaningful tokens contribute to similarity scores in late interaction retrieval systems.

## What Is the Skiplist in PyLate ColBERT?

In **late interaction retrieval**, the skiplist serves as a pre-processing filter that identifies and masks specific token IDs before they reach the scoring function. Unlike traditional embedding models that produce a single vector per document, ColBERT generates a token-level embedding matrix. The skiplist ensures that certain tokens—typically punctuation or user-defined stop words—are treated as invisible during the **MaxSim** operation, preventing them from inflating similarity scores artificially.

## How the Skiplist Is Configured

### Default Punctuation Masking

If no custom configuration is provided, PyLate automatically masks all standard punctuation symbols. In [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py), the constructor initializes `skiplist_words` using Python's `string.punctuation` constant:

```python

# pylate/models/colbert.py

self.skiplist_words = (
    skiplist_words
    if skiplist_words is not None
    else self.skiplist_words
    if self.skiplist_words is not None
    else list(string.punctuation)
)

# Convert skiplist words to token IDs

self.skiplist = [
    self.tokenizer.convert_tokens_to_ids(word) for word in self.skiplist_words
]

```

These token IDs are stored in `self.skiplist` and referenced throughout the encoding pipeline for efficient lookup.

### Customizing Skiplist Words

Users can override the default behavior by passing a custom list to the `skiplist_words` parameter. This enables domain-specific filtering, such as removing frequent stop words or special formatting tokens:

```python
from pylate import models

# Initialize with custom skiplist including stop words

model = models.ColBERT(
    model_name_or_path="sentence-transformers/all-MiniLM-L6-v2",
    device="cpu",
    skiplist_words=[",", ".", "the", "and", "is"],
)

```

## Technical Implementation in the ColBERT Model

The core masking logic resides in the static method `skiplist_mask` within [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py). This method generates a boolean tensor where `False` values indicate positions to be excluded:

```python

# pylate/models/colbert.py

@staticmethod
def skiplist_mask(input_ids: torch.Tensor, skiplist: list[int]) -> torch.Tensor:
    """Create a mask for the set of input_ids that are in the skiplist."""
    skiplist = torch.tensor(data=skiplist, dtype=torch.long, device=input_ids.device)
    mask = torch.ones_like(input=input_ids, dtype=torch.bool)  # start with all True

    for token_id in skiplist:
        mask = torch.where(
            condition=input_ids == token_id,
            input=torch.tensor(data=0, dtype=torch.bool, device=input_ids.device),
            other=mask,
        )
    return mask

```

This mask is subsequently combined with the standard `attention_mask` using a logical AND operation, ensuring that both padding tokens and skiplist tokens are excluded from gradient computation and similarity scoring.

## Applying Skiplist Masks During Training

During contrastive training, the `extract_skiplist_mask` function in [`pylate/losses/contrastive.py`](https://github.com/lightonai/pylate/blob/main/pylate/losses/contrastive.py) orchestrates the masking strategy. It preserves all tokens for the query (index 0) while applying the skiplist to all document features:

```python

# pylate/losses/contrastive.py

skiplist_masks = [
    torch.ones_like(sentence_features[0]["input_ids"], dtype=torch.bool)
]
skiplist_masks.extend(
    [
        ColBERT.skiplist_mask(
            input_ids=sentence_feature["input_ids"], skiplist=skiplist
        )
        for sentence_feature in sentence_features[1:]
    ]
)
return [
    torch.logical_and(skiplist_mask, attention_mask)
    for skiplist_mask, attention_mask in zip(skiplist_masks, attention_masks)
]

```

The resulting masks are passed to `colbert_scores`, where they prevent skiplist tokens from participating in the **late interaction** similarity calculation. The same masking approach appears in [`pylate/losses/distillation.py`](https://github.com/lightonai/pylate/blob/main/pylate/losses/distillation.py) and [`pylate/losses/cached_contrastive.py`](https://github.com/lightonai/pylate/blob/main/pylate/losses/cached_contrastive.py), ensuring consistent behavior across different training objectives.

## Performance Benefits of Using Skiplist

**Noise reduction** – Punctuation and formatting characters rarely carry semantic information for passage retrieval. Masking these tokens prevents the model from learning spurious patterns based on sentence structure rather than content meaning.

**Computational efficiency** – Fewer active tokens directly translate to fewer dot-product operations during the **MaxSim** scoring phase. This acceleration applies to both training and inference without requiring architectural modifications.

**Domain flexibility** – The configurable `skiplist_words` parameter enables practitioners to adapt masking strategies to specific corpora, such as removing citation markers in academic text or boilerplate in legal documents.

## Practical Code Examples

### Initializing ColBERT with a Custom Skiplist

Configure the model to ignore specific punctuation marks or domain-specific tokens during encoding:

```python
from pylate import models

# Initialize with custom skiplist (only commas and periods)

model = models.ColBERT(
    model_name_or_path="sentence-transformers/all-MiniLM-L6-v2",
    device="cpu",
    skiplist_words=[",", "."],
)

```

### Training with Skiplist Masking

Demonstrate how the contrastive loss automatically applies skiplist masks to documents while preserving query tokens:

```python
from pylate import models, losses

# Use default punctuation skiplist

model = models.ColBERT("sentence-transformers/all-MiniLM-L6-v2")
contrastive = losses.Contrastive(model=model)

# Tokenize query and documents

query = model.tokenize(["What is the capital of France?"], is_query=True)
doc1 = model.tokenize(["Paris, the capital of France."], is_query=False)
doc2 = model.tokenize(["Berlin is the capital of Germany."], is_query=False)

# Compute loss - commas and periods are automatically masked

loss = contrastive(sentence_features=[query, doc1, doc2])
print("Loss value:", loss.item())

```

### Multi-Process Encoding with Default Skiplist

Scale document encoding across multiple processes while maintaining punctuation masking:

```python
from pylate import models

model = models.ColBERT(
    "sentence-transformers/all-MiniLM-L6-v2",
    device="cpu",
    # Uses default list(string.punctuation) skiplist

)

pool = model.start_multi_process_pool()
embeddings = model.encode_multi_process(
    sentences=[
        "Hello, world!",  # punctuation masked

        "How does skiplist work?"
    ],
    pool=pool,
    is_query=False,  # documents → skiplist applied

)
model.stop_multi_process_pool(pool)

print(embeddings.shape)  # (2, num_tokens, embedding_dim)

```

## Summary

- The **skiplist** in PyLate's ColBERT implementation filters unwanted tokens from document embeddings by masking them in the attention mechanism.
- By default, `list(string.punctuation)` is masked, but users can customize `skiplist_words` for domain-specific requirements.
- The masking logic is implemented in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py) via the `skiplist_mask` static method, which generates boolean masks for token IDs.
- During training, [`pylate/losses/contrastive.py`](https://github.com/lightonai/pylate/blob/main/pylate/losses/contrastive.py) applies these masks to documents (not queries) through `extract_skiplist_mask`, ensuring punctuation doesn't affect similarity scores.
- This mechanism reduces semantic noise, decreases computational overhead during MaxSim scoring, and provides flexible domain adaptation.

## Frequently Asked Questions

### What tokens are included in the default PyLate skiplist?

By default, PyLate initializes the skiplist with all standard punctuation symbols using Python's `string.punctuation` constant. This includes characters such as periods, commas, exclamation marks, quotation marks, and other ASCII punctuation marks. You can inspect this default behavior in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py) where the constructor falls back to `list(string.punctuation)` when no custom `skiplist_words` are provided.

### Can I add stop words to the skiplist for domain-specific retrieval?

Yes, the `skiplist_words` parameter accepts any list of strings, allowing you to mask domain-specific stop words or special formatting tokens. For example, when processing legal documents, you might add ["the", "and", "§", "¶"] to remove both common stop words and section symbols. Simply pass your custom list when initializing the `ColBERT` model, and the tokenizer will convert these words to the appropriate token IDs for masking.

### Does the skiplist affect query embeddings or only document embeddings?

During standard contrastive training, the skiplist applies exclusively to document embeddings. In [`pylate/losses/contrastive.py`](https://github.com/lightonai/pylate/blob/main/pylate/losses/contrastive.py), the `extract_skiplist_mask` function creates a mask of all ones for the query (the first sentence feature in the batch) while building actual skiplist masks for all subsequent document features. This preserves query tokens that might include punctuation for semantic clarification while ensuring that document representations remain clean of formatting artifacts.

### How does skiplist masking impact retrieval accuracy and speed?

Skiplist masking typically improves both metrics simultaneously. By removing punctuation tokens from the MaxSim calculation, the model avoids false similarity matches based on formatting rather than content, leading to more precise semantic retrieval. Computationally, fewer active tokens directly reduce the number of dot-product operations required during the late interaction phase, accelerating both training convergence and inference latency without modifying the underlying ColBERT architecture.