# How PyLate Handles Query and Document Prefix Tokens in ColBERT

> Discover how PyLate manages query and document prefix tokens for ColBERT encoding. Learn about default and custom configurations to enhance your search results.

- Repository: [LightOn/pylate](https://github.com/lightonai/pylate)
- Tags: internals
- Published: 2026-03-06

---

**PyLate prepends dedicated prefix tokens to queries and documents during encoding, using configurable strings that default to `"[Q] "` and `"[D] "` when no custom values or checkpoint metadata are provided.**

The `lightonai/pylate` library implements a ColBERT-style retrieval model that relies on distinct prefix tokens to differentiate between query and document sequences. This design allows the transformer to learn separate representations for questions and passages while maintaining a shared vocabulary. The prefix handling strategy involves automatic token addition to the tokenizer, embedding matrix resizing, and runtime insertion during the encoding process.

## How Prefix Tokens Are Configured

PyLate determines the query and document prefix values through a hierarchical fallback system defined in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py).

### Custom User Definitions

When initializing the `ColBERT` class, users can explicitly set prefixes via the `query_prefix` and `document_prefix` arguments:

```python
from pylate import models

model = models.ColBERT(
    model_name_or_path="jinaai/jina-colbert-v2",
    query_prefix="[MyQuery]",
    document_prefix="[MyDoc]"
)

```

### Checkpoint Metadata Extraction

If no custom prefixes are provided, PyLate searches for an `artifact.metadata` file associated with Stanford-NLP-style checkpoints. When present, the model extracts `query_token_id` and `doc_token_id` values from this metadata to use as the prefix identifiers.

### Default Fallback Tokens

When neither user configuration nor checkpoint metadata is available, PyLate defaults to human-readable strings. The constructor logic (approximately lines 66-71 in [`colbert.py`](https://github.com/lightonai/pylate/blob/main/colbert.py)) implements the following fallback chain:

```python
self.query_prefix = (
    query_prefix
    if query_prefix is not None
    else self.query_prefix
    if self.query_prefix is not None
    else "[Q] "
)
self.document_prefix = (
    document_prefix
    if document_prefix is not None
    else self.document_prefix
    if self.document_prefix is not None
    else "[D] "
)

```

## Adding Prefix Tokens to the Tokenizer

Once prefix values are determined, PyLate modifies the tokenizer vocabulary and model embeddings to accommodate these new special tokens.

### Vocabulary Extension and Embedding Resizing

The `ColBERT` constructor performs the following operations (approximately lines 81-86 in [`colbert.py`](https://github.com/lightonai/pylate/blob/main/colbert.py)):

```python

# Initial resize to ensure base vocabulary size

self._first_module().auto_model.resize_token_embeddings(len(self.tokenizer))

# Add the prefix tokens to vocabulary

self.tokenizer.add_tokens([self.query_prefix, self.document_prefix])

# Second resize to account for the two new tokens

self._first_module().auto_model.resize_token_embeddings(len(self.tokenizer))

```

The double resizing ensures compatibility with tokenizers that may not support dynamic vocabulary expansion, preventing indexing errors when the model processes sequences containing the new prefix tokens.

### Token ID Caching

After adding tokens to the vocabulary, PyLate immediately caches the corresponding token IDs for runtime efficiency:

```python
self.document_prefix_id = self.tokenizer.convert_tokens_to_ids(self.document_prefix)
self.query_prefix_id = self.tokenizer.convert_tokens_to_ids(self.query_prefix)

```

These cached IDs enable O(1) lookup during the encoding phase.

## Encoding with Prefix Tokens

During inference, PyLate automatically prepends the appropriate prefix token based on the input type.

### Automatic Prefix Insertion

The `encode` method accepts an `is_query` boolean parameter that determines which prefix to apply. When `is_query=True` (the default), the query prefix token ID is prepended to the `input_ids` tensor; when `is_query=False`, the document prefix is used instead.

This logic delegates to the underlying `SentenceTransformer.tokenize` routine, which handles the actual tensor manipulation:

```python

# Queries (default behavior)

query_embeddings = model.encode(["What is PyLate?"])

# Documents (explicitly set is_query=False)

doc_embeddings = model.encode(["PyLate is a retrieval library."], is_query=False)

```

### Manual Prefix Insertion

For advanced use cases requiring direct tensor manipulation, PyLate provides the `insert_prefix_token` static method in [`colbert.py`](https://github.com/lightonai/pylate/blob/main/colbert.py) (approximately lines 70-81):

```python
@staticmethod
def insert_prefix_token(input_ids: torch.Tensor, prefix_id: int) -> torch.Tensor:
    prefix_tensor = torch.full(
        size=(input_ids.size(0), 1),
        fill_value=prefix_id,
        dtype=input_ids.dtype,
        device=input_ids.device,
    )
    return torch.cat([input_ids[:, :1], prefix_tensor, input_ids[:, 1:]], dim=1)

```

This utility inserts the prefix token immediately after the `[CLS]` token (position 0) and before the rest of the sequence, maintaining the standard BERT-style input format while marking the sequence type.

## Impact on Retrieval Architecture

The prefix token strategy enables several architectural optimizations in PyLate's retrieval pipeline.

### Differential Padding Strategies

Queries and documents receive different padding treatments after prefix insertion:

- **Queries** are padded to a uniform length to facilitate batch processing during similarity computation
- **Documents** remain unpadded by default, allowing variable-length encoding that preserves the original sequence boundaries marked by the prefix token

This distinction optimizes memory usage for document indexing while maintaining computational efficiency for query processing.

### Separation of Representations

By prepending distinct prefix tokens, PyLate enables the transformer to learn query-specific and document-specific attention patterns within a shared parameter space. The `[Q]` and `[D]` tokens (or their custom equivalents) act as type identifiers similar to token type embeddings in traditional BERT models, but with the flexibility to use arbitrary string values.

## Summary

- PyLate uses **dedicated prefix tokens** to distinguish queries from documents in ColBERT retrieval models
- Prefix values follow a **hierarchy**: user-provided arguments → checkpoint metadata → default `"[Q] "` and `"[D] "` strings
- The tokenizer vocabulary is **dynamically extended** to include prefix tokens, with the embedding matrix resized accordingly
- During encoding, the `is_query` parameter determines whether the query or document prefix is **automatically prepended** to input sequences
- The `insert_prefix_token` static method provides **manual control** for advanced tensor manipulation

## Frequently Asked Questions

### What happens if I don't specify custom prefix tokens when loading a PyLate model?

If you omit the `query_prefix` and `document_prefix` arguments, PyLate first checks for an `artifact.metadata` file in the checkpoint directory. If found, it extracts the token IDs stored as `query_token_id` and `doc_token_id`. If no metadata exists, it falls back to the default human-readable strings `"[Q] "` and `"[D] "`, which are then added to the tokenizer vocabulary.

### Can I use different prefix tokens for different types of documents?

The standard `ColBERT` class supports exactly one document prefix and one query prefix per model instance. However, you can implement custom prefix logic by manually calling `insert_prefix_token` with different prefix IDs before passing tensors to the model, or by maintaining separate model instances with different `document_prefix` configurations for distinct document types.

### Why does PyLate resize the token embeddings twice when adding prefix tokens?

PyLate calls `resize_token_embeddings` twice as a defensive programming measure. The first call ensures the embedding matrix matches the base tokenizer size before adding new tokens. After calling `tokenizer.add_tokens()` to insert the prefix strings, the second resize expands the matrix to accommodate the two new token IDs. This double-resize strategy prevents indexing errors when working with tokenizers that may not support dynamic vocabulary expansion.

### How does the `is_query` parameter affect the encoding process?

When you call `model.encode(texts, is_query=True)`, the method forwards the flag to the underlying tokenizer, which prepends the `query_prefix_id` to each sequence's `input_ids` tensor. When `is_query=False`, it prepends `document_prefix_id` instead. This automatic insertion happens before the transformer processes the tokens, ensuring the model receives properly marked sequences without manual token manipulation.