How PyLate Handles Query and Document Prefix Tokens in ColBERT
PyLate prepends dedicated prefix tokens to queries and documents during encoding, using configurable strings that default to "[Q] " and "[D] " when no custom values or checkpoint metadata are provided.
The lightonai/pylate library implements a ColBERT-style retrieval model that relies on distinct prefix tokens to differentiate between query and document sequences. This design allows the transformer to learn separate representations for questions and passages while maintaining a shared vocabulary. The prefix handling strategy involves automatic token addition to the tokenizer, embedding matrix resizing, and runtime insertion during the encoding process.
How Prefix Tokens Are Configured
PyLate determines the query and document prefix values through a hierarchical fallback system defined in pylate/models/colbert.py.
Custom User Definitions
When initializing the ColBERT class, users can explicitly set prefixes via the query_prefix and document_prefix arguments:
from pylate import models
model = models.ColBERT(
model_name_or_path="jinaai/jina-colbert-v2",
query_prefix="[MyQuery]",
document_prefix="[MyDoc]"
)
Checkpoint Metadata Extraction
If no custom prefixes are provided, PyLate searches for an artifact.metadata file associated with Stanford-NLP-style checkpoints. When present, the model extracts query_token_id and doc_token_id values from this metadata to use as the prefix identifiers.
Default Fallback Tokens
When neither user configuration nor checkpoint metadata is available, PyLate defaults to human-readable strings. The constructor logic (approximately lines 66-71 in colbert.py) implements the following fallback chain:
self.query_prefix = (
query_prefix
if query_prefix is not None
else self.query_prefix
if self.query_prefix is not None
else "[Q] "
)
self.document_prefix = (
document_prefix
if document_prefix is not None
else self.document_prefix
if self.document_prefix is not None
else "[D] "
)
Adding Prefix Tokens to the Tokenizer
Once prefix values are determined, PyLate modifies the tokenizer vocabulary and model embeddings to accommodate these new special tokens.
Vocabulary Extension and Embedding Resizing
The ColBERT constructor performs the following operations (approximately lines 81-86 in colbert.py):
# Initial resize to ensure base vocabulary size
self._first_module().auto_model.resize_token_embeddings(len(self.tokenizer))
# Add the prefix tokens to vocabulary
self.tokenizer.add_tokens([self.query_prefix, self.document_prefix])
# Second resize to account for the two new tokens
self._first_module().auto_model.resize_token_embeddings(len(self.tokenizer))
The double resizing ensures compatibility with tokenizers that may not support dynamic vocabulary expansion, preventing indexing errors when the model processes sequences containing the new prefix tokens.
Token ID Caching
After adding tokens to the vocabulary, PyLate immediately caches the corresponding token IDs for runtime efficiency:
self.document_prefix_id = self.tokenizer.convert_tokens_to_ids(self.document_prefix)
self.query_prefix_id = self.tokenizer.convert_tokens_to_ids(self.query_prefix)
These cached IDs enable O(1) lookup during the encoding phase.
Encoding with Prefix Tokens
During inference, PyLate automatically prepends the appropriate prefix token based on the input type.
Automatic Prefix Insertion
The encode method accepts an is_query boolean parameter that determines which prefix to apply. When is_query=True (the default), the query prefix token ID is prepended to the input_ids tensor; when is_query=False, the document prefix is used instead.
This logic delegates to the underlying SentenceTransformer.tokenize routine, which handles the actual tensor manipulation:
# Queries (default behavior)
query_embeddings = model.encode(["What is PyLate?"])
# Documents (explicitly set is_query=False)
doc_embeddings = model.encode(["PyLate is a retrieval library."], is_query=False)
Manual Prefix Insertion
For advanced use cases requiring direct tensor manipulation, PyLate provides the insert_prefix_token static method in colbert.py (approximately lines 70-81):
@staticmethod
def insert_prefix_token(input_ids: torch.Tensor, prefix_id: int) -> torch.Tensor:
prefix_tensor = torch.full(
size=(input_ids.size(0), 1),
fill_value=prefix_id,
dtype=input_ids.dtype,
device=input_ids.device,
)
return torch.cat([input_ids[:, :1], prefix_tensor, input_ids[:, 1:]], dim=1)
This utility inserts the prefix token immediately after the [CLS] token (position 0) and before the rest of the sequence, maintaining the standard BERT-style input format while marking the sequence type.
Impact on Retrieval Architecture
The prefix token strategy enables several architectural optimizations in PyLate's retrieval pipeline.
Differential Padding Strategies
Queries and documents receive different padding treatments after prefix insertion:
- Queries are padded to a uniform length to facilitate batch processing during similarity computation
- Documents remain unpadded by default, allowing variable-length encoding that preserves the original sequence boundaries marked by the prefix token
This distinction optimizes memory usage for document indexing while maintaining computational efficiency for query processing.
Separation of Representations
By prepending distinct prefix tokens, PyLate enables the transformer to learn query-specific and document-specific attention patterns within a shared parameter space. The [Q] and [D] tokens (or their custom equivalents) act as type identifiers similar to token type embeddings in traditional BERT models, but with the flexibility to use arbitrary string values.
Summary
- PyLate uses dedicated prefix tokens to distinguish queries from documents in ColBERT retrieval models
- Prefix values follow a hierarchy: user-provided arguments → checkpoint metadata → default
"[Q] "and"[D] "strings - The tokenizer vocabulary is dynamically extended to include prefix tokens, with the embedding matrix resized accordingly
- During encoding, the
is_queryparameter determines whether the query or document prefix is automatically prepended to input sequences - The
insert_prefix_tokenstatic method provides manual control for advanced tensor manipulation
Frequently Asked Questions
What happens if I don't specify custom prefix tokens when loading a PyLate model?
If you omit the query_prefix and document_prefix arguments, PyLate first checks for an artifact.metadata file in the checkpoint directory. If found, it extracts the token IDs stored as query_token_id and doc_token_id. If no metadata exists, it falls back to the default human-readable strings "[Q] " and "[D] ", which are then added to the tokenizer vocabulary.
Can I use different prefix tokens for different types of documents?
The standard ColBERT class supports exactly one document prefix and one query prefix per model instance. However, you can implement custom prefix logic by manually calling insert_prefix_token with different prefix IDs before passing tensors to the model, or by maintaining separate model instances with different document_prefix configurations for distinct document types.
Why does PyLate resize the token embeddings twice when adding prefix tokens?
PyLate calls resize_token_embeddings twice as a defensive programming measure. The first call ensures the embedding matrix matches the base tokenizer size before adding new tokens. After calling tokenizer.add_tokens() to insert the prefix strings, the second resize expands the matrix to accommodate the two new token IDs. This double-resize strategy prevents indexing errors when working with tokenizers that may not support dynamic vocabulary expansion.
How does the is_query parameter affect the encoding process?
When you call model.encode(texts, is_query=True), the method forwards the flag to the underlying tokenizer, which prepends the query_prefix_id to each sequence's input_ids tensor. When is_query=False, it prepends document_prefix_id instead. This automatic insertion happens before the transformer processes the tokens, ensuring the model receives properly marked sequences without manual token manipulation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →