Understanding the `pool_factor` Parameter in PyLate: A Complete Guide

The pool_factor parameter in PyLate controls how aggressively the ColBERT model compresses document embeddings by clustering consecutive tokens, keeping only 1/pool_factor of the original tokens to reduce memory usage and speed up retrieval.

When working with the lightonai/pylate library for late interaction retrieval, managing embedding dimensionality is critical for production deployments. The pool_factor parameter provides a direct lever to trade off between embedding granularity and computational efficiency, allowing you to reduce the sequence length of document representations without retraining the underlying ColBERT model.

What Is the pool_factor Parameter in PyLate?

The pool_factor is an integer argument passed to the encode() method in pylate/models/colbert.py that determines the degree of hierarchical pooling applied to document token embeddings. When pool_factor exceeds 1, the model groups consecutive token embeddings into clusters and replaces each cluster with its mean vector, effectively retaining only 1 / pool_factor of the original tokens.

  • pool_factor = 1 – Disables pooling entirely; every token embedding is preserved.
  • pool_factor = 2 – Retains approximately 50% of token embeddings.
  • pool_factor = 3 – Retains approximately 33% of token embeddings, and so on.

This compression occurs only during document encoding; query embeddings remain un-pooled to maintain dense representations necessary for accurate late interaction scoring.

How pool_factor Affects Document Embeddings

No Pooling: pool_factor = 1

When pool_factor is set to 1 (the default), the pool_embeddings_hierarchical function in pylate/models/colbert.py (lines 794-860) bypasses the clustering logic entirely. The document embedding tensor retains its original shape [sequence_length, embedding_dimension], preserving the fine-grained token-level representations that ColBERT uses for maximum retrieval accuracy.

Aggressive Compression: pool_factor > 1

Setting pool_factor to values greater than 1 triggers the hierarchical pooling mechanism:

  1. The function calculates the number of clusters as max(num_embeddings // pool_factor, 1).
  2. It performs k-means-style clustering via matrix multiplication to assign tokens to clusters.
  3. It averages the embeddings within each cluster to create representative vectors.
  4. It concatenates these pooled vectors with any protected prefix tokens.

This process dramatically reduces memory consumption for large document collections and accelerates downstream similarity search operations using FAISS or similar vector indexes.

Protected Tokens Exclusion

The pooling logic respects the protected_tokens parameter (defaulting to 1), which excludes the initial tokens—typically the CLS token—from clustering. This preservation ensures that sequence-level semantic information captured by special tokens remains intact even when aggressive pooling reduces the remaining token representations.

Internal Implementation in PyLate

The core pooling implementation resides in the ColBERT class within pylate/models/colbert.py. The pool_embeddings_hierarchical method (spanning approximately lines 794-860) orchestrates the compression:


# Conceptual flow based on pylate/models/colbert.py implementation

def pool_embeddings_hierarchical(embeddings, pool_factor, protected_tokens=1):
    # 1. Split protected prefix from poolable tokens

    protected = embeddings[:protected_tokens]
    poolable = embeddings[protected_tokens:]
    
    # 2. Calculate target number of clusters

    num_clusters = max(poolable.shape[0] // pool_factor, 1)
    
    # 3. Perform clustering and averaging (simplified representation)

    # Actual implementation uses efficient matrix operations

    pooled = cluster_and_average(poolable, num_clusters)
    
    # 4. Reconstruct final embedding

    return concatenate([protected, pooled])

The encode() method accepts pool_factor as a parameter and propagates it through the encoding pipeline, including multi-process handling defined in pylate/utils/multi_process.py for parallel document processing.

Practical Code Examples

The following example demonstrates how to encode documents with different pool_factor values to compare compression levels:

from pylate.models.colbert import ColBERT

# Initialize the ColBERT model from lightonai/pylate

model = ColBERT.from_pretrained("lightonai/colbert-base")

documents = [
    "PyLate is a Python library for efficient late interaction retrieval using ColBERT architectures.",
    "The pool_factor parameter enables memory-efficient document encoding by hierarchical pooling of token embeddings."
]

# Encode with no pooling (pool_factor=1)

full_embeddings = model.encode(
    documents,
    is_query=False,
    pool_factor=1,
    show_progress_bar=True
)

# Encode with 2x pooling (retains ~50% of tokens)

pooled_embeddings = model.encode(
    documents,
    is_query=False,
    pool_factor=2,
    show_progress_bar=True
)

# Compare embedding lengths

print(f"Original tokens per doc: {full_embeddings[0].shape[0]}")
print(f"Pooled tokens per doc:   {pooled_embeddings[0].shape[0]}")

When running this example, you will observe that pooled_embeddings contains approximately half the number of token vectors compared to full_embeddings, directly reflecting the pool_factor=2 compression ratio.

Summary

  • The pool_factor parameter in PyLate controls hierarchical pooling of document token embeddings, retaining only 1/pool_factor of the original tokens.
  • Pooling applies exclusively to documents, not queries, ensuring dense query representations for accurate late interaction scoring.
  • The implementation resides in pylate/models/colbert.py within the pool_embeddings_hierarchical method (lines 794-860), which uses k-means-style clustering via matrix operations.
  • Protected tokens (defaulting to the CLS token) are excluded from pooling to preserve sequence-level semantic information.
  • Increasing pool_factor significantly reduces memory consumption and accelerates FAISS indexing while maintaining retrieval accuracy through semantic clustering.

Frequently Asked Questions

What happens when pool_factor is set to 1 in PyLate?

When pool_factor is set to 1, PyLate disables hierarchical pooling entirely. The pool_embeddings_hierarchical function in pylate/models/colbert.py returns the original document embeddings unchanged, preserving every token vector. This configuration provides maximum retrieval accuracy at the cost of higher memory usage and slower similarity search compared to pooled configurations.

Does pool_factor affect query embeddings or only document embeddings?

The pool_factor parameter affects only document embeddings. According to the implementation in pylate/models/colbert.py, the pooling logic is bypassed when is_query=True is passed to the encode() method. Queries remain dense and uncompressed to ensure accurate late interaction scoring against potentially pooled document representations, maintaining the precision of the retrieval mechanism.

How does PyLate protect special tokens when using pool_factor?

PyLate protects special tokens through the protected_tokens parameter, which defaults to 1. In the pool_embeddings_hierarchical method, the implementation splits the embedding tensor to exclude the first protected_tokens from the clustering operation. This preserves the CLS token or other special prefix tokens that carry sequence-level semantic information, ensuring that critical contextual signals remain intact even when aggressive pooling reduces the remaining token representations.

What is the performance impact of increasing pool_factor in production?

Increasing pool_factor delivers linear reductions in memory consumption and significant speedups in FAISS indexing and retrieval operations. For example, setting pool_factor=2 reduces the number of token vectors by approximately 50%, directly decreasing storage requirements and accelerating nearest-neighbor searches. The hierarchical clustering implemented in pool_embeddings_hierarchical uses efficient matrix multiplication to minimize computational overhead during encoding, making higher pool_factor values particularly beneficial for large-scale document collections where memory constraints dominate deployment decisions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →