# Understanding the `pool_factor` Parameter in PyLate: A Complete Guide

> Master the pool_factor parameter in PyLate to optimize ColBERT model compression. Learn how to reduce memory and boost retrieval speed by intelligently clustering tokens.

- Repository: [LightOn/pylate](https://github.com/lightonai/pylate)
- Tags: deep-dive
- Published: 2026-03-06

---

**The `pool_factor` parameter in PyLate controls how aggressively the ColBERT model compresses document embeddings by clustering consecutive tokens, keeping only `1/pool_factor` of the original tokens to reduce memory usage and speed up retrieval.**

When working with the `lightonai/pylate` library for late interaction retrieval, managing embedding dimensionality is critical for production deployments. The `pool_factor` parameter provides a direct lever to trade off between embedding granularity and computational efficiency, allowing you to reduce the sequence length of document representations without retraining the underlying ColBERT model.

## What Is the `pool_factor` Parameter in PyLate?

The `pool_factor` is an integer argument passed to the `encode()` method in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py) that determines the degree of hierarchical pooling applied to document token embeddings. When `pool_factor` exceeds 1, the model groups consecutive token embeddings into clusters and replaces each cluster with its mean vector, effectively retaining only `1 / pool_factor` of the original tokens.

- **`pool_factor = 1`** – Disables pooling entirely; every token embedding is preserved.
- **`pool_factor = 2`** – Retains approximately 50% of token embeddings.
- **`pool_factor = 3`** – Retains approximately 33% of token embeddings, and so on.

This compression occurs **only during document encoding**; query embeddings remain un-pooled to maintain dense representations necessary for accurate late interaction scoring.

## How `pool_factor` Affects Document Embeddings

### No Pooling: `pool_factor = 1`

When `pool_factor` is set to 1 (the default), the `pool_embeddings_hierarchical` function in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py) (lines 794-860) bypasses the clustering logic entirely. The document embedding tensor retains its original shape `[sequence_length, embedding_dimension]`, preserving the fine-grained token-level representations that ColBERT uses for maximum retrieval accuracy.

### Aggressive Compression: `pool_factor > 1`

Setting `pool_factor` to values greater than 1 triggers the hierarchical pooling mechanism:

1. The function calculates the number of clusters as `max(num_embeddings // pool_factor, 1)`.
2. It performs k-means-style clustering via matrix multiplication to assign tokens to clusters.
3. It averages the embeddings within each cluster to create representative vectors.
4. It concatenates these pooled vectors with any protected prefix tokens.

This process dramatically reduces memory consumption for large document collections and accelerates downstream similarity search operations using FAISS or similar vector indexes.

### Protected Tokens Exclusion

The pooling logic respects the `protected_tokens` parameter (defaulting to 1), which excludes the initial tokens—typically the CLS token—from clustering. This preservation ensures that sequence-level semantic information captured by special tokens remains intact even when aggressive pooling reduces the remaining token representations.

## Internal Implementation in PyLate

The core pooling implementation resides in the `ColBERT` class within [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py). The `pool_embeddings_hierarchical` method (spanning approximately lines 794-860) orchestrates the compression:

```python

# Conceptual flow based on pylate/models/colbert.py implementation

def pool_embeddings_hierarchical(embeddings, pool_factor, protected_tokens=1):
    # 1. Split protected prefix from poolable tokens

    protected = embeddings[:protected_tokens]
    poolable = embeddings[protected_tokens:]
    
    # 2. Calculate target number of clusters

    num_clusters = max(poolable.shape[0] // pool_factor, 1)
    
    # 3. Perform clustering and averaging (simplified representation)

    # Actual implementation uses efficient matrix operations

    pooled = cluster_and_average(poolable, num_clusters)
    
    # 4. Reconstruct final embedding

    return concatenate([protected, pooled])

```

The `encode()` method accepts `pool_factor` as a parameter and propagates it through the encoding pipeline, including multi-process handling defined in [`pylate/utils/multi_process.py`](https://github.com/lightonai/pylate/blob/main/pylate/utils/multi_process.py) for parallel document processing.

## Practical Code Examples

The following example demonstrates how to encode documents with different `pool_factor` values to compare compression levels:

```python
from pylate.models.colbert import ColBERT

# Initialize the ColBERT model from lightonai/pylate

model = ColBERT.from_pretrained("lightonai/colbert-base")

documents = [
    "PyLate is a Python library for efficient late interaction retrieval using ColBERT architectures.",
    "The pool_factor parameter enables memory-efficient document encoding by hierarchical pooling of token embeddings."
]

# Encode with no pooling (pool_factor=1)

full_embeddings = model.encode(
    documents,
    is_query=False,
    pool_factor=1,
    show_progress_bar=True
)

# Encode with 2x pooling (retains ~50% of tokens)

pooled_embeddings = model.encode(
    documents,
    is_query=False,
    pool_factor=2,
    show_progress_bar=True
)

# Compare embedding lengths

print(f"Original tokens per doc: {full_embeddings[0].shape[0]}")
print(f"Pooled tokens per doc:   {pooled_embeddings[0].shape[0]}")

```

When running this example, you will observe that `pooled_embeddings` contains approximately half the number of token vectors compared to `full_embeddings`, directly reflecting the `pool_factor=2` compression ratio.

## Summary

- The **`pool_factor`** parameter in PyLate controls hierarchical pooling of document token embeddings, retaining only `1/pool_factor` of the original tokens.
- **Pooling applies exclusively to documents**, not queries, ensuring dense query representations for accurate late interaction scoring.
- The implementation resides in **[`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py)** within the `pool_embeddings_hierarchical` method (lines 794-860), which uses k-means-style clustering via matrix operations.
- **Protected tokens** (defaulting to the CLS token) are excluded from pooling to preserve sequence-level semantic information.
- Increasing `pool_factor` significantly reduces memory consumption and accelerates FAISS indexing while maintaining retrieval accuracy through semantic clustering.

## Frequently Asked Questions

### What happens when pool_factor is set to 1 in PyLate?

When `pool_factor` is set to 1, PyLate disables hierarchical pooling entirely. The `pool_embeddings_hierarchical` function in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py) returns the original document embeddings unchanged, preserving every token vector. This configuration provides maximum retrieval accuracy at the cost of higher memory usage and slower similarity search compared to pooled configurations.

### Does pool_factor affect query embeddings or only document embeddings?

The `pool_factor` parameter affects **only document embeddings**. According to the implementation in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py), the pooling logic is bypassed when `is_query=True` is passed to the `encode()` method. Queries remain dense and uncompressed to ensure accurate late interaction scoring against potentially pooled document representations, maintaining the precision of the retrieval mechanism.

### How does PyLate protect special tokens when using pool_factor?

PyLate protects special tokens through the `protected_tokens` parameter, which defaults to 1. In the `pool_embeddings_hierarchical` method, the implementation splits the embedding tensor to exclude the first `protected_tokens` from the clustering operation. This preserves the CLS token or other special prefix tokens that carry sequence-level semantic information, ensuring that critical contextual signals remain intact even when aggressive pooling reduces the remaining token representations.

### What is the performance impact of increasing pool_factor in production?

Increasing `pool_factor` delivers linear reductions in memory consumption and significant speedups in FAISS indexing and retrieval operations. For example, setting `pool_factor=2` reduces the number of token vectors by approximately 50%, directly decreasing storage requirements and accelerating nearest-neighbor searches. The hierarchical clustering implemented in `pool_embeddings_hierarchical` uses efficient matrix multiplication to minimize computational overhead during encoding, making higher `pool_factor` values particularly beneficial for large-scale document collections where memory constraints dominate deployment decisions.