# Difference Between `encode` and `encode_multi_process` in PyLate: A Complete Guide

> Understand the difference between PyLate encode and encode_multi_process for efficient batch inference. Learn how multi-process encoding accelerates large-scale tasks.

- Repository: [LightOn/pylate](https://github.com/lightonai/pylate)
- Tags: deep-dive
- Published: 2026-03-06

---

**The `encode` method performs single-process encoding on the current device, while `encode_multi_process` distributes workloads across multiple worker processes and devices using a process pool for large-scale batch inference.**

PyLate is an open-source library for efficient late interaction retrieval models, implementing the ColBERT architecture for semantic search. When encoding text passages or queries using the `ColBERT` class, you have two distinct APIs available: the standard `encode` method for straightforward single-process execution, and `encode_multi_process` for parallelized encoding across multiple GPUs or CPU cores. Understanding the architectural differences between these methods is critical for optimizing throughput in production retrieval pipelines.

## Core Architectural Differences

### Single-Process Execution (`encode`)

The `encode` method runs entirely within the current Python process, utilizing the model's default device (`self.device`) or an explicitly passed device parameter. Located in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py) at lines 83-112, this method handles batching internally via the `batch_size` parameter and returns embeddings as tensors, NumPy arrays, or lists depending on the `convert_to_tensor` and `convert_to_numpy` flags.

This approach is ideal for quick prototyping, debugging, or encoding small-to-moderate collections where process overhead would exceed the benefits of parallelization.

### Multi-Process Distribution (`encode_multi_process`)

The `encode_multi_process` method, defined in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py) at lines 28-68, implements a producer-consumer pattern using Python's multiprocessing capabilities. Rather than encoding directly, this method delegates work to a pool of independent worker processes created via `start_multi_process_pool()`.

Each worker process runs `_encode_multi_process_worker` from [`pylate/utils/multi_process.py`](https://github.com/lightonai/pylate/blob/main/pylate/utils/multi_process.py) (lines 74-112), which executes the standard `encode` method on its assigned chunk of data. The parent process splits the input into chunks (configurable via `chunk_size`), distributes them across workers, and concatenates the results using `np.concatenate` before returning a flat list of NumPy arrays.

## Performance and Scalability Comparison

### Device Handling

The `encode` method operates on a single device specified at initialization or passed as an argument. In contrast, `encode_multi_process` enables multi-device scaling: when you call `start_multi_process_pool()`, the utility function `_start_multi_process_pool` in [`pylate/utils/multi_process.py`](https://github.com/lightonai/pylate/blob/main/pylate/utils/multi_process.py) (lines 13-71) automatically detects available GPUs and creates one worker per GPU, or defaults to four CPU processes if no GPUs are present.

### Memory and Chunking

Memory management differs significantly between the two approaches. The `encode` method loads batches into GPU memory sequentially based on `batch_size`. The multi-process variant introduces an additional `chunk_size` parameter that determines how many sentences each worker processes before returning results. This chunking prevents memory exhaustion on individual workers when processing millions of documents, as each worker only materializes its assigned subset.

## Implementation Details

### Source Code Locations

The dual-API architecture is implemented across two primary files:

- **[`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py)**: Contains both `encode_multi_process` (lines 28-68) and `encode` (lines 83-112) methods of the `ColBERT` class.
- **[`pylate/utils/multi_process.py`](https://github.com/lightonai/pylate/blob/main/pylate/utils/multi_process.py)**: Houses the pool management utilities including `_start_multi_process_pool` (lines 13-71) and `_encode_multi_process_worker` (lines 74-112).

### Worker Pool Mechanics

When you invoke `start_multi_process_pool()`, the method initializes a `Pool` from Python's `multiprocessing` module. Each worker in this pool executes `_encode_multi_process_worker`, which receives a tuple containing the chunk index, sentences, and encoding parameters. The worker calls the standard `encode` method on its chunk and returns the results via the pool's queue mechanism. The parent process aggregates these chunks in order, ensuring the final output matches the input sequence despite parallel execution.

## Practical Usage Examples

### Single-Process Encoding

Use `encode` for development, testing, or small-scale inference:

```python
from pylate import models

# Initialize model on GPU

model = models.ColBERT(
    "sentence-transformers/all-MiniLM-L6-v2",
    device="cuda",
)

sentences = [
    "What is the capital of France?",
    "Explain quantum entanglement in simple terms.",
]

# Direct single-process encoding

embeddings = model.encode(
    sentences,
    batch_size=16,
    normalize_embeddings=True,
)

print(f"Encoded {len(embeddings)} documents")

```

### Multi-Process Encoding

Use `encode_multi_process` for production indexing of large corpora:

```python
from pylate import models

model = models.ColBERT(
    "sentence-transformers/all-MiniLM-L6-v2",
    device="cpu",  # Will be overridden by pool workers

)

# Initialize multi-process pool (auto-detects GPUs)

pool = model.start_multi_process_pool()

# Large corpus for indexing

large_corpus = [f"Document content number {i}" for i in range(200_000)]

# Distributed encoding across all available devices

embeddings = model.encode_multi_process(
    sentences=large_corpus,
    pool=pool,
    batch_size=32,
    chunk_size=5000,  # Each worker processes 5k docs at a time

    normalize_embeddings=True,
)

# Clean up worker processes

model.stop_multi_process_pool(pool)

print(f"Successfully encoded {len(embeddings)} vectors using multi-process distribution")

```

## Summary

- **`encode`** provides single-process execution suitable for prototyping and small datasets, operating on a single device with straightforward batching.
- **`encode_multi_process`** enables distributed encoding across multiple GPUs or CPU cores via a process pool, utilizing chunking to handle large-scale corpora efficiently.
- Both methods ultimately execute the same core encoding logic in [`pylate/models/colbert.py`](https://github.com/lightonai/pylate/blob/main/pylate/models/colbert.py), but differ in process architecture and resource utilization.
- Multi-process encoding requires explicit pool management via `start_multi_process_pool()` and `stop_multi_process_pool()`.

## Frequently Asked Questions

### When should I use `encode` versus `encode_multi_process`?

Use `encode` when processing small-to-moderate datasets (thousands of documents), during development and debugging, or when working in resource-constrained environments where spawning multiple processes adds unnecessary overhead. Use `encode_multi_process` when indexing large corpora (hundreds of thousands or millions of documents) across multiple GPUs, as the parallel processing significantly reduces total encoding time despite the overhead of process initialization.

### Does `encode_multi_process` produce different results than `encode`?

No, both methods produce identical embeddings given the same input and model weights. The `encode_multi_process` method delegates to the same underlying `encode` logic within each worker process, as implemented in `_encode_multi_process_worker` in [`pylate/utils/multi_process.py`](https://github.com/lightonai/pylate/blob/main/pylate/utils/multi_process.py). The only difference is the distribution strategy; the mathematical operations remain consistent across both APIs.

### How does device allocation work in multi-process encoding?

When you call `start_multi_process_pool()` without specifying devices, the utility function `_start_multi_process_pool` in [`pylate/utils/multi_process.py`](https://github.com/lightonai/pylate/blob/main/pylate/utils/multi_process.py) automatically detects available CUDA GPUs and creates one worker per GPU. If no GPUs are detected, it defaults to spawning four CPU processes. You can also manually specify device lists to control which hardware each worker utilizes, enabling flexible deployment across heterogeneous environments.

### What is the purpose of the `chunk_size` parameter in `encode_multi_process`?

The `chunk_size` parameter determines how many sentences each worker processes before returning results to the parent process. This parameter controls memory usage granularity: smaller chunks reduce peak memory consumption per worker but increase inter-process communication overhead, while larger chunks improve throughput but require more memory. The default behavior automatically selects an appropriate chunk size if not specified, balancing memory efficiency with encoding speed for large-scale document processing.