Difference Between `encode` and `encode_multi_process` in PyLate: A Complete Guide
The encode method performs single-process encoding on the current device, while encode_multi_process distributes workloads across multiple worker processes and devices using a process pool for large-scale batch inference.
PyLate is an open-source library for efficient late interaction retrieval models, implementing the ColBERT architecture for semantic search. When encoding text passages or queries using the ColBERT class, you have two distinct APIs available: the standard encode method for straightforward single-process execution, and encode_multi_process for parallelized encoding across multiple GPUs or CPU cores. Understanding the architectural differences between these methods is critical for optimizing throughput in production retrieval pipelines.
Core Architectural Differences
Single-Process Execution (encode)
The encode method runs entirely within the current Python process, utilizing the model's default device (self.device) or an explicitly passed device parameter. Located in pylate/models/colbert.py at lines 83-112, this method handles batching internally via the batch_size parameter and returns embeddings as tensors, NumPy arrays, or lists depending on the convert_to_tensor and convert_to_numpy flags.
This approach is ideal for quick prototyping, debugging, or encoding small-to-moderate collections where process overhead would exceed the benefits of parallelization.
Multi-Process Distribution (encode_multi_process)
The encode_multi_process method, defined in pylate/models/colbert.py at lines 28-68, implements a producer-consumer pattern using Python's multiprocessing capabilities. Rather than encoding directly, this method delegates work to a pool of independent worker processes created via start_multi_process_pool().
Each worker process runs _encode_multi_process_worker from pylate/utils/multi_process.py (lines 74-112), which executes the standard encode method on its assigned chunk of data. The parent process splits the input into chunks (configurable via chunk_size), distributes them across workers, and concatenates the results using np.concatenate before returning a flat list of NumPy arrays.
Performance and Scalability Comparison
Device Handling
The encode method operates on a single device specified at initialization or passed as an argument. In contrast, encode_multi_process enables multi-device scaling: when you call start_multi_process_pool(), the utility function _start_multi_process_pool in pylate/utils/multi_process.py (lines 13-71) automatically detects available GPUs and creates one worker per GPU, or defaults to four CPU processes if no GPUs are present.
Memory and Chunking
Memory management differs significantly between the two approaches. The encode method loads batches into GPU memory sequentially based on batch_size. The multi-process variant introduces an additional chunk_size parameter that determines how many sentences each worker processes before returning results. This chunking prevents memory exhaustion on individual workers when processing millions of documents, as each worker only materializes its assigned subset.
Implementation Details
Source Code Locations
The dual-API architecture is implemented across two primary files:
pylate/models/colbert.py: Contains bothencode_multi_process(lines 28-68) andencode(lines 83-112) methods of theColBERTclass.pylate/utils/multi_process.py: Houses the pool management utilities including_start_multi_process_pool(lines 13-71) and_encode_multi_process_worker(lines 74-112).
Worker Pool Mechanics
When you invoke start_multi_process_pool(), the method initializes a Pool from Python's multiprocessing module. Each worker in this pool executes _encode_multi_process_worker, which receives a tuple containing the chunk index, sentences, and encoding parameters. The worker calls the standard encode method on its chunk and returns the results via the pool's queue mechanism. The parent process aggregates these chunks in order, ensuring the final output matches the input sequence despite parallel execution.
Practical Usage Examples
Single-Process Encoding
Use encode for development, testing, or small-scale inference:
from pylate import models
# Initialize model on GPU
model = models.ColBERT(
"sentence-transformers/all-MiniLM-L6-v2",
device="cuda",
)
sentences = [
"What is the capital of France?",
"Explain quantum entanglement in simple terms.",
]
# Direct single-process encoding
embeddings = model.encode(
sentences,
batch_size=16,
normalize_embeddings=True,
)
print(f"Encoded {len(embeddings)} documents")
Multi-Process Encoding
Use encode_multi_process for production indexing of large corpora:
from pylate import models
model = models.ColBERT(
"sentence-transformers/all-MiniLM-L6-v2",
device="cpu", # Will be overridden by pool workers
)
# Initialize multi-process pool (auto-detects GPUs)
pool = model.start_multi_process_pool()
# Large corpus for indexing
large_corpus = [f"Document content number {i}" for i in range(200_000)]
# Distributed encoding across all available devices
embeddings = model.encode_multi_process(
sentences=large_corpus,
pool=pool,
batch_size=32,
chunk_size=5000, # Each worker processes 5k docs at a time
normalize_embeddings=True,
)
# Clean up worker processes
model.stop_multi_process_pool(pool)
print(f"Successfully encoded {len(embeddings)} vectors using multi-process distribution")
Summary
encodeprovides single-process execution suitable for prototyping and small datasets, operating on a single device with straightforward batching.encode_multi_processenables distributed encoding across multiple GPUs or CPU cores via a process pool, utilizing chunking to handle large-scale corpora efficiently.- Both methods ultimately execute the same core encoding logic in
pylate/models/colbert.py, but differ in process architecture and resource utilization. - Multi-process encoding requires explicit pool management via
start_multi_process_pool()andstop_multi_process_pool().
Frequently Asked Questions
When should I use encode versus encode_multi_process?
Use encode when processing small-to-moderate datasets (thousands of documents), during development and debugging, or when working in resource-constrained environments where spawning multiple processes adds unnecessary overhead. Use encode_multi_process when indexing large corpora (hundreds of thousands or millions of documents) across multiple GPUs, as the parallel processing significantly reduces total encoding time despite the overhead of process initialization.
Does encode_multi_process produce different results than encode?
No, both methods produce identical embeddings given the same input and model weights. The encode_multi_process method delegates to the same underlying encode logic within each worker process, as implemented in _encode_multi_process_worker in pylate/utils/multi_process.py. The only difference is the distribution strategy; the mathematical operations remain consistent across both APIs.
How does device allocation work in multi-process encoding?
When you call start_multi_process_pool() without specifying devices, the utility function _start_multi_process_pool in pylate/utils/multi_process.py automatically detects available CUDA GPUs and creates one worker per GPU. If no GPUs are detected, it defaults to spawning four CPU processes. You can also manually specify device lists to control which hardware each worker utilizes, enabling flexible deployment across heterogeneous environments.
What is the purpose of the chunk_size parameter in encode_multi_process?
The chunk_size parameter determines how many sentences each worker processes before returning results to the parent process. This parameter controls memory usage granularity: smaller chunks reduce peak memory consumption per worker but increase inter-process communication overhead, while larger chunks improve throughput but require more memory. The default behavior automatically selects an appropriate chunk size if not specified, balancing memory efficiency with encoding speed for large-scale document processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →