Difference Between `encode` and `encode_multi_process` in PyLate: A Complete Guide

The encode method performs single-process encoding on the current device, while encode_multi_process distributes workloads across multiple worker processes and devices using a process pool for large-scale batch inference.

PyLate is an open-source library for efficient late interaction retrieval models, implementing the ColBERT architecture for semantic search. When encoding text passages or queries using the ColBERT class, you have two distinct APIs available: the standard encode method for straightforward single-process execution, and encode_multi_process for parallelized encoding across multiple GPUs or CPU cores. Understanding the architectural differences between these methods is critical for optimizing throughput in production retrieval pipelines.

Core Architectural Differences

Single-Process Execution (encode)

The encode method runs entirely within the current Python process, utilizing the model's default device (self.device) or an explicitly passed device parameter. Located in pylate/models/colbert.py at lines 83-112, this method handles batching internally via the batch_size parameter and returns embeddings as tensors, NumPy arrays, or lists depending on the convert_to_tensor and convert_to_numpy flags.

This approach is ideal for quick prototyping, debugging, or encoding small-to-moderate collections where process overhead would exceed the benefits of parallelization.

Multi-Process Distribution (encode_multi_process)

The encode_multi_process method, defined in pylate/models/colbert.py at lines 28-68, implements a producer-consumer pattern using Python's multiprocessing capabilities. Rather than encoding directly, this method delegates work to a pool of independent worker processes created via start_multi_process_pool().

Each worker process runs _encode_multi_process_worker from pylate/utils/multi_process.py (lines 74-112), which executes the standard encode method on its assigned chunk of data. The parent process splits the input into chunks (configurable via chunk_size), distributes them across workers, and concatenates the results using np.concatenate before returning a flat list of NumPy arrays.

Performance and Scalability Comparison

Device Handling

The encode method operates on a single device specified at initialization or passed as an argument. In contrast, encode_multi_process enables multi-device scaling: when you call start_multi_process_pool(), the utility function _start_multi_process_pool in pylate/utils/multi_process.py (lines 13-71) automatically detects available GPUs and creates one worker per GPU, or defaults to four CPU processes if no GPUs are present.

Memory and Chunking

Memory management differs significantly between the two approaches. The encode method loads batches into GPU memory sequentially based on batch_size. The multi-process variant introduces an additional chunk_size parameter that determines how many sentences each worker processes before returning results. This chunking prevents memory exhaustion on individual workers when processing millions of documents, as each worker only materializes its assigned subset.

Implementation Details

Source Code Locations

The dual-API architecture is implemented across two primary files:

  • pylate/models/colbert.py: Contains both encode_multi_process (lines 28-68) and encode (lines 83-112) methods of the ColBERT class.
  • pylate/utils/multi_process.py: Houses the pool management utilities including _start_multi_process_pool (lines 13-71) and _encode_multi_process_worker (lines 74-112).

Worker Pool Mechanics

When you invoke start_multi_process_pool(), the method initializes a Pool from Python's multiprocessing module. Each worker in this pool executes _encode_multi_process_worker, which receives a tuple containing the chunk index, sentences, and encoding parameters. The worker calls the standard encode method on its chunk and returns the results via the pool's queue mechanism. The parent process aggregates these chunks in order, ensuring the final output matches the input sequence despite parallel execution.

Practical Usage Examples

Single-Process Encoding

Use encode for development, testing, or small-scale inference:

from pylate import models

# Initialize model on GPU

model = models.ColBERT(
    "sentence-transformers/all-MiniLM-L6-v2",
    device="cuda",
)

sentences = [
    "What is the capital of France?",
    "Explain quantum entanglement in simple terms.",
]

# Direct single-process encoding

embeddings = model.encode(
    sentences,
    batch_size=16,
    normalize_embeddings=True,
)

print(f"Encoded {len(embeddings)} documents")

Multi-Process Encoding

Use encode_multi_process for production indexing of large corpora:

from pylate import models

model = models.ColBERT(
    "sentence-transformers/all-MiniLM-L6-v2",
    device="cpu",  # Will be overridden by pool workers

)

# Initialize multi-process pool (auto-detects GPUs)

pool = model.start_multi_process_pool()

# Large corpus for indexing

large_corpus = [f"Document content number {i}" for i in range(200_000)]

# Distributed encoding across all available devices

embeddings = model.encode_multi_process(
    sentences=large_corpus,
    pool=pool,
    batch_size=32,
    chunk_size=5000,  # Each worker processes 5k docs at a time

    normalize_embeddings=True,
)

# Clean up worker processes

model.stop_multi_process_pool(pool)

print(f"Successfully encoded {len(embeddings)} vectors using multi-process distribution")

Summary

  • encode provides single-process execution suitable for prototyping and small datasets, operating on a single device with straightforward batching.
  • encode_multi_process enables distributed encoding across multiple GPUs or CPU cores via a process pool, utilizing chunking to handle large-scale corpora efficiently.
  • Both methods ultimately execute the same core encoding logic in pylate/models/colbert.py, but differ in process architecture and resource utilization.
  • Multi-process encoding requires explicit pool management via start_multi_process_pool() and stop_multi_process_pool().

Frequently Asked Questions

When should I use encode versus encode_multi_process?

Use encode when processing small-to-moderate datasets (thousands of documents), during development and debugging, or when working in resource-constrained environments where spawning multiple processes adds unnecessary overhead. Use encode_multi_process when indexing large corpora (hundreds of thousands or millions of documents) across multiple GPUs, as the parallel processing significantly reduces total encoding time despite the overhead of process initialization.

Does encode_multi_process produce different results than encode?

No, both methods produce identical embeddings given the same input and model weights. The encode_multi_process method delegates to the same underlying encode logic within each worker process, as implemented in _encode_multi_process_worker in pylate/utils/multi_process.py. The only difference is the distribution strategy; the mathematical operations remain consistent across both APIs.

How does device allocation work in multi-process encoding?

When you call start_multi_process_pool() without specifying devices, the utility function _start_multi_process_pool in pylate/utils/multi_process.py automatically detects available CUDA GPUs and creates one worker per GPU. If no GPUs are detected, it defaults to spawning four CPU processes. You can also manually specify device lists to control which hardware each worker utilizes, enabling flexible deployment across heterogeneous environments.

What is the purpose of the chunk_size parameter in encode_multi_process?

The chunk_size parameter determines how many sentences each worker processes before returning results to the parent process. This parameter controls memory usage granularity: smaller chunks reduce peak memory consumption per worker but increase inter-process communication overhead, while larger chunks improve throughput but require more memory. The default behavior automatically selects an appropriate chunk size if not specified, balancing memory efficiency with encoding speed for large-scale document processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →