Handling Large Document Ingestion Using Parallel Processing Modes in PrivateGPT
PrivateGPT provides four configurable ingestion modes—simple, batch, parallel, and pipeline—that scale from single-threaded debugging to memory-bounded, multi-process pipelines capable of ingesting massive document collections without exhausting system resources.
PrivateGPT's ingestion pipeline is architected to scale from single-file uploads to enterprise-scale document collections. When handling large document ingestion using parallel processing modes, the system leverages configurable concurrency strategies that balance CPU utilization, GPU acceleration, and memory constraints through specialized process pools and bounded queues.
The Four Parallel Processing Modes Explained
The ingestion strategy is controlled by the ingest_mode setting in EmbeddingSettings, with the implementation residing in private_gpt/components/ingest/ingest_component.py. The factory function get_ingestion_component instantiates one of four specialized classes based on your configuration.
Simple Mode
Simple mode processes documents sequentially in the calling thread. It reads, transforms, and indexes one file at a time without spawning additional processes.
Use this mode for small collections, debugging ingestion logic, or environments constrained to a single CPU core. It eliminates concurrency-related complexity, making stack traces easier to interpret when troubleshooting parser errors.
Batch Mode
Batch mode parallelizes the file-to-document parsing phase across multiple workers, then submits the resulting documents to the embedding model in large batches. This approach is optimized for GPU-accelerated embedding models that achieve higher throughput with batched inputs.
The CPU handles I/O-bound parsing concurrently, while the GPU processes embeddings efficiently. This mode is ideal when your embedding model supports large batch sizes and you want to maximize GPU utilization without CPU bottlenecks during parsing.
Parallel Mode
Parallel mode utilizes a process pool (multiprocessing.Pool) to execute both the file-to-document transformation and the embedding computation across multiple CPU cores simultaneously. It represents the fastest local setup for CPU-plus-GPU environments when sufficient RAM is available.
Key implementation details in private_gpt/components/ingest/ingest_component.py include:
- A process pool for the
_file_to_documents_work_pooltransformation step - A thread-safe lock (
_index_thread_lock) protecting vector store insertions from concurrent write conflicts - Automatic disabling of
TOKENIZERS_PARALLELISMto prevent deadlocks with Hugging Face tokenizers
Pipeline Mode
Pipeline mode implements a producer-consumer architecture designed for massive corpora where memory constraints prohibit loading all documents simultaneously. It uses bounded queues and semaphores to enforce back-pressure and prevent memory exhaustion.
The architecture consists of:
- Producer threads: Parse files into documents and place them in a bounded queue (
doc_q) - Worker threads: Transform documents into nodes (embeddings) and place results in a second bounded queue (
node_q) - Consumer thread: Batches nodes and flushes them to the vector store when reaching
NODE_FLUSH_COUNT(5,000 nodes)
A semaphore (self.doc_semaphore) limits concurrent embedding jobs to the worker count, ensuring memory usage remains predictable regardless of corpus size.
Configuring Parallel Processing for Large Document Ingestion
Selecting the Ingestion Mode
Configure your strategy via the settings.yaml file or environment variables. The EmbeddingSettings class in private_gpt/settings/settings.py defines the ingest_mode field:
embedding:
ingest_mode: parallel # Options: simple, batch, parallel, pipeline
count_workers: 8 # Match to your CPU core count
The factory function get_ingestion_component reads these settings and instantiates the appropriate component class at runtime.
Tuning Worker Counts
The count_workers parameter (defaulting to 2) controls process pool size in parallel mode and thread pool size in pipeline mode. As documented in private_gpt/settings/settings.py, you should set this value to match your physical CPU core count, but avoid exceeding it to prevent memory exhaustion and context-switching overhead.
For a machine with 8 physical cores:
embedding:
ingest_mode: parallel
count_workers: 8
Memory Safeguards
Both parallel and pipeline modes automatically disable Hugging Face tokenizer parallelism by setting os.environ["TOKENIZERS_PARALLELISM"] = "false". This prevents deadlocks that occur when tokenizer subprocesses conflict with the ingestion process pools.
In pipeline mode, memory pressure is further controlled through:
- Bounded queues (
doc_q,node_q) with maximum sizes - A semaphore limiting concurrent embedding jobs to the worker count
- Batch flushing every 5,000 nodes (
NODE_FLUSH_COUNT)
Implementation Examples
Instantiating the Ingestion Component
Use the factory function to select the appropriate component based on your configuration:
from private_gpt.components.ingest.ingest_component import get_ingestion_component
from private_gpt.settings.settings import settings
# Load configuration
cfg = settings()
# Initialize storage context and embedding model
storage_context = ...
embed_model = ...
# Get the configured ingestion component
ingest_component = get_ingestion_component(
storage_context=storage_context,
embed_model=embed_model,
transformations=[...], # Includes embedding transformation
settings=cfg,
)
The factory inspects cfg.embedding.ingest_mode and returns an instance of SimpleIngestComponent, BatchIngestComponent, ParallelIngestComponent, or PipelineIngestComponent.
Processing Documents with Parallel Mode
For maximum throughput on multi-core systems with sufficient RAM:
# Configure for parallel processing
# settings.yaml:
# embedding:
# ingest_mode: parallel
# count_workers: 8
files = [
("report.pdf", Path("/data/report.pdf")),
("article.txt", Path("/data/article.txt")),
# ... thousands of files
]
# Execute parallel ingestion
results = ingest_component.bulk_ingest(files)
print(f"Successfully ingested {len(results)} documents")
This creates a multiprocessing.Pool to parse files concurrently while a thread-safe lock (_index_thread_lock) serializes writes to the vector store.
Handling Massive Corpora with Pipeline Mode
When ingesting terabyte-scale collections with limited memory:
# Configure for pipeline processing
# settings.yaml:
# embedding:
# ingest_mode: pipeline
# count_workers: 4
# The pipeline automatically manages back-pressure
# using bounded queues and flushes every 5,000 nodes
ingest_component.bulk_ingest(file_list)
The pipeline uses two bounded queues (doc_q and node_q) and a semaphore to ensure memory usage remains constant regardless of input size, flushing nodes to the vector store in 5,000-node batches.
Summary
- PrivateGPT offers four distinct ingestion modes—simple, batch, parallel, and pipeline—to handle document collections ranging from single files to massive corpora.
- Parallel mode utilizes
multiprocessing.Pooland thread-safe locks (_index_thread_lock) to maximize CPU and GPU utilization for medium-to-large datasets with sufficient RAM. - Pipeline mode implements a producer-consumer architecture with bounded queues (
doc_q,node_q), semaphores, and batch flushing (NODE_FLUSH_COUNT = 5000) to constrain memory usage during terabyte-scale ingestion. - Configure modes via
settings.embedding.ingest_modeand tune concurrency withcount_workers, matching physical CPU cores while avoiding memory exhaustion. - Both parallel and pipeline modes automatically disable
TOKENIZERS_PARALLELISMto prevent deadlocks with Hugging Face tokenizers.
Frequently Asked Questions
What is the best parallel processing mode for ingesting millions of documents?
Pipeline mode is specifically designed for massive corpora where memory constraints prevent loading all documents simultaneously. It uses bounded queues and semaphores to enforce back-pressure, ensuring constant memory usage regardless of corpus size, while still utilizing all available CPU cores through its producer-consumer architecture.
How does parallel mode prevent data corruption when multiple processes write to the vector store?
Parallel mode implements a thread-safe lock (_index_thread_lock) that serializes write operations to the vector store. While document parsing and embedding occur concurrently across a process pool (multiprocessing.Pool), the index insertion step acquires the lock to ensure only one process writes at a time, preventing race conditions and data corruption.
Can I adjust the number of workers without modifying the source code?
Yes, worker concurrency is fully configurable through the settings file (settings.yaml) or environment variables. The count_workers parameter under embedding controls process pool size in parallel mode and thread pool size in pipeline mode. The default is 2, but you should set it to match your physical CPU core count for optimal performance, avoiding values that exceed available memory.
What happens if I run out of memory during large document ingestion?
If memory exhaustion occurs, switch from parallel mode to pipeline mode, which caps memory usage through bounded queues (doc_q, node_q) and semaphores regardless of input size. Additionally, reduce count_workers to limit concurrent embedding jobs, and ensure TOKENIZERS_PARALLELISM remains disabled (handled automatically) to prevent tokenizer subprocesses from consuming additional memory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →