# Handling Large Document Ingestion Using Parallel Processing Modes in PrivateGPT

> Discover how to efficiently process massive document collections with PrivateGPT"s parallel processing modes. Optimize your ingestion scaling for seamless large document handling.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: how-to-guide
- Published: 2026-03-06

---

**PrivateGPT provides four configurable ingestion modes—simple, batch, parallel, and pipeline—that scale from single-threaded debugging to memory-bounded, multi-process pipelines capable of ingesting massive document collections without exhausting system resources.**

PrivateGPT's ingestion pipeline is architected to scale from single-file uploads to enterprise-scale document collections. When handling large document ingestion using parallel processing modes, the system leverages configurable concurrency strategies that balance CPU utilization, GPU acceleration, and memory constraints through specialized process pools and bounded queues.

## The Four Parallel Processing Modes Explained

The ingestion strategy is controlled by the `ingest_mode` setting in `EmbeddingSettings`, with the implementation residing in [`private_gpt/components/ingest/ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/ingest/ingest_component.py). The factory function `get_ingestion_component` instantiates one of four specialized classes based on your configuration.

### Simple Mode

**Simple mode** processes documents sequentially in the calling thread. It reads, transforms, and indexes one file at a time without spawning additional processes.

Use this mode for small collections, debugging ingestion logic, or environments constrained to a single CPU core. It eliminates concurrency-related complexity, making stack traces easier to interpret when troubleshooting parser errors.

### Batch Mode

**Batch mode** parallelizes the file-to-document parsing phase across multiple workers, then submits the resulting documents to the embedding model in large batches. This approach is optimized for **GPU-accelerated embedding models** that achieve higher throughput with batched inputs.

The CPU handles I/O-bound parsing concurrently, while the GPU processes embeddings efficiently. This mode is ideal when your embedding model supports large batch sizes and you want to maximize GPU utilization without CPU bottlenecks during parsing.

### Parallel Mode

**Parallel mode** utilizes a **process pool** (`multiprocessing.Pool`) to execute both the file-to-document transformation and the embedding computation across multiple CPU cores simultaneously. It represents the fastest local setup for CPU-plus-GPU environments when sufficient RAM is available.

Key implementation details in [`private_gpt/components/ingest/ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/ingest/ingest_component.py) include:
- A process pool for the `_file_to_documents_work_pool` transformation step
- A thread-safe lock (`_index_thread_lock`) protecting vector store insertions from concurrent write conflicts
- Automatic disabling of `TOKENIZERS_PARALLELISM` to prevent deadlocks with Hugging Face tokenizers

### Pipeline Mode

**Pipeline mode** implements a **producer-consumer architecture** designed for massive corpora where memory constraints prohibit loading all documents simultaneously. It uses bounded queues and semaphores to enforce back-pressure and prevent memory exhaustion.

The architecture consists of:
- **Producer threads**: Parse files into documents and place them in a bounded queue (`doc_q`)
- **Worker threads**: Transform documents into nodes (embeddings) and place results in a second bounded queue (`node_q`)
- **Consumer thread**: Batches nodes and flushes them to the vector store when reaching `NODE_FLUSH_COUNT` (5,000 nodes)

A semaphore (`self.doc_semaphore`) limits concurrent embedding jobs to the worker count, ensuring memory usage remains predictable regardless of corpus size.

## Configuring Parallel Processing for Large Document Ingestion

### Selecting the Ingestion Mode

Configure your strategy via the [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) file or environment variables. The `EmbeddingSettings` class in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) defines the `ingest_mode` field:

```yaml
embedding:
  ingest_mode: parallel    # Options: simple, batch, parallel, pipeline

  count_workers: 8          # Match to your CPU core count

```

The factory function `get_ingestion_component` reads these settings and instantiates the appropriate component class at runtime.

### Tuning Worker Counts

The `count_workers` parameter (defaulting to 2) controls process pool size in parallel mode and thread pool size in pipeline mode. As documented in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py), you should set this value to match your physical CPU core count, but avoid exceeding it to prevent memory exhaustion and context-switching overhead.

For a machine with 8 physical cores:

```yaml
embedding:
  ingest_mode: parallel
  count_workers: 8

```

### Memory Safeguards

Both **parallel** and **pipeline** modes automatically disable Hugging Face tokenizer parallelism by setting `os.environ["TOKENIZERS_PARALLELISM"] = "false"`. This prevents deadlocks that occur when tokenizer subprocesses conflict with the ingestion process pools.

In **pipeline** mode, memory pressure is further controlled through:
- Bounded queues (`doc_q`, `node_q`) with maximum sizes
- A semaphore limiting concurrent embedding jobs to the worker count
- Batch flushing every 5,000 nodes (`NODE_FLUSH_COUNT`)

## Implementation Examples

### Instantiating the Ingestion Component

Use the factory function to select the appropriate component based on your configuration:

```python
from private_gpt.components.ingest.ingest_component import get_ingestion_component
from private_gpt.settings.settings import settings

# Load configuration

cfg = settings()

# Initialize storage context and embedding model

storage_context = ...
embed_model = ...

# Get the configured ingestion component

ingest_component = get_ingestion_component(
    storage_context=storage_context,
    embed_model=embed_model,
    transformations=[...],  # Includes embedding transformation

    settings=cfg,
)

```

The factory inspects `cfg.embedding.ingest_mode` and returns an instance of `SimpleIngestComponent`, `BatchIngestComponent`, `ParallelIngestComponent`, or `PipelineIngestComponent`.

### Processing Documents with Parallel Mode

For maximum throughput on multi-core systems with sufficient RAM:

```python

# Configure for parallel processing

# settings.yaml:

# embedding:

#   ingest_mode: parallel

#   count_workers: 8

files = [
    ("report.pdf", Path("/data/report.pdf")),
    ("article.txt", Path("/data/article.txt")),
    # ... thousands of files

]

# Execute parallel ingestion

results = ingest_component.bulk_ingest(files)
print(f"Successfully ingested {len(results)} documents")

```

This creates a `multiprocessing.Pool` to parse files concurrently while a thread-safe lock (`_index_thread_lock`) serializes writes to the vector store.

### Handling Massive Corpora with Pipeline Mode

When ingesting terabyte-scale collections with limited memory:

```python

# Configure for pipeline processing

# settings.yaml:

# embedding:

#   ingest_mode: pipeline

#   count_workers: 4

# The pipeline automatically manages back-pressure

# using bounded queues and flushes every 5,000 nodes

ingest_component.bulk_ingest(file_list)

```

The pipeline uses two bounded queues (`doc_q` and `node_q`) and a semaphore to ensure memory usage remains constant regardless of input size, flushing nodes to the vector store in 5,000-node batches.

## Summary

- **PrivateGPT** offers four distinct ingestion modes—**simple**, **batch**, **parallel**, and **pipeline**—to handle document collections ranging from single files to massive corpora.
- **Parallel mode** utilizes `multiprocessing.Pool` and thread-safe locks (`_index_thread_lock`) to maximize CPU and GPU utilization for medium-to-large datasets with sufficient RAM.
- **Pipeline mode** implements a producer-consumer architecture with bounded queues (`doc_q`, `node_q`), semaphores, and batch flushing (`NODE_FLUSH_COUNT = 5000`) to constrain memory usage during terabyte-scale ingestion.
- Configure modes via `settings.embedding.ingest_mode` and tune concurrency with `count_workers`, matching physical CPU cores while avoiding memory exhaustion.
- Both parallel and pipeline modes automatically disable `TOKENIZERS_PARALLELISM` to prevent deadlocks with Hugging Face tokenizers.

## Frequently Asked Questions

### What is the best parallel processing mode for ingesting millions of documents?

**Pipeline mode** is specifically designed for massive corpora where memory constraints prevent loading all documents simultaneously. It uses bounded queues and semaphores to enforce back-pressure, ensuring constant memory usage regardless of corpus size, while still utilizing all available CPU cores through its producer-consumer architecture.

### How does parallel mode prevent data corruption when multiple processes write to the vector store?

Parallel mode implements a **thread-safe lock** (`_index_thread_lock`) that serializes write operations to the vector store. While document parsing and embedding occur concurrently across a process pool (`multiprocessing.Pool`), the index insertion step acquires the lock to ensure only one process writes at a time, preventing race conditions and data corruption.

### Can I adjust the number of workers without modifying the source code?

Yes, worker concurrency is fully configurable through the **settings file** ([`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml)) or environment variables. The `count_workers` parameter under `embedding` controls process pool size in parallel mode and thread pool size in pipeline mode. The default is 2, but you should set it to match your physical CPU core count for optimal performance, avoiding values that exceed available memory.

### What happens if I run out of memory during large document ingestion?

If memory exhaustion occurs, switch from **parallel mode** to **pipeline mode**, which caps memory usage through bounded queues (`doc_q`, `node_q`) and semaphores regardless of input size. Additionally, reduce `count_workers` to limit concurrent embedding jobs, and ensure `TOKENIZERS_PARALLELISM` remains disabled (handled automatically) to prevent tokenizer subprocesses from consuming additional memory.