# How the PrivateGPT Document Ingestion Pipeline Works: Simple, Batch, Parallel, and Pipeline Modes

> Discover how PrivateGPT's document ingestion pipeline optimizes processing with simple, batch, parallel, and pipeline modes. Learn about its pluggable architecture and configuration settings.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: internals
- Published: 2026-03-06

---

**PrivateGPT's document ingestion pipeline uses a pluggable architecture that switches between four processing modes—simple, batch, parallel, and pipeline—based on the `embedding.ingest_mode` configuration setting, allowing optimization for everything from debugging small files to high-throughput GPU-accelerated ingestion of large document collections.**

The document ingestion pipeline in PrivateGPT (zylon-ai/private-gpt) provides a flexible, configuration-driven system for processing documents into vector embeddings. Built around a service façade and pluggable backend components, the pipeline adapts its execution strategy—from single-threaded processing to multi-stage streaming—based on the `embedding.ingest_mode` setting defined in your YAML configuration.

## How the Document Ingestion Pipeline Selects Processing Modes

### Configuration-Driven Component Selection

The mode selection logic resides in [`private_gpt/components/ingest/ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/ingest/ingest_component.py) within the `get_ingestion_component` function (lines 88-113). This function inspects the `settings.embedding.ingest_mode` value and instantiates the appropriate concrete component:

```python
ingest_mode = settings.embedding.ingest_mode
if ingest_mode == "batch":
    # Returns BatchIngestComponent

elif ingest_mode == "parallel":
    # Returns ParallelizedIngestComponent

elif ingest_mode == "pipeline":
    # Returns PipelineIngestComponent

else:
    # Default → SimpleIngestComponent

```

Changing the `embedding.ingest_mode` value in your [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) to `simple`, `batch`, `parallel`, or `pipeline` swaps the underlying implementation without requiring code changes elsewhere in the application.

### The IngestService Façade

All document ingestion pipeline modes expose a uniform API through `IngestService` in [`private_gpt/server/ingest/ingest_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/ingest/ingest_service.py). This service receives the selected component at construction time and delegates operations to it:

- **`ingest_file`** – Processes a single file from disk
- **`ingest_text`** – Creates a temporary file from raw text and ingests it
- **`bulk_ingest`** – Processes multiple files in one operation
- **`list_ingested`** – Queries the vector store for stored documents
- **`delete`** – Removes a document from the index

All components inherit from `BaseIngestComponentWithIndex`, which handles index initialization, persistence, and provides a thread lock (`_index_thread_lock`) to protect concurrent writes.

## The Four Document Ingestion Modes Explained

### Simple Mode: Single-Threaded Processing

`SimpleIngestComponent` provides the most straightforward document ingestion pipeline implementation, suitable for debugging and small workloads.

**Workflow:**
1. Parse the file into `Document` objects using `IngestionHelper.transform_file_into_documents`
2. Insert each document into the index via `self._index.insert`
3. Persist the index and docstore

**Source:** `SimpleIngestComponent.ingest` in [`private_gpt/components/ingest/ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/ingest/ingest_component.py) (lines 20-28) and the `_save_docs` helper (lines 38-46).

This mode runs single-threaded and processes files sequentially, making it predictable but slower for large collections.

### Batch Mode: CPU-Parallel Parsing with GPU Batching

`BatchIngestComponent` accelerates the document ingestion pipeline by parallelizing file parsing and batching embedding computations.

**Workflow:**
1. Transform files in parallel using a `multiprocessing.Pool` (`self._file_to_documents_work_pool`)
2. Collect all `Document` objects from the pool
3. Run transformations (including embedding generation) on the entire batch via `run_transformations`
4. Insert nodes in bulk using `self._index.insert_nodes`

**Configuration:** Set `count_workers` in your settings to control the multiprocessing pool size.

**Source:** `BatchIngestComponent.__init__` (lines 66-73) and bulk ingestion implementation (lines 87-100) in [`private_gpt/components/ingest/ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/ingest/ingest_component.py).

This mode maximizes GPU utilization by submitting large embedding batches while using CPU cores for file I/O and parsing.

### Parallel Mode: Overlapping I/O and GPU Work

`ParallelizedIngestComponent` creates a document ingestion pipeline that overlaps file parsing with embedding computation to reduce wall-clock time.

**Workflow:**
1. File-to-documents conversion runs in a dedicated process pool (`_file_to_documents_work_pool`)
2. The resulting `Document` list feeds into a thread pool (`_ingest_work_pool`) that calls the component's own `ingest` method for each file
3. This keeps CPU cores busy with parsing while the GPU processes embeddings simultaneously
4. Nodes insert as in the batch component

**Source:** `ParallelizedIngestComponent.__init__` (lines 40-50) and `bulk_ingest` implementation (lines 73-81) in [`private_gpt/components/ingest/ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/ingest/ingest_component.py).

This mode excels with mixed-size workloads where file parsing times vary, ensuring the GPU never waits for the CPU to finish reading files.

### Pipeline Mode: Streamed Processing with Bounded Memory

`PipelineIngestComponent` implements a fully streamed document ingestion pipeline that maintains constant memory usage regardless of collection size.

**Workflow:**
1. **Reader thread** parses files into `Document` chunks and pushes them onto `doc_q` (a queue)
2. **Embedding workers** (a `ThreadPool`) pull from `doc_q`, apply the full transformation pipeline (including embeddings), and push resulting nodes onto `node_q`
3. **Writer thread** drains `node_q`, accumulating nodes until reaching `NODE_FLUSH_COUNT` (default 5000), then flushes them to the index in one I/O-heavy batch

**Source:** `PipelineIngestComponent.__init__` (lines 36-48) for queue and thread setup, and the core loops in `_doc_to_node`, `_doc_to_node_worker`, `_write_nodes`, and `_save_docs` (lines 77-112) in [`private_gpt/components/ingest/ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/ingest/ingest_component.py).

This mode is ideal for ingesting large document collections (millions of pages) where loading everything into memory would exhaust RAM.

## Implementing the Document Ingestion Pipeline in Your Application

### Configuration Example

Switch modes by modifying your [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) file without changing application code:

```yaml
embedding:
  ingest_mode: pipeline  # Options: simple, batch, parallel, pipeline

  count_workers: 4       # Controls parallel workers for batch/parallel modes

```

For pipeline mode, the `NODE_FLUSH_COUNT` constant (default 5000) controls how many nodes accumulate before writing to disk, which you can modify in the source if needed.

### Programmatic Usage

The dependency injection container handles component selection automatically based on your configuration:

```python
from private_gpt.di import global_injector
from private_gpt.server.ingest.ingest_service import IngestService
from pathlib import Path

# Resolve the service (DI container builds it with the selected mode)

ingest_srv: IngestService = global_injector.get(IngestService)

# Ingest a single PDF (the service routes to the configured component)

docs = ingest_srv.ingest_file("contract.pdf", Path("/data/contracts/contract.pdf"))
print(f"Ingested {len(docs)} documents")

# Bulk ingest a folder (behavior depends on your settings.yaml mode)

folder = Path("/data/reports")
files = [(p.name, p) for p in folder.rglob("*.*") if p.is_file()]
bulk_docs = ingest_srv.bulk_ingest(files)
print(f"Bulk-ingested {len(bulk_docs)} documents")

```

The `global_injector` instantiates the appropriate component—whether `SimpleIngestComponent`, `BatchIngestComponent`, `ParallelizedIngestComponent`, or `PipelineIngestComponent`—based solely on the `embedding.ingest_mode` value in your settings.

## Summary

- PrivateGPT's document ingestion pipeline uses a **pluggable component architecture** selected via the `embedding.ingest_mode` configuration setting.
- **Simple mode** processes files sequentially in a single thread, ideal for debugging and small workloads.
- **Batch mode** parallelizes file parsing across CPU cores while batching embeddings for GPU efficiency.
- **Parallel mode** overlaps I/O-bound file parsing with GPU-bound embedding generation using separate process and thread pools.
- **Pipeline mode** implements a three-stage streaming architecture (reader → embedding workers → writer) that maintains constant memory usage regardless of collection size.
- All modes expose the same API through `IngestService` in [`private_gpt/server/ingest/ingest_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/ingest/ingest_service.py), allowing seamless switching via configuration changes alone.

## Frequently Asked Questions

### What is the difference between batch and parallel document ingestion modes?

**Batch mode** uses a multiprocessing pool to parse files in parallel, then collects all documents before running batched transformations and inserting nodes in bulk. **Parallel mode** adds a second layer of concurrency by using a thread pool to process documents while the process pool continues parsing files, effectively overlapping CPU-bound parsing with GPU-bound embedding generation. Choose batch mode for maximum GPU utilization when files are similar in size, and parallel mode when file parsing times vary significantly.

### When should I use pipeline mode for document ingestion?

Use **pipeline mode** when ingesting large document collections that would exhaust available memory if loaded entirely into RAM. This mode implements a streaming architecture with a reader thread, embedding worker threads, and a writer thread that flushes nodes to disk in batches of 5000 (controlled by `NODE_FLUSH_COUNT`). It maintains constant memory usage regardless of input size, making it ideal for production deployments processing millions of pages.

### How do I configure the number of workers for parallel processing?

Set the `count_workers` parameter in your [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) file under the `embedding` section. This value controls the size of the multiprocessing pool used by batch and parallel modes for file parsing, and the thread pool size for embedding workers in pipeline mode. For example, setting `count_workers: 4` creates four parallel workers. Adjust this based on your CPU core count and GPU memory capacity.

### Can I change the document ingestion mode without restarting the application?

No, the ingestion mode is selected at application startup when the dependency injection container instantiates the `IngestService`. The `get_ingestion_component` function in [`private_gpt/components/ingest/ingest_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/ingest/ingest_component.py) evaluates `settings.embedding.ingest_mode` once during initialization to determine whether to inject `SimpleIngestComponent`, `BatchIngestComponent`, `ParallelizedIngestComponent`, or `PipelineIngestComponent`. To switch modes, modify your [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml) and restart the application.