How the PrivateGPT Document Ingestion Pipeline Works: Simple, Batch, Parallel, and Pipeline Modes

PrivateGPT's document ingestion pipeline uses a pluggable architecture that switches between four processing modes—simple, batch, parallel, and pipeline—based on the embedding.ingest_mode configuration setting, allowing optimization for everything from debugging small files to high-throughput GPU-accelerated ingestion of large document collections.

The document ingestion pipeline in PrivateGPT (zylon-ai/private-gpt) provides a flexible, configuration-driven system for processing documents into vector embeddings. Built around a service façade and pluggable backend components, the pipeline adapts its execution strategy—from single-threaded processing to multi-stage streaming—based on the embedding.ingest_mode setting defined in your YAML configuration.

How the Document Ingestion Pipeline Selects Processing Modes

Configuration-Driven Component Selection

The mode selection logic resides in private_gpt/components/ingest/ingest_component.py within the get_ingestion_component function (lines 88-113). This function inspects the settings.embedding.ingest_mode value and instantiates the appropriate concrete component:

ingest_mode = settings.embedding.ingest_mode
if ingest_mode == "batch":
    # Returns BatchIngestComponent

elif ingest_mode == "parallel":
    # Returns ParallelizedIngestComponent

elif ingest_mode == "pipeline":
    # Returns PipelineIngestComponent

else:
    # Default → SimpleIngestComponent

Changing the embedding.ingest_mode value in your settings.yaml to simple, batch, parallel, or pipeline swaps the underlying implementation without requiring code changes elsewhere in the application.

The IngestService Façade

All document ingestion pipeline modes expose a uniform API through IngestService in private_gpt/server/ingest/ingest_service.py. This service receives the selected component at construction time and delegates operations to it:

  • ingest_file – Processes a single file from disk
  • ingest_text – Creates a temporary file from raw text and ingests it
  • bulk_ingest – Processes multiple files in one operation
  • list_ingested – Queries the vector store for stored documents
  • delete – Removes a document from the index

All components inherit from BaseIngestComponentWithIndex, which handles index initialization, persistence, and provides a thread lock (_index_thread_lock) to protect concurrent writes.

The Four Document Ingestion Modes Explained

Simple Mode: Single-Threaded Processing

SimpleIngestComponent provides the most straightforward document ingestion pipeline implementation, suitable for debugging and small workloads.

Workflow:

  1. Parse the file into Document objects using IngestionHelper.transform_file_into_documents
  2. Insert each document into the index via self._index.insert
  3. Persist the index and docstore

Source: SimpleIngestComponent.ingest in private_gpt/components/ingest/ingest_component.py (lines 20-28) and the _save_docs helper (lines 38-46).

This mode runs single-threaded and processes files sequentially, making it predictable but slower for large collections.

Batch Mode: CPU-Parallel Parsing with GPU Batching

BatchIngestComponent accelerates the document ingestion pipeline by parallelizing file parsing and batching embedding computations.

Workflow:

  1. Transform files in parallel using a multiprocessing.Pool (self._file_to_documents_work_pool)
  2. Collect all Document objects from the pool
  3. Run transformations (including embedding generation) on the entire batch via run_transformations
  4. Insert nodes in bulk using self._index.insert_nodes

Configuration: Set count_workers in your settings to control the multiprocessing pool size.

Source: BatchIngestComponent.__init__ (lines 66-73) and bulk ingestion implementation (lines 87-100) in private_gpt/components/ingest/ingest_component.py.

This mode maximizes GPU utilization by submitting large embedding batches while using CPU cores for file I/O and parsing.

Parallel Mode: Overlapping I/O and GPU Work

ParallelizedIngestComponent creates a document ingestion pipeline that overlaps file parsing with embedding computation to reduce wall-clock time.

Workflow:

  1. File-to-documents conversion runs in a dedicated process pool (_file_to_documents_work_pool)
  2. The resulting Document list feeds into a thread pool (_ingest_work_pool) that calls the component's own ingest method for each file
  3. This keeps CPU cores busy with parsing while the GPU processes embeddings simultaneously
  4. Nodes insert as in the batch component

Source: ParallelizedIngestComponent.__init__ (lines 40-50) and bulk_ingest implementation (lines 73-81) in private_gpt/components/ingest/ingest_component.py.

This mode excels with mixed-size workloads where file parsing times vary, ensuring the GPU never waits for the CPU to finish reading files.

Pipeline Mode: Streamed Processing with Bounded Memory

PipelineIngestComponent implements a fully streamed document ingestion pipeline that maintains constant memory usage regardless of collection size.

Workflow:

  1. Reader thread parses files into Document chunks and pushes them onto doc_q (a queue)
  2. Embedding workers (a ThreadPool) pull from doc_q, apply the full transformation pipeline (including embeddings), and push resulting nodes onto node_q
  3. Writer thread drains node_q, accumulating nodes until reaching NODE_FLUSH_COUNT (default 5000), then flushes them to the index in one I/O-heavy batch

Source: PipelineIngestComponent.__init__ (lines 36-48) for queue and thread setup, and the core loops in _doc_to_node, _doc_to_node_worker, _write_nodes, and _save_docs (lines 77-112) in private_gpt/components/ingest/ingest_component.py.

This mode is ideal for ingesting large document collections (millions of pages) where loading everything into memory would exhaust RAM.

Implementing the Document Ingestion Pipeline in Your Application

Configuration Example

Switch modes by modifying your settings.yaml file without changing application code:

embedding:
  ingest_mode: pipeline  # Options: simple, batch, parallel, pipeline

  count_workers: 4       # Controls parallel workers for batch/parallel modes

For pipeline mode, the NODE_FLUSH_COUNT constant (default 5000) controls how many nodes accumulate before writing to disk, which you can modify in the source if needed.

Programmatic Usage

The dependency injection container handles component selection automatically based on your configuration:

from private_gpt.di import global_injector
from private_gpt.server.ingest.ingest_service import IngestService
from pathlib import Path

# Resolve the service (DI container builds it with the selected mode)

ingest_srv: IngestService = global_injector.get(IngestService)

# Ingest a single PDF (the service routes to the configured component)

docs = ingest_srv.ingest_file("contract.pdf", Path("/data/contracts/contract.pdf"))
print(f"Ingested {len(docs)} documents")

# Bulk ingest a folder (behavior depends on your settings.yaml mode)

folder = Path("/data/reports")
files = [(p.name, p) for p in folder.rglob("*.*") if p.is_file()]
bulk_docs = ingest_srv.bulk_ingest(files)
print(f"Bulk-ingested {len(bulk_docs)} documents")

The global_injector instantiates the appropriate component—whether SimpleIngestComponent, BatchIngestComponent, ParallelizedIngestComponent, or PipelineIngestComponent—based solely on the embedding.ingest_mode value in your settings.

Summary

  • PrivateGPT's document ingestion pipeline uses a pluggable component architecture selected via the embedding.ingest_mode configuration setting.
  • Simple mode processes files sequentially in a single thread, ideal for debugging and small workloads.
  • Batch mode parallelizes file parsing across CPU cores while batching embeddings for GPU efficiency.
  • Parallel mode overlaps I/O-bound file parsing with GPU-bound embedding generation using separate process and thread pools.
  • Pipeline mode implements a three-stage streaming architecture (reader → embedding workers → writer) that maintains constant memory usage regardless of collection size.
  • All modes expose the same API through IngestService in private_gpt/server/ingest/ingest_service.py, allowing seamless switching via configuration changes alone.

Frequently Asked Questions

What is the difference between batch and parallel document ingestion modes?

Batch mode uses a multiprocessing pool to parse files in parallel, then collects all documents before running batched transformations and inserting nodes in bulk. Parallel mode adds a second layer of concurrency by using a thread pool to process documents while the process pool continues parsing files, effectively overlapping CPU-bound parsing with GPU-bound embedding generation. Choose batch mode for maximum GPU utilization when files are similar in size, and parallel mode when file parsing times vary significantly.

When should I use pipeline mode for document ingestion?

Use pipeline mode when ingesting large document collections that would exhaust available memory if loaded entirely into RAM. This mode implements a streaming architecture with a reader thread, embedding worker threads, and a writer thread that flushes nodes to disk in batches of 5000 (controlled by NODE_FLUSH_COUNT). It maintains constant memory usage regardless of input size, making it ideal for production deployments processing millions of pages.

How do I configure the number of workers for parallel processing?

Set the count_workers parameter in your settings.yaml file under the embedding section. This value controls the size of the multiprocessing pool used by batch and parallel modes for file parsing, and the thread pool size for embedding workers in pipeline mode. For example, setting count_workers: 4 creates four parallel workers. Adjust this based on your CPU core count and GPU memory capacity.

Can I change the document ingestion mode without restarting the application?

No, the ingestion mode is selected at application startup when the dependency injection container instantiates the IngestService. The get_ingestion_component function in private_gpt/components/ingest/ingest_component.py evaluates settings.embedding.ingest_mode once during initialization to determine whether to inject SimpleIngestComponent, BatchIngestComponent, ParallelizedIngestComponent, or PipelineIngestComponent. To switch modes, modify your settings.yaml and restart the application.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →