How the Stirling-PDF Pipeline System Works for Automated PDF Processing Workflows

The Stirling-PDF pipeline system enables automated PDF processing workflows by monitoring watched folders for incoming files, executing a chain of REST API operations defined in JSON configuration files, and delivering results to a finished folder.

The open-source Stirling-Tools/Stirling-PDF repository includes a powerful pipeline subsystem that transforms standalone PDF operations into fully automated sequences. Administrators define processing workflows—such as repair, sanitize, and compress—using simple JSON configuration files, while the backend handles file watching, execution orchestration, and result delivery without manual intervention.

Architecture Overview

The pipeline implementation centers on three core components working in concert to process documents asynchronously.

Pipeline Configuration stored in JSON files that conform to the PipelineConfig POJO ([PipelineConfig.java](https://github.com/Stirling-Tools/Stirling-PDF/blob/main/app/core/src/main/java/stirling/software/SPDF/model/PipelineConfig.java)) define the ordered sequence of operations.

File-Watching and Execution handled by background services such as [TelegramPipelineBot.java](https://github.com/Stirling-Tools/Stirling-PDF/blob/main/app/core/src/main/java/stirling/software/SPDF/service/telegram/TelegramPipelineBot.java) monitor inbox directories, trigger processing, and poll for completion.

Path Resolution managed by [RuntimePathConfig.java](https://github.com/Stirling-Tools/Stirling-PDF/blob/main/app/common/src/main/java/stirling/software/common/configuration/RuntimePathConfig.java) converts application properties into concrete filesystem locations for watched folders, finished folders, and default configurations.

Defining Pipeline Workflows with JSON

A pipeline definition is a JSON object that mirrors the PipelineConfig model structure. Each file specifies an ordered array of operations that map directly to Stirling-PDF's REST API endpoints.

The configuration schema includes:

  • name – Human-readable identifier for the workflow
  • pipeline – Array of PipelineOperation objects containing operation (API endpoint) and parameters
  • outputDir – Target subdirectory for processed files
  • outputFileName – Pattern for naming results, supporting placeholders like {originalName}
{
  "name": "Prepare-pdf-for-email",
  "pipeline": [
    { "operation": "/api/v1/misc/repair", "parameters": {} },
    { "operation": "/api/v1/security/sanitize-pdf", "parameters": { "removeJavaScript": true } },
    { "operation": "/api/v1/misc/compress-pdf", "parameters": { "optimizeLevel": 2 } }
  ],
  "outputDir": "my-output",
  "outputFileName": "final-{originalName}"
}

Pipeline JSON files reside in the default UI config directory (pipeline/defaultWebUIConfigs). During startup, GeneralUtils.extractPipeline() (source lines 945-964) extracts built-in workflow templates from the classpath to this location. Administrators can add custom JSON files here to create new automated workflows without modifying source code.

Directory Structure and Path Resolution

When the Spring container initializes, RuntimePathConfig resolves three critical filesystem paths (lines 60-78) that govern pipeline behavior:

this.pipelineWatchedFoldersPath =   // e.g., <install>/pipeline/watchedFolders
this.pipelineFinishedFoldersPath = // e.g., <install>/pipeline/finishedFolders  
this.pipelineDefaultWebUiConfigs = // e.g., <install>/pipeline/defaultWebUIConfigs

These paths resolve from application.yml under the system.customPaths.pipeline.* hierarchy or fall back to default subdirectories within the installation root. The separation of concerns allows the system to ingest files from watched locations, process them asynchronously, and deposit results into distinct finished folders for retrieval.

Execution and File Watching Mechanisms

The pipeline system uses a polling-based execution model that supports multiple ingestion methods, including Telegram bots and automated drop-folders.

The Telegram Bot Consumer

TelegramPipelineBot provides a concrete implementation of the pipeline consumer pattern. When users send PDFs via Telegram, the bot executes the following sequence:

  1. Inbox Resolution – getInboxFolder(chatId) constructs the path <watchedFolder>/<telegramInboxFolder> (default "telegram") and creates missing directories
  2. Configuration Validation – hasJsonConfig(chatId) verifies that at least one *.json pipeline definition exists in the target inbox
  3. File Download – Incoming files are saved to the inbox directory via downloadFile
  4. Result Polling – waitForPipelineOutputs(info) polls the finished folder using a synchronized pipelinePollMonitor until a file matching the unique base name (original filename + UUID) appears or processingTimeoutSeconds expires

The polling loop in waitForPipelineOutputs (lines 46-78) respects configurable pollingIntervalMillis values from application.yml:

while (Duration.between(start, Instant.now()).compareTo(timeout) <= 0) {
    // list finished folder → filter by base name & modification time
}

Global Auto-Pipeline Mode

For deployments without Telegram, the ApplicationProperties.AutoPipeline object (lines 151-165) enables a drop-folder mode that processes any PDF placed into watched directories automatically. This configuration uses the same directory layout and polling logic but triggers without chat-based interaction.

API Integration and Pipeline Discovery

The web interface consumes pipeline configurations through the UIDataController REST endpoint. The getPipelineData() method (lines 110-161) scans runtimePathConfig.getPipelineDefaultWebUiConfigs() for all *.json files and returns two data structures:

  • pipelineConfigs – Raw JSON strings for frontend rendering
  • pipelineConfigsWithNames – Structured pairs of { name, json } extracted from the JSON name field or filename

Access available workflows via:

curl -X GET http://localhost:8080/api/v1/ui-data/pipeline \
     -H "Accept: application/json"

Response format:

{
  "pipelineConfigsWithNames": [
    { "name": "Prepare-pdfs-for-email", "json": "{...}" },
    { "name": "OCR-and-Compress", "json": "{...}" }
  ]
}

This endpoint enables the pipeline editor UI to display available automated workflows without direct filesystem access.

Result Handling and Resource Cleanup

The PipelineResult class implements AutoCloseable to manage temporary file lifecycles during multi-stage processing. As operations execute, the result object tracks intermediate TempFile instances and ensures cleanup via the close() method upon completion or error.

The result object also exposes metadata properties:

  • hasErrors – Boolean flag indicating processing failures
  • filtersApplied – Tracking for applied transformations

This design prevents temporary file leaks during long-running automated workflows and providesstructured error reporting for downstream consumers.

Summary

  • JSON-based configuration – Define automated workflows in pipeline/defaultWebUIConfigs using the PipelineConfig schema that chains REST API operations
  • Filesystem-driven execution – The system monitors pipeline/watchedFolders for incoming PDFs and deposits results into pipeline/finishedFolders
  • Flexible ingestion – Process files via Telegram bot (TelegramPipelineBot), global auto-pipeline mode, or custom drop-folder implementations
  • Resource safety – PipelineResult implements AutoCloseable to ensure temporary files clean up automatically after processing
  • RESTful management – Query available pipelines via /api/v1/ui-data/pipeline and configure paths through application.yml under system.customPaths.pipeline.*

Frequently Asked Questions

How do I create a custom automated PDF workflow in Stirling-PDF?

Create a JSON file following the PipelineConfig schema and place it in <SPDF_INSTALL>/pipeline/defaultWebUIConfigs/. The JSON must include a name field, a pipeline array containing PipelineOperation objects with operation (API endpoint) and parameters, plus optional outputDir and outputFileName fields. The system automatically detects new configurations on startup or via the UI upload function.

What triggers the pipeline system to start processing a PDF?

The pipeline triggers when any PDF file appears in a configured watched folder. For Telegram integration, TelegramPipelineBot.handleIncomingFile saves uploaded documents to <watchedFolder>/telegram/<chatId>/ and begins polling. For auto-pipeline mode, dropping a file directly into the watched folder triggers the same processing loop without user interaction, provided a JSON configuration exists in the target directory.

Where does Stirling-PDF store processed pipeline results?

Processed files appear in the pipeline/finishedFolders directory configured in RuntimePathConfig. The system constructs output paths using the outputDir and outputFileName patterns from your JSON configuration. The Telegram bot specifically monitors this location via waitForPipelineOutputs to detect when files matching the expected UUID-patterned basename are ready for retrieval.

Can I run the pipeline system without using the Telegram bot?

Yes. Enable the autoPipeline configuration in application.yml by setting autoPipeline.enabled: true. This activates a global watched folder mode that processes any PDF placed into the configured directories automatically using the first matching JSON configuration, eliminating the need for chat-based interaction while maintaining the same backend execution logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →