How to Create Custom Pipeline Workflows Using the PipelineController in Stirling-PDF
You can create custom pipeline workflows by POSTing a JSON configuration describing your desired operations together with PDF files to the /api/v1/pipeline/handleData endpoint, which the PipelineController processes sequentially using the internal PipelineProcessor.
Stirling-PDF enables you to chain multiple PDF-processing operations—such as repair, OCR, compression, and watermarking—into a single automated workflow. The PipelineController (PipelineController.java) provides the REST API entry point for these custom pipelines, allowing you to execute complex document transformations without writing glue code.
Understanding the Pipeline Architecture
The pipeline system consists of a controller layer, configuration models, and an internal processor that orchestrates execution across existing API endpoints.
API Entry Point and Routing
The PipelineController class is decorated with the @PipelineApi meta-annotation, which automatically registers it as a @RestController under the base path /api/v1/pipeline【/app/common/src/main/java/stirling/software/common/annotations/api/PipelineApi.java#L19-L21】.
The primary method handleData (lines 48-66 in PipelineController.java) accepts multipart requests at POST /api/v1/pipeline/handleData【/app/core/src/main/java/stirling/software/SPDF/controller/api/pipeline/PipelineController.java#L48-L66】. This method expects:
fileInput: One or more PDF files asMultipartFile[]json: A JSON string describing the pipeline configuration
Configuration Models
The JSON payload deserializes into a PipelineConfig object, which contains:
- name: Human-readable identifier for the workflow
- outputDir: Optional override for the output directory
- outputPattern: Optional naming pattern for output files
- operations: Ordered list of
PipelineOperationobjects【/app/core/src/main/java/stirling/software/SPDF/model/PipelineConfig.java#L9-L20】
Each PipelineOperation specifies:
- operation: The API endpoint path (e.g.,
/api/v1/misc/repair) - parameters: Map of arguments passed directly to that endpoint【/app/core/src/main/java/stirling/software/SPDF/model/PipelineOperation.java#L7-L11】
Creating a Custom Pipeline JSON
To define a workflow, create a JSON file that lists your desired operations in execution order. The following example processes invoices through repair, OCR, compression, and watermarking:
{
"name": "Invoice Processing",
"pipeline": [
{
"operation": "/api/v1/misc/repair",
"parameters": {}
},
{
"operation": "/api/v1/ocr/ocr-pdf",
"parameters": {
"language": "eng",
"outputFileName": "ocred"
}
},
{
"operation": "/api/v1/misc/compress-pdf",
"parameters": {
"optimizeLevel": 2
}
},
{
"operation": "/api/v1/watermark/add",
"parameters": {
"text": "CONFIDENTIAL",
"fontSize": 30,
"color": "#FF0000"
}
}
],
"outputDir": "processed",
"outputFileName": "processed_invoice.pdf"
}
Key fields explained:
- operation: Must match an existing Stirling-PDF API endpoint path
- parameters: JSON object containing the exact parameters that endpoint expects
- outputDir: Optional directory name for organizing results
- outputFileName: Optional base name for the final output (the system appends counters for duplicates using
GeneralUtils.generateFilename)
Executing Pipelines via API
Once your JSON configuration is ready, invoke the pipeline by POSTing to the handleData endpoint with your files and configuration:
curl -X POST "http://localhost:8080/api/v1/pipeline/handleData" \
-F "fileInput=@/path/to/invoice1.pdf" \
-F "fileInput=@/path/to/invoice2.pdf" \
-F "json=$(cat invoice-pipeline.json)" \
-H "Accept: application/octet-stream" \
-o processed_output.zip
Execution flow:
- The
PipelineControllerparses the multipart request and deserializes the JSON into aPipelineConfigobject processor.generateInputFiles(files)converts uploadedMultipartFiles into SpringResourceobjectsprocessor.runPipelineAgainstFiles(inputFiles, config)iterates throughconfig.getOperations(), internally invoking each specified REST endpoint in sequence- The system packages results into a ZIP archive if multiple files are processed, using
GeneralUtils.generateFilenameto deduplicate names【/app/core/src/main/java/stirling/software/SPDF/controller/api/pipeline/PipelineController.java#L99-L110】
Integrating with the Web UI
To make custom pipelines available through the Stirling-PDF web interface, place your JSON definition files in the default web UI configurations directory:
<runtime-path>/pipeline/defaultWebUIConfigs/
The exact location is resolved by runtimePathConfig.getPipelineDefaultWebUiConfigs() as defined in RuntimePathConfig.java【/app/common/src/main/java/stirling/software/common/configuration/RuntimePathConfig.java#L47-L55】. The UIDataController.getPipelineData() method exposes these files to the frontend, making them selectable from the Automation → Pipeline menu【/app/core/src/main/java/stirling/software/SPDF/controller/api/UIDataController.java#L110-L124】.
Summary
- PipelineController provides the REST endpoint at
/api/v1/pipeline/handleDatafor executing custom workflows - PipelineConfig and PipelineOperation define the JSON structure for specifying operation sequences and parameters
- Submit workflows via multipart POST requests containing PDF files and a JSON configuration describing the operation chain
- Store reusable pipeline definitions in
<runtime-path>/pipeline/defaultWebUIConfigs/to expose them through the web interface - The system automatically handles duplicate filenames using
GeneralUtils.generateFilenamewhen packaging multiple outputs into ZIP archives
Frequently Asked Questions
Where do I store custom pipeline JSON files to make them available in the web UI?
Place your JSON files in the directory resolved by runtimePathConfig.getPipelineDefaultWebUiConfigs(), typically located at <runtime-path>/pipeline/defaultWebUIConfigs/. The UIDataController serves these files to the frontend pipeline selector, allowing users to choose predefined workflows from the Automation menu.
How does the PipelineController handle duplicate filenames when processing multiple files?
When packaging multiple outputs into a ZIP archive, the controller uses GeneralUtils.generateFilename to deduplicate filenames by appending numeric counters (e.g., file(1).pdf, file(2).pdf). This logic executes inside the ZIP-building loop in PipelineController.java to ensure unique entries within the archive.
What happens if a pipeline operation fails during execution?
If any operation in the sequence fails, the PipelineProcessor propagates the exception back to the PipelineController, which logs the error and returns an HTTP 500 response. The pipeline does not implement automatic rollback or partial output cleanup, so you should implement client-side error handling and consider adding validation steps at the beginning of your pipeline JSON.
Can I override output directories and filenames in the pipeline configuration?
Yes. Include the optional outputDir and outputFileName fields in your PipelineConfig JSON. The outputDir specifies a subdirectory for organizing results, while outputFileName sets the base name for the final output. The system applies these settings during the result packaging phase in PipelineController.handleData.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →