How to Use PaddleOCR with Different Image Formats: JPG, PNG, PDF, and GIF Support
PaddleOCR natively supports JPG, PNG, BMP, TIFF, GIF, and PDF files through built-in utilities in ppocr/utils/utility.py that handle format detection, validation, and decoding automatically without requiring manual conversion.
PaddleOCR's inference pipeline accepts diverse image containers out-of-the-box, eliminating the need for pre-processing conversion scripts. Whether you are processing single photos, multi-page PDFs, or animated GIFs, the framework's internal helpers in the PaddlePaddle/PaddleOCR repository manage format-specific loading transparently. This guide explains the architecture behind this flexibility and provides practical examples for both the command-line interface and Python API.
Supported Image Formats and Validation
PaddleOCR recognizes nine distinct file extensions through the _check_image_file helper function located at lines 62-64 of ppocr/utils/utility.py. The supported set includes:
- Standard images:
.jpg,.jpeg,.png,.bmp,.rgb,.tif,.tiff - Special formats:
.gif(first frame only),.pdf(all pages)
When you provide an input path, the framework validates the extension against this set. Any file outside these formats triggers an informative exception stating not found any img file, preventing silent failures during batch processing.
Core Utilities for Image Discovery and Loading
The flexibility to handle directories, list files, and format-specific decoding stems from three complementary utilities in ppocr/utils/utility.py:
get_image_file_list (Lines 67-94)
This function resolves user input into a sorted list of valid image paths. It accepts:
- A single image file path
- A directory containing multiple images
- A text file (
.txt) containing one image path per line
For each candidate path, it validates file existence and filters by the supported extension set using _check_image_file.
check_and_read (Lines 19-52)
This helper handles special decoding requirements for non-standard formats:
- GIF files: Opens via
cv2.VideoCaptureand extracts the first frame as a BGR numpy array - PDF files: Opens with PyMuPDF (
fitz), renders each page at 2× scaling (capped at 2000px width), and returns a list of page images as BGR numpy arrays - Standard formats: Returns
Noneflags, indicating the caller should fall back to standardcv2.imread
How the Inference Pipeline Uses These Utilities
The inference scripts tools/infer/predict_rec.py and tools/infer/predict_det.py demonstrate the operational workflow. They import the utilities:
from ppocr.utils.utility import get_image_file_list, check_and_read
The main processing loop follows this pattern:
- Discovery:
get_image_file_list(args.image_dir)resolves the input into a flat list of image paths - Loading:
check_and_read(img_path)returns(img, is_gif, is_pdf)flags - Iteration: For PDFs, the code iterates over the returned list of page images; for single images, it processes the array directly
- Preprocessing: Images pass through
ResizeImage,NormalizeImage, and other transforms before reaching the model predictor
This architecture means users never manually convert PDFs or extract GIF frames—the framework handles decoding transparently.
Code Examples for Different Image Formats
Command-Line Interface
The paddleocr entry-point automatically invokes the utility functions, supporting all formats without extra flags:
# Single image (any supported format)
paddleocr --image_dir ./invoice.png
# Directory of mixed formats
paddleocr --image_dir ./batch_input/
# Text file listing paths (one per line)
paddleocr --image_dir ./image_paths.txt
Python API: Single Image (PNG/JPG)
For standard image files, the API returns a nested list of bounding boxes and recognized text:
from paddleocr import PaddleOCR
ocr = PaddleOCR(use_angle_cls=True, lang='en')
result = ocr.ocr('documents/page_1.png', cls=True)
# Extract text and confidence
for line in result:
bbox, (text, confidence) = line
print(f"{text}: {confidence:.2f}")
Python API: Multi-Page PDF
When processing PDFs, the API returns a list where each element corresponds to one page:
from paddleocr import PaddleOCR
ocr = PaddleOCR(use_angle_cls=True, lang='en')
pdf_path = 'reports/document.pdf'
# Returns list of pages, each containing OCR results
pages = ocr.ocr(pdf_path, cls=True)
for page_idx, page_results in enumerate(pages):
print(f"--- Page {page_idx + 1} ---")
for line in page_results:
bbox, (text, confidence) = line
print(f"{text} ({confidence:.2f})")
Python API: Animated GIF
GIF processing extracts only the first frame, treating it as a static image:
from paddleocr import PaddleOCR
ocr = PaddleOCR(lang='en')
result = ocr.ocr('input/animation.gif')
# Result contains OCR from first frame only
print(result[0][1][0]) # Print first detected text string
Summary
- PaddleOCR supports nine formats natively—JPG, PNG, BMP, RGB, TIF, TIFF, GIF, and PDF—validated via
_check_image_fileinppocr/utils/utility.py - PDF handling uses PyMuPDF to rasterize pages at 2× scaling, returning a list of page images for sequential processing
- GIF handling extracts the first frame using
cv2.VideoCapture, discarding subsequent animation frames - Input flexibility allows passing single files, directories, or
.txtlist files viaget_image_file_list - Zero conversion required—the
check_and_readutility manages all format-specific decoding before images enter the pre-processing pipeline
Frequently Asked Questions
Does PaddleOCR support PDF files with multiple pages?
Yes. According to the source code in ppocr/utils/utility.py, PaddleOCR uses PyMuPDF (fitz) to open PDF documents and renders each page into a BGR numpy array. The check_and_read function returns a list of images (one per page), and inference scripts process each page sequentially. When using the Python API, ocr.ocr('file.pdf') returns a list where each element contains the OCR results for that specific page.
Can PaddleOCR process animated GIFs?
Yes, but with limitations. The framework opens GIF files using cv2.VideoCapture and extracts only the first frame for OCR processing. Subsequent frames are ignored. This behavior is implemented in the check_and_read function at lines 28-32 of ppocr/utils/utility.py. If you need to OCR multiple frames, you must manually split the GIF into individual images before processing.
What happens if I pass an unsupported image format?
PaddleOCR raises an exception with the message not found any img file. The _check_image_file function validates extensions against the supported set {jpg, bmp, png, jpeg, rgb, tif, tiff, gif, pdf}. This validation occurs within get_image_file_list, which filters the input list before any inference begins, preventing the model from attempting to load unsupported binary data.
How does PaddleOCR handle directories containing mixed image types?
The get_image_file_list function in ppocr/utils/utility.py scans directories and automatically filters files by supported extensions. It returns a sorted list containing only valid image paths, silently skipping unsupported files. This allows you to point the --image_dir argument at a folder containing JPGs, PNGs, and PDFs together, and PaddleOCR will process all compatible files without manual sorting.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →