How to Use PaddleOCR with Different Image Formats: JPG, PNG, PDF, and GIF Support

PaddleOCR natively supports JPG, PNG, BMP, TIFF, GIF, and PDF files through built-in utilities in ppocr/utils/utility.py that handle format detection, validation, and decoding automatically without requiring manual conversion.

PaddleOCR's inference pipeline accepts diverse image containers out-of-the-box, eliminating the need for pre-processing conversion scripts. Whether you are processing single photos, multi-page PDFs, or animated GIFs, the framework's internal helpers in the PaddlePaddle/PaddleOCR repository manage format-specific loading transparently. This guide explains the architecture behind this flexibility and provides practical examples for both the command-line interface and Python API.

Supported Image Formats and Validation

PaddleOCR recognizes nine distinct file extensions through the _check_image_file helper function located at lines 62-64 of ppocr/utils/utility.py. The supported set includes:

  • Standard images: .jpg, .jpeg, .png, .bmp, .rgb, .tif, .tiff
  • Special formats: .gif (first frame only), .pdf (all pages)

When you provide an input path, the framework validates the extension against this set. Any file outside these formats triggers an informative exception stating not found any img file, preventing silent failures during batch processing.

Core Utilities for Image Discovery and Loading

The flexibility to handle directories, list files, and format-specific decoding stems from three complementary utilities in ppocr/utils/utility.py:

get_image_file_list (Lines 67-94)

This function resolves user input into a sorted list of valid image paths. It accepts:

  • A single image file path
  • A directory containing multiple images
  • A text file (.txt) containing one image path per line

For each candidate path, it validates file existence and filters by the supported extension set using _check_image_file.

check_and_read (Lines 19-52)

This helper handles special decoding requirements for non-standard formats:

  • GIF files: Opens via cv2.VideoCapture and extracts the first frame as a BGR numpy array
  • PDF files: Opens with PyMuPDF (fitz), renders each page at 2× scaling (capped at 2000px width), and returns a list of page images as BGR numpy arrays
  • Standard formats: Returns None flags, indicating the caller should fall back to standard cv2.imread

How the Inference Pipeline Uses These Utilities

The inference scripts tools/infer/predict_rec.py and tools/infer/predict_det.py demonstrate the operational workflow. They import the utilities:

from ppocr.utils.utility import get_image_file_list, check_and_read

The main processing loop follows this pattern:

  1. Discovery: get_image_file_list(args.image_dir) resolves the input into a flat list of image paths
  2. Loading: check_and_read(img_path) returns (img, is_gif, is_pdf) flags
  3. Iteration: For PDFs, the code iterates over the returned list of page images; for single images, it processes the array directly
  4. Preprocessing: Images pass through ResizeImage, NormalizeImage, and other transforms before reaching the model predictor

This architecture means users never manually convert PDFs or extract GIF frames—the framework handles decoding transparently.

Code Examples for Different Image Formats

Command-Line Interface

The paddleocr entry-point automatically invokes the utility functions, supporting all formats without extra flags:


# Single image (any supported format)

paddleocr --image_dir ./invoice.png

# Directory of mixed formats

paddleocr --image_dir ./batch_input/

# Text file listing paths (one per line)

paddleocr --image_dir ./image_paths.txt

Python API: Single Image (PNG/JPG)

For standard image files, the API returns a nested list of bounding boxes and recognized text:

from paddleocr import PaddleOCR

ocr = PaddleOCR(use_angle_cls=True, lang='en')
result = ocr.ocr('documents/page_1.png', cls=True)

# Extract text and confidence

for line in result:
    bbox, (text, confidence) = line
    print(f"{text}: {confidence:.2f}")

Python API: Multi-Page PDF

When processing PDFs, the API returns a list where each element corresponds to one page:

from paddleocr import PaddleOCR

ocr = PaddleOCR(use_angle_cls=True, lang='en')
pdf_path = 'reports/document.pdf'

# Returns list of pages, each containing OCR results

pages = ocr.ocr(pdf_path, cls=True)

for page_idx, page_results in enumerate(pages):
    print(f"--- Page {page_idx + 1} ---")
    for line in page_results:
        bbox, (text, confidence) = line
        print(f"{text} ({confidence:.2f})")

Python API: Animated GIF

GIF processing extracts only the first frame, treating it as a static image:

from paddleocr import PaddleOCR

ocr = PaddleOCR(lang='en')
result = ocr.ocr('input/animation.gif')

# Result contains OCR from first frame only

print(result[0][1][0])  # Print first detected text string

Summary

  • PaddleOCR supports nine formats natively—JPG, PNG, BMP, RGB, TIF, TIFF, GIF, and PDF—validated via _check_image_file in ppocr/utils/utility.py
  • PDF handling uses PyMuPDF to rasterize pages at 2× scaling, returning a list of page images for sequential processing
  • GIF handling extracts the first frame using cv2.VideoCapture, discarding subsequent animation frames
  • Input flexibility allows passing single files, directories, or .txt list files via get_image_file_list
  • Zero conversion required—the check_and_read utility manages all format-specific decoding before images enter the pre-processing pipeline

Frequently Asked Questions

Does PaddleOCR support PDF files with multiple pages?

Yes. According to the source code in ppocr/utils/utility.py, PaddleOCR uses PyMuPDF (fitz) to open PDF documents and renders each page into a BGR numpy array. The check_and_read function returns a list of images (one per page), and inference scripts process each page sequentially. When using the Python API, ocr.ocr('file.pdf') returns a list where each element contains the OCR results for that specific page.

Can PaddleOCR process animated GIFs?

Yes, but with limitations. The framework opens GIF files using cv2.VideoCapture and extracts only the first frame for OCR processing. Subsequent frames are ignored. This behavior is implemented in the check_and_read function at lines 28-32 of ppocr/utils/utility.py. If you need to OCR multiple frames, you must manually split the GIF into individual images before processing.

What happens if I pass an unsupported image format?

PaddleOCR raises an exception with the message not found any img file. The _check_image_file function validates extensions against the supported set {jpg, bmp, png, jpeg, rgb, tif, tiff, gif, pdf}. This validation occurs within get_image_file_list, which filters the input list before any inference begins, preventing the model from attempting to load unsupported binary data.

How does PaddleOCR handle directories containing mixed image types?

The get_image_file_list function in ppocr/utils/utility.py scans directories and automatically filters files by supported extensions. It returns a sorted list containing only valid image paths, silently skipping unsupported files. This allows you to point the --image_dir argument at a folder containing JPGs, PNGs, and PDFs together, and PaddleOCR will process all compatible files without manual sorting.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →