How to Use PaddleOCR for End-to-End OCR: Complete Python and CLI Guide

Use the PaddleOCR class to instantiate a PaddleX pipeline that performs text detection and recognition in a single inference call, returning structured results with built-in visualization methods.

PaddleOCR is an open-source optical character recognition system maintained by the PaddlePaddle community. To use PaddleOCR for end-to-end OCR, developers initialize a high-level wrapper that automatically configures detection, recognition, and post-processing models. This pipeline architecture abstracts hardware-specific optimizations while exposing a unified interface for both Python scripting and command-line batch processing.

Quick Start with the Python API

The fastest way to run end-to-end OCR is through the PaddleOCR class defined in paddleocr/_pipelines/ocr.py. This class inherits from PaddleXPipelineWrapper (paddleocr/_pipelines/base.py) and handles all model initialization internally.

from paddleocr import PaddleOCR

# Initialize with default models (PP-OCRv5, English/Chinese mixed)

ocr = PaddleOCR(
    use_doc_orientation_classify=False,
    use_doc_unwarping=False,
    use_textline_orientation=False,
)

# Run inference on a local image or URL

results = ocr.predict("docs/example.png")

# Handle outputs

for line in results:
    line.print()           # Console output: text, confidence, bbox

    line.save_to_img("vis")    # Visualize bounding boxes

    line.save_to_json("out")   # Export structured JSON

During instantiation, the constructor builds a _params dictionary and calls load_pipeline_config("OCR") to fetch default PaddleX configurations. User overrides are merged via _get_paddlex_config_overrides before create_pipeline initializes the actual inference engine.

Internal Architecture and Data Flow

Understanding the internal flow helps debug configuration issues and optimize performance. The end-to-end pipeline follows this execution path:

  1. API Entry: PaddleOCR.__init__ (lines 54-70 in paddleocr/_pipelines/ocr.py) collects model arguments and normalizes deprecated parameters (lines 71-144).

  2. Wrapper Initialization: The base class PaddleXPipelineWrapper (lines 54-71 in paddleocr/_pipelines/base.py) loads the default OCR pipeline configuration.

  3. Config Merging: User-provided dictionaries are deep-merged with defaults using _get_merged_paddlex_config and _merge_dicts (lines 90-101).

  4. Pipeline Creation: The system calls create_pipeline (lines 102-106) to instantiate the hardware-specific inference engine.

  5. Inference: PaddleOCR.predict (lines 98-104) forwards inputs directly to self.paddlex_pipeline.predict, which executes detection and recognition sequentially.

  6. Result Handling: Output objects expose print(), save_to_img(), and save_to_json() methods for immediate consumption.

Hardware selection (CPU, GPU, XPU, NPU) occurs automatically based on the installed PaddlePaddle runtime and the optional device argument defined in paddleocr/_common_args.py.

Command-Line Interface for Batch Processing

For production workflows or shell scripting, PaddleOCR exposes the same functionality through a CLI implemented in paddleocr/_utils/cli.py. The interface uses add_simple_inference_args to parse arguments and perform_simple_inference to execute the pipeline.

paddleocr ocr -i https://github.com/PaddlePaddle/PaddleOCR/raw/main/docs/images/ocr_demo.png \
    --use_doc_orientation_classify False \
    --use_doc_unwarping False \
    --use_textline_orientation False

CLI arguments map directly to the Python API parameters, ensuring consistent behavior across interfaces.

Customizing Models and Languages

To use PaddleOCR for end-to-end OCR with non-default languages or specific model versions, override the lang and ocr_version parameters. The private method _get_ocr_model_names resolves these arguments to specific checkpoint files from the PaddleOCR model zoo.

from paddleocr import PaddleOCR

# Configure for Korean text recognition using PP-OCRv5

ocr = PaddleOCR(
    lang="ko",
    ocr_version="PP-OCRv5",
)

result = ocr.predict("korean_sample.jpg")

Available options for lang include standard language codes (e.g., "en", "ch", "fr", "german"), while ocr_version selects between model generations like "PP-OCRv4" and "PP-OCRv5".

Processing Results Programmatically

Result objects returned by predict() behave like iterables containing dictionaries with text, score, and bbox keys. Extract pure text strings for downstream natural language processing or retrieval-augmented generation (RAG) pipelines:


# Concatenate all detected text lines

full_text = "\n".join([line["text"] for line in result])
print("Extracted content:\n", full_text)

Each result object maintains metadata about the source image and detection coordinates, enabling precise mapping between recognized text and original document regions.

Summary

  • Primary Entry Point: The PaddleOCR class in paddleocr/_pipelines/ocr.py provides the main user interface for end-to-end OCR.
  • Pipeline Architecture: Internally wraps PaddleX pipelines via PaddleXPipelineWrapper (paddleocr/_pipelines/base.py), handling model loading and hardware abstraction.
  • Flexible Interfaces: Supports both Python API and CLI (paddleocr/_utils/cli.py) with identical parameter sets.
  • Rich Outputs: Result objects include print(), save_to_img(), and save_to_json() methods for immediate visualization and data export.
  • Multi-Language Support: Configure via lang parameter with automatic model resolution for 80+ languages.

Frequently Asked Questions

What hardware accelerators does PaddleOCR support?

PaddleOCR automatically detects and utilizes available hardware including NVIDIA GPUs, Intel CPUs with MKLDNN optimization, Baidu XPU, and Huawei NPU. Specify the target device explicitly using the device argument (e.g., device="gpu:0") or allow auto-detection based on your PaddlePaddle installation.

How do I disable document preprocessing steps?

Set use_doc_orientation_classify=False to skip document-level rotation detection, use_doc_unwarping=False to disable geometric distortion correction, and use_textline_orientation=False to turn off line-level orientation classification. These parameters reduce latency when processing pre-aligned images.

Can I use custom-trained models instead of the defaults?

Yes. While the standard API downloads official PP-OCR models, you can pass a custom pipeline configuration YAML to the underlying PaddleX pipeline. Use the utilities in paddleocr/_pipelines/utils.py to generate valid config structures from custom training outputs.

What is the difference between PP-OCRv4 and PP-OCRv5?

PP-OCRv5 represents the latest model generation with improved accuracy for multilingual text and curved line detection. Specify ocr_version="PP-OCRv5" during initialization to access newer architectures, or use "PP-OCRv4" for legacy compatibility with existing deployments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →