What Are the Different OCR Pipelines in PaddleOCR? A Complete Technical Guide

PaddleOCR provides ten specialized OCR pipelines—from classic text detection to multimodal visual-language models—all inheriting from the PaddleXPipelineWrapper base class and exposing unified predict() and CLI interfaces.

The PaddlePaddle/PaddleOCR repository organizes its optical character recognition capabilities into a modular hierarchy of pipeline classes. These wrappers combine low-level modules like text detection, layout parsing, and VLMs into single-call APIs, all residing under paddleocr/_pipelines/ and registered for both local and service-based execution.

Core Pipeline Architecture

Every OCR pipeline in the repository extends the abstract base class PaddleXPipelineWrapper defined in paddleocr/_pipelines/base.py. This foundation provides three critical abstractions:

  • A unified predict and predict_iter interface for batch or streaming inference.
  • Automatic conversion of user-friendly arguments into internal PaddleX configurations.
  • A PipelineCLISubcommandExecutor that exposes each pipeline through the paddleocr command-line tool.

This architecture ensures that whether you are running classic OCR or a multimodal document analysis workflow, the API surface remains consistent.

The Ten OCR Pipelines Explained

1. Classic End-to-End OCR: PaddleOCR

The PaddleOCR pipeline in paddleocr/_pipelines/ocr.py implements the standard detection-to-recognition flow. It chains TextDetection and TextRecognition modules, optionally inserting DocOrientationClassify and DocUnwarping for preprocessing skewed or curved documents.

from paddleocr import PaddleOCR

ocr = PaddleOCR(lang="en", ocr_version="PP-OCRv5")
result = ocr.predict("doc/example.png")
print(result[0]["text_lines"])  # List of (text, confidence, bbox) tuples

2. Layout-Aware Document Parsing: PPStructureV3

For complex documents containing tables, formulas, and mixed layouts, the PPStructureV3 class (source: paddleocr/_pipelines/pp_structurev3.py) combines DocPreprocessor, LayoutParsing, and the OCR sub-pipeline to return structured Markdown with embedded base64 images.

from paddleocr import PPStructureV3

doc = PPStructureV3()
markdown_output = doc.predict("doc/example.pdf")

3. Visual-Language Model Integration: PaddleOCRVL

The PaddleOCRVL pipeline (paddleocr/_pipelines/paddleocr_vl.py) bridges a large visual-language model backbone with the OCR sub-pipeline. It requires an optional mllm service deployment and outputs structured dictionaries containing both OCR text and element metadata.

from paddleocr import PaddleOCRVL

vl = PaddleOCRVL(pipeline_version="v1")
structured_content = vl.predict("doc/complex_page.png")
print(structured_content["markdown"])

4. Conversational Document Analysis: PPChatOCRv4Doc

PPChatOCRv4Doc (paddleocr/_pipelines/pp_chatocrv4_doc.py) wraps PPStructureV3 with a chat LLM interface (ERNIE-Bot), enabling question-answering over document content through a conversational API.

5. Automated Document Translation: PPDocTranslation

The PPDocTranslation pipeline (paddleocr/_pipelines/pp_doctranslation.py) chains PPStructureV3 with ERNIE 4.5 translation models to perform end-to-end document translation, converting source PDFs into translated Markdown.

from paddleocr import PPDocTranslation

translator = PPDocTranslation(target_lang="en")
translated_md = translator.predict("doc/chinese_doc.pdf")

6. Table Recognition: TableRecognitionPipelineV2

TableRecognitionPipelineV2 (paddleocr/_pipelines/table_recognition_v2.py) provides dedicated table detection and cell extraction, combining a table detector with OCR for reading individual cell contents.

7. Seal and Stamp Recognition: SealRecognition

Official document processing relies on SealRecognition (paddleocr/_pipelines/seal_recognition.py), which detects seals or stamps and applies OCR to extract their text content.

8. Mathematical Formula Extraction: FormulaRecognitionPipeline

For academic documents, FormulaRecognitionPipeline (paddleocr/_pipelines/formula_recognition.py) detects mathematical expressions and outputs LaTeX-formatted strings.

9. Document Preprocessing: DocPreprocessor

The low-level DocPreprocessor pipeline (paddleocr/_pipelines/doc_preprocessor.py) exposes DocOrientationClassify and DocUnwarping as reusable components for correcting document skew and perspective distortion before OCR.

10. Generic Document Understanding: DocUnderstanding

DocUnderstanding (paddleocr/_pipelines/doc_understanding.py) acts as a flexible wrapper that combines OCR, PPStructureV3, and custom downstream modules into a single document-understanding workflow.

MCP Server and CLI Integration

All pipelines are registered in the Model Control Protocol (MCP) server via mcp_server/paddleocr_mcp/pipelines.py. The _PIPELINE_HANDLERS dictionary maps human-readable names to handler classes that support both local execution (instantiating Python classes) and service execution (remote HTTP endpoints):

_PIPELINE_HANDLERS = {
    "OCR": OCRHandler,
    "PP-StructureV3": PPStructureV3Handler,
    "PaddleOCR-VL": PaddleOCRVLHandler,
    "PaddleOCR-VL-1.5": PaddleOCRVLHandler,
}

Each pipeline also exposes CLI sub-commands through dedicated executor classes (e.g., PaddleOCRCLISubcommandExecutor):

paddleocr ocr --image_dir ./images --lang en --ocr_version PP-OCRv5

Summary

  • Base abstraction: All pipelines inherit from PaddleXPipelineWrapper in paddleocr/_pipelines/base.py, ensuring consistent predict() and CLI interfaces.
  • Classic OCR: The PaddleOCR class handles standard text detection and recognition.
  • Document intelligence: PPStructureV3, PPChatOCRv4Doc, and PPDocTranslation provide layout parsing, conversational AI, and translation capabilities.
  • Multimodal AI: PaddleOCRVL integrates visual-language models for complex document understanding.
  • Specialized tasks: Dedicated pipelines exist for tables (TableRecognitionPipelineV2), seals (SealRecognition), formulas (FormulaRecognitionPipeline), and preprocessing (DocPreprocessor).
  • Deployment flexibility: The MCP server registration in mcp_server/paddleocr_mcp/pipelines.py enables both local and remote service execution modes.

Frequently Asked Questions

What is the difference between the PaddleOCR and PPStructureV3 pipelines?

The PaddleOCR pipeline performs classic line-by-line text detection and recognition, returning raw text and bounding boxes. PPStructureV3 adds layout analysis to identify document elements like headings, tables, and formulas, outputting structured Markdown rather than plain text sequences.

How do I execute pipelines in service mode versus local mode?

The MCP handler classes defined in mcp_server/paddleocr_mcp/pipelines.py abstract execution mode. When instantiated locally, handlers create the Python pipeline class directly (e.g., PaddleOCR()). In service mode, the same handler routes calls to a remote HTTP endpoint, allowing identical code to run against either local GPUs or remote servers.

Which pipeline should I use for PDF documents containing tables and math formulas?

Use PPStructureV3 for mixed-content PDFs. It automatically routes tables to TableRecognitionPipelineV2 and formulas to FormulaRecognitionPipeline while maintaining document structure. For pure translation workflows, PPDocTranslation chains this parsing with ERNIE 4.5 translation models.

Can I preprocess documents before running OCR?

Yes. The DocPreprocessor pipeline (paddleocr/_pipelines/doc_preprocessor.py) provides orientation classification and geometric unwarping. Alternatively, the main PaddleOCR class accepts parameters to enable these preprocessing steps automatically before text detection occurs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →