# What Are the Different OCR Pipelines in PaddleOCR? A Complete Technical Guide

> Explore the ten OCR pipelines in PaddleOCR, from classic detection to multimodal VLM. This guide details their unified interfaces for efficient text recognition.

- Repository: [PaddlePaddle/PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)
- Tags: deep-dive
- Published: 2026-03-03

---

**PaddleOCR provides ten specialized OCR pipelines—from classic text detection to multimodal visual-language models—all inheriting from the `PaddleXPipelineWrapper` base class and exposing unified `predict()` and CLI interfaces.**

The PaddlePaddle/PaddleOCR repository organizes its optical character recognition capabilities into a modular hierarchy of pipeline classes. These wrappers combine low-level modules like text detection, layout parsing, and VLMs into single-call APIs, all residing under `paddleocr/_pipelines/` and registered for both local and service-based execution.

## Core Pipeline Architecture

Every OCR pipeline in the repository extends the abstract base class **`PaddleXPipelineWrapper`** defined in [`paddleocr/_pipelines/base.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/base.py). This foundation provides three critical abstractions:

- A unified **`predict`** and **`predict_iter`** interface for batch or streaming inference.
- Automatic conversion of user-friendly arguments into internal PaddleX configurations.
- A **`PipelineCLISubcommandExecutor`** that exposes each pipeline through the `paddleocr` command-line tool.

This architecture ensures that whether you are running classic OCR or a multimodal document analysis workflow, the API surface remains consistent.

## The Ten OCR Pipelines Explained

### 1. Classic End-to-End OCR: `PaddleOCR`

The **`PaddleOCR`** pipeline in [`paddleocr/_pipelines/ocr.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/ocr.py) implements the standard detection-to-recognition flow. It chains **`TextDetection`** and **`TextRecognition`** modules, optionally inserting **`DocOrientationClassify`** and **`DocUnwarping`** for preprocessing skewed or curved documents.

```python
from paddleocr import PaddleOCR

ocr = PaddleOCR(lang="en", ocr_version="PP-OCRv5")
result = ocr.predict("doc/example.png")
print(result[0]["text_lines"])  # List of (text, confidence, bbox) tuples

```

### 2. Layout-Aware Document Parsing: `PPStructureV3`

For complex documents containing tables, formulas, and mixed layouts, the **`PPStructureV3`** class (source: [`paddleocr/_pipelines/pp_structurev3.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/pp_structurev3.py)) combines **`DocPreprocessor`**, **`LayoutParsing`**, and the OCR sub-pipeline to return structured Markdown with embedded base64 images.

```python
from paddleocr import PPStructureV3

doc = PPStructureV3()
markdown_output = doc.predict("doc/example.pdf")

```

### 3. Visual-Language Model Integration: `PaddleOCRVL`

The **`PaddleOCRVL`** pipeline ([`paddleocr/_pipelines/paddleocr_vl.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/paddleocr_vl.py)) bridges a large visual-language model backbone with the OCR sub-pipeline. It requires an optional **mllm** service deployment and outputs structured dictionaries containing both OCR text and element metadata.

```python
from paddleocr import PaddleOCRVL

vl = PaddleOCRVL(pipeline_version="v1")
structured_content = vl.predict("doc/complex_page.png")
print(structured_content["markdown"])

```

### 4. Conversational Document Analysis: `PPChatOCRv4Doc`

**`PPChatOCRv4Doc`** ([`paddleocr/_pipelines/pp_chatocrv4_doc.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/pp_chatocrv4_doc.py)) wraps `PPStructureV3` with a chat LLM interface (ERNIE-Bot), enabling question-answering over document content through a conversational API.

### 5. Automated Document Translation: `PPDocTranslation`

The **`PPDocTranslation`** pipeline ([`paddleocr/_pipelines/pp_doctranslation.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/pp_doctranslation.py)) chains `PPStructureV3` with ERNIE 4.5 translation models to perform end-to-end document translation, converting source PDFs into translated Markdown.

```python
from paddleocr import PPDocTranslation

translator = PPDocTranslation(target_lang="en")
translated_md = translator.predict("doc/chinese_doc.pdf")

```

### 6. Table Recognition: `TableRecognitionPipelineV2`

**`TableRecognitionPipelineV2`** ([`paddleocr/_pipelines/table_recognition_v2.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/table_recognition_v2.py)) provides dedicated table detection and cell extraction, combining a table detector with OCR for reading individual cell contents.

### 7. Seal and Stamp Recognition: `SealRecognition`

Official document processing relies on **`SealRecognition`** ([`paddleocr/_pipelines/seal_recognition.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/seal_recognition.py)), which detects seals or stamps and applies OCR to extract their text content.

### 8. Mathematical Formula Extraction: `FormulaRecognitionPipeline`

For academic documents, **`FormulaRecognitionPipeline`** ([`paddleocr/_pipelines/formula_recognition.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/formula_recognition.py)) detects mathematical expressions and outputs LaTeX-formatted strings.

### 9. Document Preprocessing: `DocPreprocessor`

The low-level **`DocPreprocessor`** pipeline ([`paddleocr/_pipelines/doc_preprocessor.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/doc_preprocessor.py)) exposes **`DocOrientationClassify`** and **`DocUnwarping`** as reusable components for correcting document skew and perspective distortion before OCR.

### 10. Generic Document Understanding: `DocUnderstanding`

**`DocUnderstanding`** ([`paddleocr/_pipelines/doc_understanding.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/doc_understanding.py)) acts as a flexible wrapper that combines OCR, `PPStructureV3`, and custom downstream modules into a single document-understanding workflow.

## MCP Server and CLI Integration

All pipelines are registered in the Model Control Protocol (MCP) server via [`mcp_server/paddleocr_mcp/pipelines.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/mcp_server/paddleocr_mcp/pipelines.py). The `_PIPELINE_HANDLERS` dictionary maps human-readable names to handler classes that support both **local** execution (instantiating Python classes) and **service** execution (remote HTTP endpoints):

```python
_PIPELINE_HANDLERS = {
    "OCR": OCRHandler,
    "PP-StructureV3": PPStructureV3Handler,
    "PaddleOCR-VL": PaddleOCRVLHandler,
    "PaddleOCR-VL-1.5": PaddleOCRVLHandler,
}

```

Each pipeline also exposes CLI sub-commands through dedicated executor classes (e.g., `PaddleOCRCLISubcommandExecutor`):

```bash
paddleocr ocr --image_dir ./images --lang en --ocr_version PP-OCRv5

```

## Summary

- **Base abstraction**: All pipelines inherit from `PaddleXPipelineWrapper` in [`paddleocr/_pipelines/base.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/base.py), ensuring consistent `predict()` and CLI interfaces.
- **Classic OCR**: The `PaddleOCR` class handles standard text detection and recognition.
- **Document intelligence**: `PPStructureV3`, `PPChatOCRv4Doc`, and `PPDocTranslation` provide layout parsing, conversational AI, and translation capabilities.
- **Multimodal AI**: `PaddleOCRVL` integrates visual-language models for complex document understanding.
- **Specialized tasks**: Dedicated pipelines exist for tables (`TableRecognitionPipelineV2`), seals (`SealRecognition`), formulas (`FormulaRecognitionPipeline`), and preprocessing (`DocPreprocessor`).
- **Deployment flexibility**: The MCP server registration in [`mcp_server/paddleocr_mcp/pipelines.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/mcp_server/paddleocr_mcp/pipelines.py) enables both local and remote service execution modes.

## Frequently Asked Questions

### What is the difference between the `PaddleOCR` and `PPStructureV3` pipelines?

The **`PaddleOCR`** pipeline performs classic line-by-line text detection and recognition, returning raw text and bounding boxes. **`PPStructureV3`** adds layout analysis to identify document elements like headings, tables, and formulas, outputting structured Markdown rather than plain text sequences.

### How do I execute pipelines in service mode versus local mode?

The MCP handler classes defined in [`mcp_server/paddleocr_mcp/pipelines.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/mcp_server/paddleocr_mcp/pipelines.py) abstract execution mode. When instantiated locally, handlers create the Python pipeline class directly (e.g., `PaddleOCR()`). In service mode, the same handler routes calls to a remote HTTP endpoint, allowing identical code to run against either local GPUs or remote servers.

### Which pipeline should I use for PDF documents containing tables and math formulas?

Use **`PPStructureV3`** for mixed-content PDFs. It automatically routes tables to `TableRecognitionPipelineV2` and formulas to `FormulaRecognitionPipeline` while maintaining document structure. For pure translation workflows, **`PPDocTranslation`** chains this parsing with ERNIE 4.5 translation models.

### Can I preprocess documents before running OCR?

Yes. The **`DocPreprocessor`** pipeline ([`paddleocr/_pipelines/doc_preprocessor.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/doc_preprocessor.py)) provides orientation classification and geometric unwarping. Alternatively, the main `PaddleOCR` class accepts parameters to enable these preprocessing steps automatically before text detection occurs.