# How to Install PaddleOCR from Source: Complete Development Setup Guide

> Install PaddleOCR from source using the PaddlePaddle/PaddleOCR repo. Clone, install dependencies, and run pip install -e . for a development setup to track changes.

- Repository: [PaddlePaddle/PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)
- Tags: how-to-guide
- Published: 2026-03-03

---

**Install PaddleOCR from source by cloning the PaddlePaddle/PaddleOCR repository, installing dependencies from [`requirements.txt`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/requirements.txt), and running `pip install -e .` for an editable installation that tracks your code changes.**

Installing PaddleOCR from the source tree gives you access to the latest development branch, unreleased model definitions, and the ability to modify the library without waiting for PyPI updates. This guide walks through the complete workflow using the official repository structure and packaging configuration.

## Prerequisites

Before installing from source, ensure you have **Python 3.8+** and **Git** installed on your system. You should also have the **PaddlePaddle** deep learning framework installed (CPU or GPU version), as PaddleOCR depends on it for model inference and training.

## Step-by-Step Source Installation

### Clone the Repository

Start by cloning the official repository and optionally checking out a stable release branch:

```bash
git clone https://github.com/PaddlePaddle/PaddleOCR
cd PaddleOCR

# Optional: checkout a specific release branch (e.g., release/3.2)

git checkout release/3.2

```

The project structure follows a modular layout with core pipelines in `paddleocr/_pipelines/`, model definitions in `ppocr/`, and document analysis tools in `ppstructure/`.

### Install Core Dependencies

Install the minimal runtime packages required for OCR inference:

```bash
python -m pip install -r requirements.txt

```

The [`requirements.txt`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/requirements.txt) file at the repository root lists essential packages including `opencv-python`, `numpy`, `shapely`, and `Pillow`. These support the image processing and geometric operations used in `ppocr/utils/` and the detection pipelines.

### Install Optional Feature Groups

PaddleOCR provides **optional dependency groups** for specialized workflows. Install only the components you need:

```bash

# Install all optional features (document parsing, information extraction, translation)

python -m pip install "paddleocr[all]"

# Install only document parsing capabilities

python -m pip install "paddleocr[doc-parser]"

# Install only information extraction models

python -m pip install "paddleocr[ie]"

```

These groups are defined in the packaging metadata and pull additional libraries for layout analysis, table recognition, and visual-language models hosted in `ppstructure/` and [`paddleocr/_pipelines/paddleocr_vl.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/paddleocr_vl.py).

### Install PaddleOCR in Editable Mode

From the repository root, install the package in **editable mode** (`-e` flag):

```bash
python -m pip install -e .

```

This creates a link to your local source tree rather than copying files to your Python environment. Changes made to [`paddleocr/_pipelines/ocr.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/ocr.py) or `ppocr/modeling/` are reflected immediately without reinstallation. The installation uses [`setup.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/setup.py) at the repository root, which configures the package entry points and module discovery.

### Verify the Installation

Confirm the installation works by importing the `PaddleOCR` class and running inference:

```python
from paddleocr import PaddleOCR

ocr = PaddleOCR(use_angle_cls=False, lang='en')
result = ocr.predict('https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/general_ocr_002.png')
print(result[0].text)

```

If the script executes without `ModuleNotFoundError` and returns recognized text, your source installation is functional.

## Understanding the Repository Structure

After installing from source, you can navigate these key directories to customize functionality:

- **[`paddleocr/_pipelines/ocr.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/ocr.py)** – Contains the `PaddleOCR` class implementing the high-level OCR pipeline with text detection and recognition.
- **[`paddleocr/_pipelines/paddleocr_vl.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/paddleocr_vl.py)** – Implements the Vision-Language pipeline (`PaddleOCRVL`) for multimodal document understanding.
- **`ppocr/modeling/`** – Houses backbone architectures, detection heads, and recognition heads defined in `backbones/`, `heads/`, and `necks/`.
- **`ppstructure/`** – Document structure analysis modules including layout extraction, table reconstruction, and key information extraction (KIE).
- **`tools/`** – Training, evaluation, and model export scripts ([`train.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/train.py), [`export_model.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/export_model.py)).
- **`ppocr/utils/`** – Shared utilities for visualization, logging, and data preprocessing.

## Working with the Source Code

### Complete Installation Script

Automate the entire workflow with this bash script:

```bash
#!/usr/bin/env bash
set -e

# Clone and enter repository

git clone https://github.com/PaddlePaddle/PaddleOCR
cd PaddleOCR

# Checkout stable branch (optional)

git checkout release/3.2

# Install dependencies

python -m pip install -r requirements.txt
python -m pip install "paddleocr[all]"

# Editable install

python -m pip install -e .

# Verification test

python - <<'PY'
from paddleocr import PaddleOCR
ocr = PaddleOCR(use_angle_cls=False, lang='en')
res = ocr.predict('https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/general_ocr_002.png')
print(f"Installation successful. Sample text: {res[0].text}")
PY

```

### Customizing the Pipeline

With a source installation, you can modify pipeline behavior directly. To override default preprocessing in [`paddleocr/_pipelines/ocr.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/ocr.py):

```python
from paddleocr import PaddleOCR
from paddleocr._pipelines.ocr import create_config_from_structure

# Load default configuration

config = create_config_from_structure(PaddleOCR)

# Inject custom preprocessing

def resize_to_1024(img):
    import cv2
    return cv2.resize(img, (1024, 1024))

config["preprocess"] = resize_to_1024

# Initialize with modified config

ocr = PaddleOCR.from_config(config)
result = ocr.predict('custom_image.png')

```

## Summary

- **Clone** the PaddlePaddle/PaddleOCR repository and checkout a release branch for stability.
- **Install core dependencies** using [`requirements.txt`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/requirements.txt) to get `opencv-python`, `numpy`, and other inference requirements.
- **Add optional features** via dependency groups like `[all]`, `[doc-parser]`, or `[ie]` for specialized document processing.
- **Use editable mode** (`pip install -e .`) to create a development installation that reflects code changes immediately.
- **Verify** the installation by importing `PaddleOCR` from [`paddleocr/_pipelines/ocr.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/ocr.py) and running test inference.
- **Modify source files** in `ppocr/` or `ppstructure/` to customize models, preprocessing, or post-processing logic.

## Frequently Asked Questions

### What is the difference between installing from PyPI versus source?

Installing from PyPI (`pip install paddleocr`) provides the latest stable release with fixed functionality. Installing from source gives you the **development branch**, access to bleeding-edge features in `ppstructure/`, and the ability to modify [`paddleocr/_pipelines/ocr.py`](https://github.com/PaddlePaddle/PaddleOCR/blob/main/paddleocr/_pipelines/ocr.py) or model definitions in `ppocr/modeling/` without waiting for official releases.

### How do I switch between release branches after cloning?

Navigate to your local repository and use `git checkout` to switch branches:

```bash
cd PaddleOCR
git fetch origin
git checkout release/3.1  # or any specific tag/branch

pip install -e . --force-reinstall --no-deps

```

The `--force-reinstall` flag ensures Python picks up the new branch's code changes.

### How do I uninstall a source-based PaddleOCR installation?

Because source installations use `pip install -e .`, you can remove the package using pip while leaving the repository intact:

```bash
pip uninstall paddleocr

```

This removes the Python environment link but preserves your local `PaddleOCR/` directory for future reinstallation or reference.

### Can I contribute changes back to the repository after installing from source?

Yes. After installing from source with the editable flag, create a new branch for your changes, commit your modifications (e.g., to `ppocr/utils/` or `ppstructure/`), and push to your fork. The editable installation ensures your changes are live for testing before you submit a pull request to the main PaddlePaddle/PaddleOCR repository.