How to Train a Custom PaddleOCR Model: Complete Fine-Tuning Guide
Train a custom PaddleOCR model by preparing your dataset in the required format, copying and editing a task-specific YAML configuration file, and running tools/train.py with the -c flag to point to your config and -o flags to override dataset paths and hyperparameters.
PaddleOCR implements a modular, configuration-driven training architecture that cleanly separates data preparation, model construction, loss calculation, and optimization. The repository allows you to fine-tune text detection, recognition, classification, or key information extraction models on custom datasets without modifying Python source code, leveraging YAML configuration files and command-line overrides.
Understanding the Training Architecture
The training system orchestrates several specialized modules through a central entry point. When you launch training, tools/train.py initializes the environment and delegates execution to specialized builders and the core training loop.
Key components include:
tools/train.py— The main entry point that parses CLI arguments, loads the YAML configuration, initializes distributed training if enabled (paddle.distributed.init_parallel_env), and builds the data pipeline.ppocr/data/build_dataloader.py— Constructs train and validation dataloaders, applying augmentations defined in the config.ppocr/modeling/architectures.py— Factory that assembles the model architecture (detection, recognition, or classification) based on theArchitecturesection of your YAML file.ppocr/optimizer/learning_rate.py— Implements learning-rate schedules (Piecewise, Cosine, etc.) used by the optimizer.tools/program.py— Contains theprogram.trainfunction that executes the actual training loop, handles checkpointing, evaluation, logging, and VisualDL integration.
All hyperparameters—including batch size, learning-rate schedules, data augmentation policies, and dataset paths—are defined in a YAML configuration file (e.g., configs/rec/rec_mv3_none_none_ctc.yml) and can be overridden at runtime using the -o command-line option.
Preparing Your Dataset
PaddleOCR requires strict dataset formatting that varies by task. Store images in a dedicated directory and create a label file where each line maps an image to its annotation.
For Text Recognition:
- Format:
<image_path>\t<label> - Example line:
train/img_001.jpg\tHelloWorld - Place all training images in a single directory (e.g.,
/data/my_dataset/rec/images/).
For Text Detection:
- Format:
image_path\t[{"transcription": "text", "points": [[x1,y1], [x2,y2], ...]}, ...] - Polygons must be defined as coordinate lists.
Configuring the Training Pipeline
Step 1: Copy a base configuration that matches your task from the configs/ directory. Recognition configs reside in configs/rec/, detection in configs/det/, and table/layout in configs/table/ or configs/structure/.
Step 2: Edit the Train and Eval sections to point to your data:
Train:
dataset:
image_dir: /data/my_dataset/rec/images
label_file_list:
- /data/my_dataset/rec/labels.txt
loader:
batch_size_per_card: 64
Eval:
dataset:
image_dir: /data/my_dataset/rec/val_images
label_file_list:
- /data/my_dataset/rec/val_labels.txt
Step 3: Set Global.pretrained_model to a PaddleOCR checkpoint path to enable transfer learning, and specify Global.save_model_dir for output checkpoints.
Step 4: Adjust the PostProcess configuration (e.g., CTCLabelDecode or SARLabelDecode) to match your character dictionary; the post-process class automatically adjusts the model head's output channels based on the character set size.
Running the Training Script
Single-GPU Training
Execute tools/train.py with the config path and any necessary overrides:
python tools/train.py \
-c configs/rec/rec_mv3_none_none_ctc.yml \
-o Global.pretrained_model=./pretrain_models/MobileNetV3_large_x0_5_pretrained \
Global.save_model_dir=./my_custom_rec \
Global.epoch_num=50 \
Global.use_amp=True \
Train.loader.batch_size_per_card=64 \
Train.dataset.label_file_list=["/data/my_dataset/rec/labels.txt"] \
Train.dataset.image_dir="/data/my_dataset/rec/images"
The -o flag accepts dot-notation keys to override any YAML value without editing the file.
Multi-GPU Distributed Training
For multi-GPU training, use Paddle's distributed launcher. The script automatically detects the distributed environment and initializes init_parallel_env:
python -m paddle.distributed.launch --gpus=0,1,2 \
tools/train.py \
-c configs/det/det_mv3_db.yml \
-o Global.pretrained_model=./pretrain_models/MobileNetV3_large_x0_5_pretrained \
Global.save_model_dir=./my_det_finetune \
Global.epoch_num=40 \
Train.loader.batch_size_per_card=16 \
Train.dataset.label_file_list=["/data/custom_det/label.txt"] \
Train.dataset.image_dir="/data/custom_det/images"
Enabling Advanced Features
- Automatic Mixed Precision (AMP): Set
Global.use_amp: Truein the config or pass-o Global.use_amp=Trueto enable FP16 training with automatic loss scaling. - Static Graph Optimization: The training script automatically applies
apply_to_staticwhen configured, optimizing the graph for deployment. - Pretrained Backbone Loading: The function
ppocr.utils.save_load.load_modelloads weights into the backbone before training begins, allowing fine-tuning from general-purpose vision features.
Exporting the Inference Model
After training completes, convert the saved checkpoints to a static inference model using tools/export_model.py:
python tools/export_model.py \
-c configs/rec/rec_mv3_none_none_ctc.yml \
-o Global.checkpoints=./my_custom_rec/latest \
Global.save_inference_dir=./inference/custom_rec
This generates .pdmodel, .pdiparams, and .pdiparams.info files optimized for the Paddle Inference engine.
Summary
- Configuration-driven workflow: All training parameters live in YAML files (e.g.,
configs/rec/rec_mv3_none_none_ctc.yml), editable via text editors or CLI-ooverrides. - Modular code structure:
tools/train.pyorchestratesbuild_dataloader.pyfor data,architectures.pyfor models, andprogram.pyfor the training loop. - Flexible compute: The same training script supports single-GPU, multi-GPU distributed, and mixed-precision training without code changes.
- Dataset requirements: Recognition uses
\tdelimited label files; detection uses JSON polygon annotations. - Transfer learning: Load pretrained weights via
Global.pretrained_modelto accelerate convergence on small custom datasets.
Frequently Asked Questions
What file format does PaddleOCR expect for custom datasets?
PaddleOCR uses tab-separated text files. For recognition, each line contains image_path\tlabel. For detection, each line contains image_path\t[{\"transcription\": \"text\", \"points\": [[x1,y1],...]}]. The dataset builder in ppocr/data/build_dataloader.py parses these formats according to the dataset class specified in your YAML (e.g., SimpleDataSet or LMDBDataSet).
Can I train on multiple GPUs without changing the training script?
Yes. Use the paddle.distributed.launch module as shown in the multi-GPU example. The training script in tools/train.py automatically detects Global.distributed: True and initializes the parallel environment via paddle.distributed.init_parallel_env(), partitioning the batch across GPUs.
How do I resume training from a checkpoint?
Set Global.checkpoints to the path of your saved model directory (containing best_accuracy.pdparams or latest.pdparams) using the -o flag:
python tools/train.py -c configs/rec/rec_mv3_none_none_ctc.yml -o Global.checkpoints=./output/latest
The load_model utility in ppocr/utils/save_load.py restores both model weights and optimizer states.
Why should I configure the PostProcess section when training recognition?
The PostProcess class (e.g., CTCLabelDecode) defines the character dictionary used to map model outputs to text strings. During initialization in tools/train.py, the model head's output channels are automatically adjusted to match the number of characters in this dictionary. Mismatches here cause training failures or incorrect inference results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →