How to Fine-Tune PaddleOCR on a Custom Dataset: A Complete Training Guide
Fine-tuning PaddleOCR on a specific dataset requires preparing your data in the image_path\tlabel format, editing a YAML configuration file to point to a pre-trained checkpoint via Global.pretrained_model, and executing tools/train.py to resume training on your custom images.
The PaddlePaddle/PaddleOCR repository provides a modular optical character recognition framework that separates text detection and recognition into distinct, trainable components. To adapt these pre-trained models to specialized domains—such as handwritten documents, low-quality scans, or industry-specific fonts—you must fine-tune PaddleOCR using its configuration-driven pipeline. This guide details the exact workflow to prepare your data, modify training configurations, and deploy production-ready models.
PaddleOCR Architecture and Training Pipeline
PaddleOCR implements a modular pipeline that isolates text detection, text recognition, and post-processing into separate deep-learning models. Each module is defined by a YAML configuration file and trained using the generic script tools/train.py located in the repository root.
According to the PaddleOCR source code, tools/train.py orchestrates the entire training lifecycle: it loads the configuration, initializes the data loader via build_dataloader, constructs the model architecture via create_model, optionally loads pre-trained weights from Global.pretrained_model, and executes the forward-backward loop across epochs. This unified entry point handles both detection models (such as DBNet and EAST) and recognition models (such as SVTR and CRNN).
- Text Detection: Localizes text boxes within images. Configuration example:
configs/det/PP-OCRv3/PP-OCRv3_mobile_det.yml. - Text Recognition: Converts cropped text regions into character sequences. Configuration example:
configs/rec/PP-OCRv3/PP-OCRv3_mobile_rec_distillation.yml.
Dataset Preparation Format
Before fine-tuning, you must structure your dataset according to PaddleOCR’s strict label file specification. The framework expects a text file where each line follows the format:
image_path\tlabel
The tab character (\t) separates the relative image path from its annotation.
- For detection tasks, the label field contains bounding-box coordinates and the transcription (if available).
- For recognition tasks, the label field contains only the text transcription.
Split your data into train and val directories, then specify the label file paths in the Train.dataset.label_file_list field of your YAML configuration.
Step-by-Step Fine-Tuning Workflow
1. Download Pre-Trained Weights
Begin by downloading an official pre-trained checkpoint that matches your target architecture. These checkpoints provide the initialization weights that dramatically accelerate convergence on small custom datasets.
- Detection: Download
ch_PP-OCRv3_det_distill_train.tarand extractstudent.pdparams. - Recognition: Download
ch_PP-OCRv3_rec_train.tarand extractbest_accuracy.pdparams.
2. Configure the Training YAML
Copy the official configuration file to a custom location, then edit three critical sections:
First, set Global.pretrained_model to the absolute or relative path of your downloaded .pdparams file:
Global:
pretrained_model: ./ch_PP-OCRv3_det_distill_train/student.pdparams
save_model_dir: ./output/det_finetune/
Second, update Train.dataset.label_file_list and Eval.dataset.label_file_list to point to your custom label files:
Train:
dataset:
label_file_list: ["./my_data/train_list.txt"]
data_dir: ./my_data/train_imgs/
Third, ensure the Architecture section matches the checkpoint type (detection vs. recognition) to avoid shape mismatch errors during weight loading.
3. Adjust Hyperparameters for Your Hardware
Scale your training parameters linearly with your available GPU count to prevent out-of-memory errors or unstable convergence.
- Batch size: Set
Train.loader.batch_size_per_cardto a value your GPU memory can accommodate (e.g.,8for a single 16GB GPU). - Learning rate: For single-GPU fine-tuning, use approximately
1e-4for detection and5e-5to1e-4for recognition. Multi-GPU training requires proportionally higher rates. - Gradient accumulation: If you must reduce batch size due to memory constraints, enable
Global.accumulate_stepsto maintain effective batch size.
4. Execute Training with tools/train.py
Launch the training process by passing your edited configuration to the training script:
python tools/train.py -c configs/det/PP-OCRv3/PP-OCRv3_mobile_det.yml
For recognition tasks, substitute the configuration path with your recognition YAML. The script automatically loads the pre-trained weights specified in Global.pretrained_model, initializes the optimizer, and begins iterating over your custom dataset. Checkpoints and logs are saved to the directory specified in Global.save_model_dir (default: output/).
5. Export the Fine-Tuned Model for Inference
After training converges, convert the dynamic training checkpoint into a static inference model using tools/export_model.py:
python tools/export_model.py \
-c configs/det/PP-OCRv3/PP-OCRv3_mobile_det.yml \
-o Global.pretrained_model=output/det_finetune/best_accuracy.pdparams \
-o Global.save_inference_dir=./inference/det_finetuned
This generates infer.pdmodel and infer.pdiparams files optimized for deployment. Run inference using the exported model:
python tools/infer/predict_det.py \
-c configs/det/PP-OCRv3/PP-OCRv3_mobile_det.yml \
-o Global.infer_img=./test_images/sample.jpg \
-o Global.pretrained_model=./inference/det_finetuned/infer.pdmodel
Common Pitfalls and Troubleshooting
When you fine-tune PaddleOCR, avoid these frequent configuration errors:
No Images in train dataseterror: Verify thatlabel_file_listpaths are correct relative to your execution directory and that the label file lines strictly follow theimage_path\tlabelformat without extra spaces.- Architecture mismatch errors: Ensure the checkpoint type matches the config’s
Architecturesection. Loading a detection checkpoint into a recognition model causes dimension mismatch failures duringcreate_model. - Out-of-memory (OOM) errors: Reduce
batch_size_per_cardimmediately if you encounter CUDA memory errors. Alternatively, increaseGlobal.accumulate_stepsto simulate larger batches across multiple forward passes. - Unstable training or NaN losses: If using a single GPU, ensure you scaled down the learning rate from the multi-GPU default (e.g., reduce by a factor of 8 if the default config assumes 8 GPUs).
Summary
- PaddleOCR uses a modular architecture where detection and recognition models are trained separately via
tools/train.pyusing YAML configurations. - Dataset labels must follow the strict
image_path\tlabelformat, with bounding-box coordinates included for detection tasks. - Fine-tuning requires setting
Global.pretrained_modelin your config to point to an official.pdparamscheckpoint, then adjustingbatch_size_per_cardand learning rate for your hardware. - Export trained models using
tools/export_model.pyto generate deployment-ready inference files. - Always match the checkpoint architecture (detection vs. recognition) with the configuration file to prevent loading errors.
Frequently Asked Questions
Do I need to fine-tune both the detection and recognition models?
No. You can fine-tune only the component that underperforms on your data. If text boxes are detected correctly but characters are misread, fine-tune only the recognition model using configs/rec/PP-OCRv3/PP-OCRv3_mobile_rec_distillation.yml. Conversely, if detection fails on your specific document layout, fine-tune only the detection model.
What learning rate should I use for single-GPU fine-tuning?
For single-GPU fine-tuning, reduce the default learning rate proportionally. Use approximately 1e-4 for detection tasks and 5e-5 to 1e-4 for recognition tasks. The default configs often assume 8-GPU training with higher base rates, so linear scaling is essential to prevent divergence.
How do I format bounding box coordinates for detection fine-tuning?
In the label file for detection, each line should contain: image_path\t[{"transcription": "text", "points": [[x1,y1], [x2,y2], [x3,y3], [x4,y4]]}]. The points array represents the four corners of the text box in clockwise order. For recognition-only fine-tuning, omit the JSON structure and provide only the transcription string after the tab character.
Can I resume training if it is interrupted?
Yes. Set Global.checkpoints in your YAML file to the path of your last saved checkpoint (e.g., output/det_finetune/latest.pdparams) instead of Global.pretrained_model. The training script will restore both model weights and optimizer state, allowing you to resume from the exact iteration where training stopped.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →