How to Configure PaddleOCR Models: YAML Files, CLI Overrides, and Runtime Options
PaddleOCR employs a layered configuration system that merges static YAML definitions with dynamic command-line overrides to control training and inference pipelines.
PaddleOCR, maintained by PaddlePaddle, separates global training settings from model-specific architecture details through a hierarchical configuration framework. To configure PaddleOCR models for custom datasets or deployment scenarios, you interact with three distinct layers: base YAML files, CLI arguments, and runtime inference flags.
The Three-Layer Configuration Architecture
PaddleOCR's configuration stack processes settings through distinct phases, allowing flexible customization without modifying source code.
Base YAML Configuration Files
The foundation of every PaddleOCR workflow is a YAML file loaded by tools/program.py. The load_config() function asserts that files carry a .yml or .yaml extension and parses them using PyYAML according to the implementation in tools/program.py.
These files organize parameters into logical sections:
| Section | Purpose | Example Keys |
|---|---|---|
Global |
Training/inference flags, device settings, paths | use_gpu, epoch_num, save_model_dir |
Architecture |
Model type, algorithm, backbone, neck, head | model_type, algorithm, Backbone.name |
Loss |
Loss function selection and hyperparameters | name, balance_loss |
Optimizer |
Optimizer choice and learning rate schedules | name, lr.learning_rate |
PostProcess |
Algorithm-specific post-processing | name, thresh (for DBPostProcess) |
Metric |
Evaluation metrics | name, main_indicator |
Train / Eval |
Dataset definitions and data transforms | dataset, transforms, num_workers |
A typical detection configuration like configs/det/det_r50_vd_db.yml demonstrates this structure, defining a DB detector with ResNet backbone and associated training parameters.
Command-Line Override Mechanism
Users can override any YAML value without editing files using the -o key=value syntax. In tools/program.py, the ArgsParser._parse_opt() method parses these strings by passing them through yaml.load to convert values into appropriate Python objects (integers, floats, booleans, or strings).
The merge_config() function then recursively merges these CLI-provided values into the loaded configuration dictionary, as implemented at tools/program.py#L88. This supports hierarchical key syntax using dot notation:
python tools/train.py -c configs/det/det_r50_vd_db.yml \
-o Global.epoch_num=800 \
-o Optimizer.lr.learning_rate=0.0005 \
-o Architecture.Backbone.layers=50
Runtime Inference Configuration
During inference, configuration loading follows a different path. Scripts like tools/infer/predict_det.py load a model-specific inference.yml from the exported model directory using utility.load_config() (predict_det.py#L38).
Additionally, tools/infer/utility.py defines shared runtime flags parsed by ArgsParser, including:
--use_gpu– Enable GPU inference (defaultTrue)--precision– Selectfp16,int8, orfp32for TensorRT backends--use_onnx– Load ONNX format models instead of Paddle--gpu_memand--gpu_id– Control GPU memory allocation and device selection- Algorithm-specific flags (e.g.,
--det_db_thresh,--det_box_type) – Override post-processing thresholds directly
The create_predictor() function in utility.py merges these runtime flags with the loaded inference.yml configuration before initializing the Paddle Inference engine.
Practical Configuration Examples
Training with Custom Hyperparameters
Override training epochs and learning rate without modifying the base config:
python tools/train.py \
-c configs/det/det_r50_vd_db.yml \
-o Global.use_gpu=true \
-o Global.epoch_num=600 \
-o Optimizer.lr.learning_rate=0.0008
Exporting Models for Deployment
Export a trained checkpoint and generate the inference.yml required for inference scripts:
python tools/export_model.py \
-c configs/det/det_r50_vd_db.yml \
-o Global.save_inference_dir=./inference/det_db \
-o Global.save_model_dir=./output/det_db
Inference with Runtime Threshold Overrides
Run detection using an exported model while adjusting the DB confidence threshold via CLI:
python tools/infer/predict_det.py \
--det_model_dir ./inference/det_db \
--det_algorithm DB \
--det_db_thresh 0.5
Dynamic Architecture Modifications
Switch backbones dynamically during training initialization:
python tools/train.py \
-c configs/det/det_mv3_db.yml \
-o Architecture.Backbone.name=MobileNetV3 \
-o Architecture.Backbone.layers=1
Key Source Files and Functions
Understanding these core files enables advanced configuration debugging:
tools/program.py– Containsload_config()for YAML parsing andmerge_config()for recursive dictionary merging of CLI overrides.tools/train.py– Main training entry point; consumes the fully merged configuration dictionary to initialize the data loader, model builder, and optimizer.tools/export_model.py– Exports trained checkpoints and writes theinference.ymlfile used by prediction scripts.tools/infer/utility.py– Implementsload_config()for inference configs andcreate_predictor()to merge runtime flags with model configurations.tools/infer/predict_det.py– Demonstrates runtime configuration loading for detection tasks, including model-specific parameter application.configs/– Directory containing default configurations (e.g.,configs/det/det_r50_vd_db.yml) serving as templates for custom setups.
Summary
- PaddleOCR configurations are defined in YAML files organized into
Global,Architecture,Loss,Optimizer,PostProcess,Metric, and data sections. - Command-line overrides use the
-o key=valuesyntax processed byArgsParser._parse_opt()and merged viamerge_config()intools/program.py. - Hierarchical keys (e.g.,
Optimizer.lr.learning_rate) allow precise parameter targeting without file editing. - Inference pipelines load
inference.ymlfrom exported model directories, then merge runtime flags like--use_gpuand--det_db_threshthroughutility.create_predictor(). - Source files in
tools/program.pyandtools/infer/utility.pyimplement the core configuration parsing and merging logic.
Frequently Asked Questions
How do I change the learning rate without editing the YAML file?
Use the -o flag with dot notation for nested keys: -o Optimizer.lr.learning_rate=0.0001. The ArgsParser class in tools/program.py parses this string using yaml.load and merge_config() updates the configuration dictionary before training begins.
What is the difference between training configs and inference configs?
Training configurations (e.g., configs/det/det_r50_vd_db.yml) contain full specifications for model architecture, optimization, and data pipelines. When you run tools/export_model.py, PaddleOCR generates an inference.yml containing only the parameters necessary for prediction. During inference, tools/infer/utility.py loads this file and merges it with runtime command-line flags.
Can I switch model architectures using command-line arguments?
Yes. Override the Architecture section using hierarchical syntax. For example, to change a backbone from ResNet to MobileNetV3: -o Architecture.Backbone.name=MobileNetV3 -o Architecture.Backbone.layers=1. This modifies the configuration dictionary before the model builder instantiates the network in tools/train.py.
Where does PaddleOCR load configuration during inference?
Inference scripts like tools/infer/predict_det.py call utility.load_config() to read inference.yml from the directory specified by --det_model_dir. The create_predictor() function in tools/infer/utility.py then merges these YAML settings with runtime flags (e.g., --use_gpu, --precision) to configure the Paddle Inference predictor.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →