How to Configure PaddleOCR Models: YAML Files, CLI Overrides, and Runtime Options

PaddleOCR employs a layered configuration system that merges static YAML definitions with dynamic command-line overrides to control training and inference pipelines.

PaddleOCR, maintained by PaddlePaddle, separates global training settings from model-specific architecture details through a hierarchical configuration framework. To configure PaddleOCR models for custom datasets or deployment scenarios, you interact with three distinct layers: base YAML files, CLI arguments, and runtime inference flags.

The Three-Layer Configuration Architecture

PaddleOCR's configuration stack processes settings through distinct phases, allowing flexible customization without modifying source code.

Base YAML Configuration Files

The foundation of every PaddleOCR workflow is a YAML file loaded by tools/program.py. The load_config() function asserts that files carry a .yml or .yaml extension and parses them using PyYAML according to the implementation in tools/program.py.

These files organize parameters into logical sections:

Section Purpose Example Keys
Global Training/inference flags, device settings, paths use_gpu, epoch_num, save_model_dir
Architecture Model type, algorithm, backbone, neck, head model_type, algorithm, Backbone.name
Loss Loss function selection and hyperparameters name, balance_loss
Optimizer Optimizer choice and learning rate schedules name, lr.learning_rate
PostProcess Algorithm-specific post-processing name, thresh (for DBPostProcess)
Metric Evaluation metrics name, main_indicator
Train / Eval Dataset definitions and data transforms dataset, transforms, num_workers

A typical detection configuration like configs/det/det_r50_vd_db.yml demonstrates this structure, defining a DB detector with ResNet backbone and associated training parameters.

Command-Line Override Mechanism

Users can override any YAML value without editing files using the -o key=value syntax. In tools/program.py, the ArgsParser._parse_opt() method parses these strings by passing them through yaml.load to convert values into appropriate Python objects (integers, floats, booleans, or strings).

The merge_config() function then recursively merges these CLI-provided values into the loaded configuration dictionary, as implemented at tools/program.py#L88. This supports hierarchical key syntax using dot notation:

python tools/train.py -c configs/det/det_r50_vd_db.yml \
    -o Global.epoch_num=800 \
    -o Optimizer.lr.learning_rate=0.0005 \
    -o Architecture.Backbone.layers=50

Runtime Inference Configuration

During inference, configuration loading follows a different path. Scripts like tools/infer/predict_det.py load a model-specific inference.yml from the exported model directory using utility.load_config() (predict_det.py#L38).

Additionally, tools/infer/utility.py defines shared runtime flags parsed by ArgsParser, including:

  • --use_gpu – Enable GPU inference (default True)
  • --precision – Select fp16, int8, or fp32 for TensorRT backends
  • --use_onnx – Load ONNX format models instead of Paddle
  • --gpu_mem and --gpu_id – Control GPU memory allocation and device selection
  • Algorithm-specific flags (e.g., --det_db_thresh, --det_box_type) – Override post-processing thresholds directly

The create_predictor() function in utility.py merges these runtime flags with the loaded inference.yml configuration before initializing the Paddle Inference engine.

Practical Configuration Examples

Training with Custom Hyperparameters

Override training epochs and learning rate without modifying the base config:

python tools/train.py \
    -c configs/det/det_r50_vd_db.yml \
    -o Global.use_gpu=true \
    -o Global.epoch_num=600 \
    -o Optimizer.lr.learning_rate=0.0008

Exporting Models for Deployment

Export a trained checkpoint and generate the inference.yml required for inference scripts:

python tools/export_model.py \
    -c configs/det/det_r50_vd_db.yml \
    -o Global.save_inference_dir=./inference/det_db \
    -o Global.save_model_dir=./output/det_db

Inference with Runtime Threshold Overrides

Run detection using an exported model while adjusting the DB confidence threshold via CLI:

python tools/infer/predict_det.py \
    --det_model_dir ./inference/det_db \
    --det_algorithm DB \
    --det_db_thresh 0.5

Dynamic Architecture Modifications

Switch backbones dynamically during training initialization:

python tools/train.py \
    -c configs/det/det_mv3_db.yml \
    -o Architecture.Backbone.name=MobileNetV3 \
    -o Architecture.Backbone.layers=1

Key Source Files and Functions

Understanding these core files enables advanced configuration debugging:

  • tools/program.py – Contains load_config() for YAML parsing and merge_config() for recursive dictionary merging of CLI overrides.
  • tools/train.py – Main training entry point; consumes the fully merged configuration dictionary to initialize the data loader, model builder, and optimizer.
  • tools/export_model.py – Exports trained checkpoints and writes the inference.yml file used by prediction scripts.
  • tools/infer/utility.py – Implements load_config() for inference configs and create_predictor() to merge runtime flags with model configurations.
  • tools/infer/predict_det.py – Demonstrates runtime configuration loading for detection tasks, including model-specific parameter application.
  • configs/ – Directory containing default configurations (e.g., configs/det/det_r50_vd_db.yml) serving as templates for custom setups.

Summary

  • PaddleOCR configurations are defined in YAML files organized into Global, Architecture, Loss, Optimizer, PostProcess, Metric, and data sections.
  • Command-line overrides use the -o key=value syntax processed by ArgsParser._parse_opt() and merged via merge_config() in tools/program.py.
  • Hierarchical keys (e.g., Optimizer.lr.learning_rate) allow precise parameter targeting without file editing.
  • Inference pipelines load inference.yml from exported model directories, then merge runtime flags like --use_gpu and --det_db_thresh through utility.create_predictor().
  • Source files in tools/program.py and tools/infer/utility.py implement the core configuration parsing and merging logic.

Frequently Asked Questions

How do I change the learning rate without editing the YAML file?

Use the -o flag with dot notation for nested keys: -o Optimizer.lr.learning_rate=0.0001. The ArgsParser class in tools/program.py parses this string using yaml.load and merge_config() updates the configuration dictionary before training begins.

What is the difference between training configs and inference configs?

Training configurations (e.g., configs/det/det_r50_vd_db.yml) contain full specifications for model architecture, optimization, and data pipelines. When you run tools/export_model.py, PaddleOCR generates an inference.yml containing only the parameters necessary for prediction. During inference, tools/infer/utility.py loads this file and merges it with runtime command-line flags.

Can I switch model architectures using command-line arguments?

Yes. Override the Architecture section using hierarchical syntax. For example, to change a backbone from ResNet to MobileNetV3: -o Architecture.Backbone.name=MobileNetV3 -o Architecture.Backbone.layers=1. This modifies the configuration dictionary before the model builder instantiates the network in tools/train.py.

Where does PaddleOCR load configuration during inference?

Inference scripts like tools/infer/predict_det.py call utility.load_config() to read inference.yml from the directory specified by --det_model_dir. The create_predictor() function in tools/infer/utility.py then merges these YAML settings with runtime flags (e.g., --use_gpu, --precision) to configure the Paddle Inference predictor.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →