How to Evaluate PaddleOCR Model Performance: A Complete Guide
Use tools/eval.py to run standardized evaluation on validation data, or call program.eval() directly in Python to compute task-specific metrics like accuracy for recognition or H-mean for detection.
PaddleOCR provides a unified evaluation framework that handles both single-task models (text detection, recognition, classification) and complex end-to-end document understanding pipelines. The evaluation system is configuration-driven, allowing you to assess model quality using YAML config files without modifying source code, while the underlying implementation in tools/program.py provides a reusable loop for custom integrations.
Overview of the Evaluation Architecture
The PaddleOCR evaluation flow follows a strict pipeline: data loading → inference → post-processing → metric aggregation. This architecture is implemented across several core modules:
tools/eval.py– The command-line entry point that orchestrates the entire process (lines 35-78).tools/program.py– Contains the genericeval()function (lines 61-74) that executes the inference loop.ppocr/metrics/– Task-specific metric classes (RecMetric,DetMetric,E2EMetric) that compute final scores.tools/end2end/eval_end2end.py– Specialized evaluator for full pipelines that calculates polygon IoU and edit distance.
For single-task evaluation, the system uses build_dataloader() to prepare validation data, build_model() to construct the architecture, and build_metric() to instantiate the appropriate scorer based on the config file's Metric section.
Single-Task Model Evaluation
Single-task evaluation applies to text detection, text recognition, and angle classification models. You can run evaluation via CLI for quick assessment or use the Python API for custom workflows.
Command-Line Evaluation
The fastest way to evaluate a trained model is using the provided evaluation script with a configuration file:
# Install PaddleOCR
pip install "paddleocr[all]"
# Run evaluation (example for recognition model)
python tools/eval.py -c configs/rec/rec_ch_PP-OCRv3.yml
Ensure your configuration file contains a valid Eval dataset section pointing to your validation images and annotations. The script automatically loads the checkpoint specified in the config, runs inference on the validation set, and outputs metrics:
metric eval ***************
accuracy:0.9578
norm_edit_dis:0.0412
fps: 123.4
The tools/eval.py script handles device placement, mixed precision (AMP) scaling when enabled, and checkpoint loading via load_model() from ppocr/utils/save_load.py.
Python API Evaluation
For integration into custom pipelines or Jupyter notebooks, import the evaluation components directly:
import yaml
import paddle
from ppocr.data import build_dataloader
from ppocr.modeling.architectures import build_model
from ppocr.postprocess import build_post_process
from ppocr.metrics import build_metric
from ppocr.utils.save_load import load_model
from tools.program import eval as paddleocr_eval
# Load configuration
config = yaml.safe_load(open('configs/rec/rec_ch_PP-OCRv3.yml', 'rb'))
device = paddle.set_device('gpu' if paddle.is_compiled_with_cuda() else 'cpu')
# Build components
val_loader = build_dataloader(config, 'Eval', device, logger=None)
model = build_model(config['Architecture'])
load_model(config, model, model_type=config['Architecture']['model_type'])
post_process = build_post_process(config['PostProcess'], config['Global'])
metric = build_metric(config['Metric'])
# Execute evaluation
results = paddleocr_eval(
model,
val_loader,
post_process,
metric,
model_type=config['Architecture'].get('model_type')
)
print(f"Final accuracy: {results}")
This approach gives you full control over the data loader, model initialization, and metric computation while maintaining compatibility with PaddleOCR's standardized evaluation logic.
Key Evaluation Components
Understanding these factory functions helps debug evaluation issues or implement custom metrics:
build_dataloader(config, 'Eval', device, ...)– Constructs the validation data loader fromppocr/data/__init__.pyusing the dataset paths and transforms defined in the config'sEvalsection.build_model(config['Architecture'])– Instantiates the neural network architecture; for distillation models, this automatically adjusts head output channels.build_post_process(config['PostProcess'], global_config)– Converts raw model outputs (logits or heatmaps) into interpretable predictions like text strings or bounding boxes.build_metric(config['Metric'])– Creates the metric calculator that aggregates batch-wise statistics into final scores such as accuracy for recognition or H-mean for detection.
End-to-End Document Evaluation
For complete OCR pipelines that include layout analysis, text detection, and recognition in one system, use the specialized end-to-end evaluator:
python tools/end2end/eval_end2end.py \
/path/to/ground_truth \
/path/to/predictions \
--ignore_blank=True
This script computes polygon IoU (Intersection over Union) between predicted and ground-truth text regions using the polygon_iou() function, then calculates character-level edit distance via the ed() function. It aggregates these into precision, recall, and F-measure scores:
hit, dt_count, gt_count 1250 1300 1240
character_acc: 93.21%
avg_edit_dist_field: 1.84
precision: 96.15%
recall: 92.31%
fmeasure: 94.19%
The e2e_eval() function (lines 70-80 in tools/end2end/eval_end2end.py) handles the complex matching logic required when multiple text regions overlap or when layout elements must be considered alongside text content.
Understanding Evaluation Metrics
Different tasks use different primary indicators:
- Text Recognition: Uses
RecMetricto calculate accuracy (exact match) and normalized edit distance (character error rate). - Text Detection: Uses
DetMetricto compute precision, recall, and H-mean (F1-score) based on box IoU thresholds. - End-to-End: Uses
E2EMetricor the standaloneeval_end2end.pyto combine detection IoU with transcription accuracy, producing an overall F-measure that reflects both localization and recognition quality.
All metric classes implement a consistent interface: they accumulate batch results via __call__() and return a dictionary of final values through get_metric().
Summary
- Primary entry point: Use
tools/eval.pyfor standard single-task evaluation andtools/end2end/eval_end2end.pyfor full document pipelines. - Configuration-driven: Evaluation requires a YAML config file with a properly defined
Evaldataset section andMetricspecification. - Core workflow:
build_dataloader()→build_model()→load_model()→program.eval()→ metric aggregation. - Key files:
tools/program.pycontains the reusable evaluation loop;ppocr/metrics/contains task-specific scoring logic. - Metrics vary by task: Recognition uses accuracy, detection uses H-mean, and end-to-end uses F-measure combined with edit distance.
Frequently Asked Questions
How do I evaluate a model on a custom dataset?
Prepare a YAML configuration file that specifies your validation image paths and label files in the Eval.dataset section, then run python tools/eval.py -c your_config.yml. The data loader will automatically use your custom annotations as long as they follow the format expected by the dataset class (e.g., SimpleDataSet or LMDBDataSet).
What is the difference between single-task and end-to-end evaluation?
Single-task evaluation (tools/eval.py) assesses one component at a time—either detection, recognition, or classification—using task-specific metrics like accuracy or H-mean. End-to-end evaluation (tools/end2end/eval_end2end.py) assesses the complete pipeline by matching predicted text regions against ground truth using polygon IoU, then checking transcription accuracy, providing a holistic F-measure score.
Can I use mixed precision (AMP) during evaluation?
Yes. When running python tools/eval.py, the evaluation loop in tools/program.py automatically detects and uses Automatic Mixed Precision if configured in your YAML file's Global section. This speeds up inference on supported GPUs without affecting metric accuracy.
How do I add a custom metric for a new task?
Create a new metric class in ppocr/metrics/ (e.g., custom_metric.py) that implements __call__(self, preds, batch) to process batch results and get_metric() to return final scores. Register it in ppocr/metrics/__init__.py, then reference it in your config file's Metric section. The existing tools/eval.py driver will automatically use your custom metric without requiring changes to the evaluation script.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →