How to Evaluate PaddleOCR Model Performance: A Complete Guide

Use tools/eval.py to run standardized evaluation on validation data, or call program.eval() directly in Python to compute task-specific metrics like accuracy for recognition or H-mean for detection.

PaddleOCR provides a unified evaluation framework that handles both single-task models (text detection, recognition, classification) and complex end-to-end document understanding pipelines. The evaluation system is configuration-driven, allowing you to assess model quality using YAML config files without modifying source code, while the underlying implementation in tools/program.py provides a reusable loop for custom integrations.

Overview of the Evaluation Architecture

The PaddleOCR evaluation flow follows a strict pipeline: data loading → inference → post-processing → metric aggregation. This architecture is implemented across several core modules:

  • tools/eval.py – The command-line entry point that orchestrates the entire process (lines 35-78).
  • tools/program.py – Contains the generic eval() function (lines 61-74) that executes the inference loop.
  • ppocr/metrics/ – Task-specific metric classes (RecMetric, DetMetric, E2EMetric) that compute final scores.
  • tools/end2end/eval_end2end.py – Specialized evaluator for full pipelines that calculates polygon IoU and edit distance.

For single-task evaluation, the system uses build_dataloader() to prepare validation data, build_model() to construct the architecture, and build_metric() to instantiate the appropriate scorer based on the config file's Metric section.

Single-Task Model Evaluation

Single-task evaluation applies to text detection, text recognition, and angle classification models. You can run evaluation via CLI for quick assessment or use the Python API for custom workflows.

Command-Line Evaluation

The fastest way to evaluate a trained model is using the provided evaluation script with a configuration file:


# Install PaddleOCR

pip install "paddleocr[all]"

# Run evaluation (example for recognition model)

python tools/eval.py -c configs/rec/rec_ch_PP-OCRv3.yml

Ensure your configuration file contains a valid Eval dataset section pointing to your validation images and annotations. The script automatically loads the checkpoint specified in the config, runs inference on the validation set, and outputs metrics:


metric eval ***************
accuracy:0.9578
norm_edit_dis:0.0412
fps: 123.4

The tools/eval.py script handles device placement, mixed precision (AMP) scaling when enabled, and checkpoint loading via load_model() from ppocr/utils/save_load.py.

Python API Evaluation

For integration into custom pipelines or Jupyter notebooks, import the evaluation components directly:

import yaml
import paddle
from ppocr.data import build_dataloader
from ppocr.modeling.architectures import build_model
from ppocr.postprocess import build_post_process
from ppocr.metrics import build_metric
from ppocr.utils.save_load import load_model
from tools.program import eval as paddleocr_eval

# Load configuration

config = yaml.safe_load(open('configs/rec/rec_ch_PP-OCRv3.yml', 'rb'))
device = paddle.set_device('gpu' if paddle.is_compiled_with_cuda() else 'cpu')

# Build components

val_loader = build_dataloader(config, 'Eval', device, logger=None)
model = build_model(config['Architecture'])
load_model(config, model, model_type=config['Architecture']['model_type'])
post_process = build_post_process(config['PostProcess'], config['Global'])
metric = build_metric(config['Metric'])

# Execute evaluation

results = paddleocr_eval(
    model,
    val_loader,
    post_process,
    metric,
    model_type=config['Architecture'].get('model_type')
)
print(f"Final accuracy: {results}")

This approach gives you full control over the data loader, model initialization, and metric computation while maintaining compatibility with PaddleOCR's standardized evaluation logic.

Key Evaluation Components

Understanding these factory functions helps debug evaluation issues or implement custom metrics:

  • build_dataloader(config, 'Eval', device, ...) – Constructs the validation data loader from ppocr/data/__init__.py using the dataset paths and transforms defined in the config's Eval section.
  • build_model(config['Architecture']) – Instantiates the neural network architecture; for distillation models, this automatically adjusts head output channels.
  • build_post_process(config['PostProcess'], global_config) – Converts raw model outputs (logits or heatmaps) into interpretable predictions like text strings or bounding boxes.
  • build_metric(config['Metric']) – Creates the metric calculator that aggregates batch-wise statistics into final scores such as accuracy for recognition or H-mean for detection.

End-to-End Document Evaluation

For complete OCR pipelines that include layout analysis, text detection, and recognition in one system, use the specialized end-to-end evaluator:

python tools/end2end/eval_end2end.py \
    /path/to/ground_truth \
    /path/to/predictions \
    --ignore_blank=True

This script computes polygon IoU (Intersection over Union) between predicted and ground-truth text regions using the polygon_iou() function, then calculates character-level edit distance via the ed() function. It aggregates these into precision, recall, and F-measure scores:


hit, dt_count, gt_count  1250 1300 1240
character_acc: 93.21%
avg_edit_dist_field: 1.84
precision: 96.15%
recall: 92.31%
fmeasure: 94.19%

The e2e_eval() function (lines 70-80 in tools/end2end/eval_end2end.py) handles the complex matching logic required when multiple text regions overlap or when layout elements must be considered alongside text content.

Understanding Evaluation Metrics

Different tasks use different primary indicators:

  • Text Recognition: Uses RecMetric to calculate accuracy (exact match) and normalized edit distance (character error rate).
  • Text Detection: Uses DetMetric to compute precision, recall, and H-mean (F1-score) based on box IoU thresholds.
  • End-to-End: Uses E2EMetric or the standalone eval_end2end.py to combine detection IoU with transcription accuracy, producing an overall F-measure that reflects both localization and recognition quality.

All metric classes implement a consistent interface: they accumulate batch results via __call__() and return a dictionary of final values through get_metric().

Summary

  • Primary entry point: Use tools/eval.py for standard single-task evaluation and tools/end2end/eval_end2end.py for full document pipelines.
  • Configuration-driven: Evaluation requires a YAML config file with a properly defined Eval dataset section and Metric specification.
  • Core workflow: build_dataloader() → build_model() → load_model() → program.eval() → metric aggregation.
  • Key files: tools/program.py contains the reusable evaluation loop; ppocr/metrics/ contains task-specific scoring logic.
  • Metrics vary by task: Recognition uses accuracy, detection uses H-mean, and end-to-end uses F-measure combined with edit distance.

Frequently Asked Questions

How do I evaluate a model on a custom dataset?

Prepare a YAML configuration file that specifies your validation image paths and label files in the Eval.dataset section, then run python tools/eval.py -c your_config.yml. The data loader will automatically use your custom annotations as long as they follow the format expected by the dataset class (e.g., SimpleDataSet or LMDBDataSet).

What is the difference between single-task and end-to-end evaluation?

Single-task evaluation (tools/eval.py) assesses one component at a time—either detection, recognition, or classification—using task-specific metrics like accuracy or H-mean. End-to-end evaluation (tools/end2end/eval_end2end.py) assesses the complete pipeline by matching predicted text regions against ground truth using polygon IoU, then checking transcription accuracy, providing a holistic F-measure score.

Can I use mixed precision (AMP) during evaluation?

Yes. When running python tools/eval.py, the evaluation loop in tools/program.py automatically detects and uses Automatic Mixed Precision if configured in your YAML file's Global section. This speeds up inference on supported GPUs without affecting metric accuracy.

How do I add a custom metric for a new task?

Create a new metric class in ppocr/metrics/ (e.g., custom_metric.py) that implements __call__(self, preds, batch) to process batch results and get_metric() to return final scores. Register it in ppocr/metrics/__init__.py, then reference it in your config file's Metric section. The existing tools/eval.py driver will automatically use your custom metric without requiring changes to the evaluation script.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →