How to Use OCR Models for Text Detection and Recognition with ailia-models
Use the PP-OCR or EasyOCR pipelines in axinc-ai/ailia-models to detect text regions with DB/CRAFT detectors and recognize characters with CRNN models via the ailia.Net() runtime.
The axinc-ai/ailia-models repository provides production-ready OCR implementations that combine text detection and text recognition into unified pipelines. This guide explains how to use OCR models for text detection and recognition using the PP-OCR (PaddleOCR) and EasyOCR implementations, including automated model setup, inference configuration, and result visualization.
PP-OCR and EasyOCR Pipeline Architecture
The repository bundles two distinct OCR implementations that share a common six-stage architecture:
- Model download –
check_and_download_models()pulls ONNX weights fromhttps://storage.googleapis.com/ailia-models/ - Model loading –
ailia.Net()initializes the runtime inference engine - Text detection – Identifies bounding boxes (quadrilaterals) for text regions
- Angle classification – Optional rotation correction using a lightweight classifier
- Text recognition – CRNN-based decoding using language-specific dictionaries
- Visualization –
draw_ocr_box_txt()overlays results on the source image
PP-OCR (PaddleOCR) uses a DB (Differentiable Binarization) detector for fast, lightweight text detection paired with a CRNN recognizer. This implementation supports Japanese, English, Chinese, German, French, and Korean. The primary entry point is text_recognition/paddleocr/paddleocr.py.
EasyOCR employs a CRAFT detector for high-quality region proposals and G2-style recognizers. This pipeline supports Chinese, Japanese, English, French, Korean, and Thai. The entry point is text_recognition/easyocr/easyocr.py, with utility functions defined in text_recognition/easyocr/easyocr_utils.py.
Model Management and Setup
Automatic Model Download
All OCR scripts use check_and_download_models() from util/model_utils.py to verify local model files against the remote Google Cloud bucket. If the local weight file is missing or outdated, the utility automatically downloads the required ONNX and Prototxt files.
from model_utils import check_and_download_models
REMOTE_PATH = 'https://storage.googleapis.com/ailia-models/paddle_ocr/'
check_and_download_models(DETECTOR_WEIGHT_PATH, DETECTOR_MODEL_PATH, REMOTE_PATH)
Loading ONNX Models with ailia
The ailia.Net() class abstracts the underlying inference engine, supporting CPU, GPU, or NPU execution environments. All OCR scripts expose the --env_id argument to select the target device.
import ailia
detector = ailia.Net(DETECTOR_MODEL_PATH, DETECTOR_WEIGHT_PATH, env_id=args.env_id)
recognizer = ailia.Net(REC_MODEL_PATH, REC_WEIGHT_PATH, env_id=args.env_id)
Text Detection Implementation
PP-OCR DB Detector
In text_recognition/paddleocr/paddleocr.py, the detector_predict() function (around lines 600-700) implements the DB detection algorithm. The function normalizes the input image, runs detector.run() to generate a raw score map, and applies configurable thresholds:
det_db_thresh– Binarization threshold for the probability mapdet_db_box_thresh– Confidence threshold for bounding box filteringdet_db_unclip_ratio– Expansion ratio for text region dilation
The output is a list of quadrilateral bounding boxes representing text line regions.
EasyOCR CRAFT Detector
The EasyOCR implementation in text_recognition/easyocr/easyocr_utils.py uses detector_predict() to invoke the CRAFT detector. This produces both horizontal and free-form text box lists optimized for curved or irregular text layouts.
Text Recognition and Post-Processing
Optional Angle Classification
PP-OCR includes a lightweight angle classifier configured via cls_model_path. When use_angle_cls is enabled (around line 630 in paddleocr.py), the classifier predicts whether the image requires 180° rotation to ensure upright text before recognition.
CRNN Recognition Pipeline
Both implementations use CRNN-style recognizers that follow a four-step process:
- Crop detected regions using quadrilateral vertices
- Resize to the recognizer's input shape (
rec_image_shape) - Inference via
recognizer.run()to generate logits - Decode using language-specific dictionaries (
*_dict.txt) to map logits to glyphs
In paddleocr.py, the recognizer_predict() function (around line 850) handles this pipeline, filtering low-confidence results using the drop_score parameter. The EasyOCR equivalent resides in easyocr_utils.py (around line 300).
Visualization
The draw_ocr_box_txt() function renders bounding boxes and recognized text on the original image using Pillow (ImageDraw, ImageFont) for high-resolution text rendering and OpenCV for final output. This function is implemented in both paddleocr.py and easyocr_utils.py.
Command-Line Usage Examples
PP-OCR Quick Start
Detect and recognize text in a single image using the default Japanese mobile model:
python text_recognition/paddleocr/paddleocr.py -i input.jpg -s output.png
Specify English language and server-size model for higher accuracy:
python text_recognition/paddleocr/paddleocr.py \
-i input.jpg -s output.png \
-l english -c server
Process a live video stream from camera index 0:
python text_recognition/paddleocr/paddleocr.py -v 0 -s out.mp4
All arguments are defined in the arg_utils parser (see get_base_parser in util/arg_utils.py). The main execution block resides at lines 1390-1420 in paddleocr.py.
EasyOCR Quick Start
Detect and recognize Chinese text (default):
python text_recognition/easyocr/easyocr.py -i input.jpg -s result.png
Switch to Japanese language support:
python text_recognition/easyocr/easyocr.py -l japanese -i input.jpg -s result.png
Process a video file with side-by-side visualization:
python text_recognition/easyocr/easyocr.py -v sample.mp4 -s out.mp4
The entry point logic is located at lines 45-55 in easyocr.py.
Programmatic Python API Integration
Embed OCR capabilities directly into Python applications by importing the pipeline functions:
import ailia
import cv2
from text_recognition.paddleocr.paddleocr import (
get_default_config, set_config, transform,
draw_ocr_box_txt, check_and_download_models
)
# 1️⃣ Configure model paths and parameters
cfg = get_default_config()
cfg = set_config(
cfg,
weight_path_det='chi_eng_num_sym_server_det_org.onnx',
det_algorithm='DB',
weight_path_rec='jpn_eng_num_sym_server_rec_add.onnx',
dict_path_rec='./dict/jpn_eng_num_sym_add.txt',
weight_path_cls='chi_eng_num_sym_mobile_cls_org.onnx'
)
# 2️⃣ Download models if not present
REMOTE = 'https://storage.googleapis.com/ailia-models/paddle_ocr/'
check_and_download_models(cfg['det_model_path'], cfg['det_model_path'], REMOTE)
check_and_download_models(cfg['rec_model_path'], cfg['rec_model_path'], REMOTE)
# 3️⃣ Initialize inference engines
detector = ailia.Net(cfg['det_model_path'], cfg['det_model_path'])
recognizer = ailia.Net(cfg['rec_model_path'], cfg['rec_model_path'])
# 4️⃣ Run inference (simplified - see predict() function for full implementation)
img = cv2.imread('input.jpg')
# boxes = detector.run(img) # Actual implementation uses preprocessing helpers
# results = recognizer.run(transform(img, ...))
# 5️⃣ Visualize results
# vis = draw_ocr_box_txt(img, results)
# cv2.imwrite('annotated.png', vis)
For complete preprocessing and post-processing logic, reference the predict() function in paddleocr.py (lines 640-690) and the utility functions in easyocr_utils.py.
Summary
- Two pipeline options: PP-OCR (DB + CRNN) for speed and multi-language support, or EasyOCR (CRAFT + G2) for high-quality curved text detection.
- Automated setup:
check_and_download_models()inutil/model_utils.pyhandles model verification and downloading from Google Cloud Storage. - Flexible deployment: Use
ailia.Net()with--env_idto target CPU, GPU, or NPU execution environments. - Complete workflow: Detection → optional angle classification → recognition → visualization via
draw_ocr_box_txt(). - Multiple interfaces: Command-line scripts support batch processing, video streams, and language selection via
-land-cflags.
Frequently Asked Questions
What is the difference between PP-OCR and EasyOCR in ailia-models?
PP-OCR uses a DB (Differentiable Binarization) detector optimized for speed and lightweight deployment, supporting Japanese, English, Chinese, German, French, and Korean. EasyOCR employs a CRAFT detector that excels at detecting curved or irregular text regions, supporting Chinese, Japanese, English, French, Korean, and Thai. Both use CRNN-based recognizers, but PP-OCR offers mobile and server model variants via the -c flag.
How do I add support for a new language not included in the default dictionaries?
Create a character dictionary file (e.g., custom_dict.txt) containing all glyphs the recognizer should output, then pass the path via dict_path_rec in the configuration. You must also obtain or train a CRNN recognizer model (ONNX format) that outputs logits matching your dictionary length. Place both files in the appropriate text_recognition/paddleocr/dict/ directory or specify absolute paths in set_config().
Can I run these OCR models on edge devices without GPUs?
Yes. The ailia.Net() runtime supports CPU inference and exposes the --env_id parameter to select execution backends. Use -c mobile with PP-OCR for quantized, lightweight models optimized for ARM CPUs and edge NPUs. The DB detector and mobile CRNN models are specifically designed for real-time inference on resource-constrained devices.
How do I adjust detection sensitivity for dense or small text?
Modify the DB detector parameters in the PP-OCR configuration: decrease det_db_thresh to detect fainter text, reduce det_db_box_thresh to accept lower-confidence regions, or increase det_db_unclip_ratio to expand tight bounding boxes around text clusters. These parameters are passed through set_config() or command-line arguments when available.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →