How to Use PaddleOCR for Text Detection: CLI and Python API Guide
PaddleOCR performs text detection by loading a pre-trained inference model, preprocessing images according to size constraints, running neural network inference, and post-processing results to return bounding box coordinates with confidence scores.
PaddleOCR is an open-source optical character recognition framework maintained by Baidu's PaddlePaddle team. To use PaddleOCR for text detection, you can invoke the TextDetection class directly through Python or execute the text_detection CLI subcommand. This article explores the modular pipeline architecture, essential tuning parameters, and practical code implementations based on the source code in the PaddlePaddle/PaddleOCR repository.
Core Architecture of the Detection Pipeline
PaddleOCR follows a modular design where text detection operates as a distinct stage from recognition. The detection component locates text-line bounding boxes and outputs their coordinates together with confidence scores.
Pipeline Components
| Component | Description | Source File |
|---|---|---|
| User-Facing Wrapper | TextDetection inherits from TextDetectionMixin (argument handling) and PaddleXPredictorWrapper (model loading). |
paddleocr/_models/text_detection.py |
| Argument Mixin | TextDetectionMixin stores hyper-parameters including limit_side_len, thresh, box_thresh, unclip_ratio, and input_shape. |
paddleocr/_models/_text_detection.py |
| CLI Executor | TextDetectionSubcommandExecutor registers the text_detection subcommand and forwards arguments to perform_simple_inference. |
paddleocr/_models/text_detection.py |
| Inference Engine | PaddleXPredictorWrapper wraps paddle.inference.create_predictor to build the predictor from the supplied model directory. |
paddleocr/_models/base.py |
| Post-Processing | PicoDetPostProcess (or similar) thresholds the probability map, filters boxes by confidence, and applies the unclip operation. |
ppocr/postprocess/picodet_postprocess.py |
Five-Step Detection Flow
- Load model – Initialize the predictor using the specified weights and configuration.
- Pre-process image – Scale the image according to
limit_side_lenandlimit_typewhile preserving aspect ratio. - Run inference – Execute the neural network forward pass through the Paddle inference engine.
- Post-process – Apply pixel-level thresholding (
thresh), box-level filtering (box_thresh), non-maximum suppression, and box expansion viaunclip_ratio. - Return boxes – Output a list of coordinates
[x_min, y_min, x_max, y_max]with associated confidence scores.
Running Text Detection via Command Line
The CLI provides the fastest way to use PaddleOCR for text detection on single images or batches without writing Python code.
paddleocr text_detection \
-i https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/general_ocr_002.png \
--limit_side_len 960 \
--thresh 0.3 \
--box_thresh 0.5 \
--unclip_ratio 2.0
The text_detection subcommand invokes TextDetectionSubcommandExecutor, which processes the image and outputs JSON containing bounding boxes and scores. Adjust --limit_side_len to control memory usage and inference speed for high-resolution images.
Using the Python API for Stand-Alone Detection
For embedding detection into custom scripts, initialize PaddleOCR with recognition disabled to run only the detector.
from paddleocr import PaddleOCR
# Initialize detector only (no recognizer)
ocr = PaddleOCR(
det=True, # explicitly enable detection (default)
rec=False, # disable recognition
use_angle_cls=False, # skip orientation classification
limit_side_len=960,
thresh=0.3,
box_thresh=0.5,
unclip_ratio=2.0,
)
# Run detection on a remote or local image
result = ocr.ocr(
"https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/general_ocr_002.png",
det=True,
rec=False
)
# Parse results: list of dicts with 'box' and 'score' keys
for line in result:
print(f"Box: {line['box']}, Score: {line['score']}")
The constructor forwards detection arguments to TextDetectionMixin, which validates and stores them for the inference stage. Setting rec=False skips the recognition model entirely, reducing latency when you only need bounding box coordinates.
Integrating Detection into Full OCR Pipelines
To use PaddleOCR for text detection as part of a complete workflow that includes recognition, enable both stages in the constructor.
from paddleocr import PaddleOCR
# Full pipeline: detection + recognition
ocr = PaddleOCR(
use_angle_cls=False,
limit_side_len=960,
thresh=0.3,
box_thresh=0.5,
unclip_ratio=2.0,
)
# Process a local image
result = ocr.ocr("docs/images/demo_image/general_ocr_002.png")
# Each element contains the bounding box and recognized text tuple
for line in result:
box, (text, score) = line
print(f"Box: {box}, Text: '{text}', Score: {score:.3f}")
In this configuration, PaddleOCR first runs the detection model to generate candidate boxes, then feeds cropped regions to the recognition model. The output structure follows the [bounding_box, (text, confidence)] format documented in the repository README.
Key Detection Parameters Explained
Tuning these parameters in paddleocr/_models/_text_detection.py directly affects accuracy and inference speed.
| Parameter | Function | Typical Range |
|---|---|---|
| limit_side_len | Maximum dimension (width or height) after resizing; larger values improve small-text detection but increase compute. | 640–960 |
| limit_type | Determines whether limit_side_len constrains the max (longer) or min (shorter) image side. |
max or min |
| thresh | Pixel-level confidence threshold on the probability map; higher values reduce false positives. | 0.2–0.4 |
| box_thresh | Minimum average confidence inside a detected box; filters low-quality regions. | 0.5–0.7 |
| unclip_ratio | Multiplier expanding the shrunk polygon back to the full text region; increase for dense text. | 1.5–2.5 |
| input_shape | Static (C, H, W) tensor shape for optimized inference engines like TensorRT or ONNX. | (3, 640, 640) |
As implemented in ppocr/postprocess/picodet_postprocess.py, the post-processor first applies thresh to binarize the probability map, then averages pixel scores within each connected component to enforce box_thresh, and finally dilates the contour by unclip_ratio to recover the full text boundary.
Summary
- PaddleOCR text detection is implemented through the
TextDetectionclass, which combinesTextDetectionMixinfor configuration andPaddleXPredictorWrapperfor inference. - The pipeline executes five distinct stages: model loading, image preprocessing, neural network inference, post-processing (thresholding and unclipping), and result formatting.
- You can invoke detection via the
paddleocr text_detectionCLI for quick tasks or through the Python API by settingrec=Falsefor stand-alone detection. - Critical tuning parameters—
thresh,box_thresh,unclip_ratio, andlimit_side_len—control the trade-off between detection recall and precision. - Source files
paddleocr/_models/text_detection.pyandppocr/postprocess/picodet_postprocess.pycontain the core logic for model execution and result refinement.
Frequently Asked Questions
How do I adjust the sensitivity of text detection in PaddleOCR?
Modify the thresh and box_thresh parameters when initializing the PaddleOCR class. Lower thresh values (e.g., 0.2) detect faint or low-contrast text but may introduce noise, while higher box_thresh values (e.g., 0.6) filter out uncertain detections.
What is the difference between rec=False and det=True in the Python API?
Setting rec=False disables the recognition model entirely, causing ocr.ocr() to return only bounding boxes and confidence scores. Setting det=True (the default) ensures the detection model runs; if both det=True and rec=True, the method returns recognized text alongside each box.
Can I run PaddleOCR text detection without GPU support?
Yes, the PaddleXPredictorWrapper automatically falls back to CPU execution if PaddlePaddle is installed without CUDA support. Specify --use_gpu false in CLI commands or ensure no GPU flag is passed in Python to force CPU inference.
Where is the bounding box expansion logic implemented?
The unclipping algorithm that expands detected boxes by the unclip_ratio multiplier resides in ppocr/postprocess/picodet_postprocess.py within the PicoDetPostProcess class. This post-processor converts the shrunk text regions predicted by the neural network back to the original text boundaries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →