How to Use PaddleOCR for Text Detection: CLI and Python API Guide

PaddleOCR performs text detection by loading a pre-trained inference model, preprocessing images according to size constraints, running neural network inference, and post-processing results to return bounding box coordinates with confidence scores.

PaddleOCR is an open-source optical character recognition framework maintained by Baidu's PaddlePaddle team. To use PaddleOCR for text detection, you can invoke the TextDetection class directly through Python or execute the text_detection CLI subcommand. This article explores the modular pipeline architecture, essential tuning parameters, and practical code implementations based on the source code in the PaddlePaddle/PaddleOCR repository.

Core Architecture of the Detection Pipeline

PaddleOCR follows a modular design where text detection operates as a distinct stage from recognition. The detection component locates text-line bounding boxes and outputs their coordinates together with confidence scores.

Pipeline Components

Component Description Source File
User-Facing Wrapper TextDetection inherits from TextDetectionMixin (argument handling) and PaddleXPredictorWrapper (model loading). paddleocr/_models/text_detection.py
Argument Mixin TextDetectionMixin stores hyper-parameters including limit_side_len, thresh, box_thresh, unclip_ratio, and input_shape. paddleocr/_models/_text_detection.py
CLI Executor TextDetectionSubcommandExecutor registers the text_detection subcommand and forwards arguments to perform_simple_inference. paddleocr/_models/text_detection.py
Inference Engine PaddleXPredictorWrapper wraps paddle.inference.create_predictor to build the predictor from the supplied model directory. paddleocr/_models/base.py
Post-Processing PicoDetPostProcess (or similar) thresholds the probability map, filters boxes by confidence, and applies the unclip operation. ppocr/postprocess/picodet_postprocess.py

Five-Step Detection Flow

  1. Load model – Initialize the predictor using the specified weights and configuration.
  2. Pre-process image – Scale the image according to limit_side_len and limit_type while preserving aspect ratio.
  3. Run inference – Execute the neural network forward pass through the Paddle inference engine.
  4. Post-process – Apply pixel-level thresholding (thresh), box-level filtering (box_thresh), non-maximum suppression, and box expansion via unclip_ratio.
  5. Return boxes – Output a list of coordinates [x_min, y_min, x_max, y_max] with associated confidence scores.

Running Text Detection via Command Line

The CLI provides the fastest way to use PaddleOCR for text detection on single images or batches without writing Python code.

paddleocr text_detection \
    -i https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/general_ocr_002.png \
    --limit_side_len 960 \
    --thresh 0.3 \
    --box_thresh 0.5 \
    --unclip_ratio 2.0

The text_detection subcommand invokes TextDetectionSubcommandExecutor, which processes the image and outputs JSON containing bounding boxes and scores. Adjust --limit_side_len to control memory usage and inference speed for high-resolution images.

Using the Python API for Stand-Alone Detection

For embedding detection into custom scripts, initialize PaddleOCR with recognition disabled to run only the detector.

from paddleocr import PaddleOCR

# Initialize detector only (no recognizer)

ocr = PaddleOCR(
    det=True,                # explicitly enable detection (default)

    rec=False,               # disable recognition

    use_angle_cls=False,     # skip orientation classification

    limit_side_len=960,
    thresh=0.3,
    box_thresh=0.5,
    unclip_ratio=2.0,
)

# Run detection on a remote or local image

result = ocr.ocr(
    "https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/general_ocr_002.png",
    det=True,
    rec=False
)

# Parse results: list of dicts with 'box' and 'score' keys

for line in result:
    print(f"Box: {line['box']}, Score: {line['score']}")

The constructor forwards detection arguments to TextDetectionMixin, which validates and stores them for the inference stage. Setting rec=False skips the recognition model entirely, reducing latency when you only need bounding box coordinates.

Integrating Detection into Full OCR Pipelines

To use PaddleOCR for text detection as part of a complete workflow that includes recognition, enable both stages in the constructor.

from paddleocr import PaddleOCR

# Full pipeline: detection + recognition

ocr = PaddleOCR(
    use_angle_cls=False,
    limit_side_len=960,
    thresh=0.3,
    box_thresh=0.5,
    unclip_ratio=2.0,
)

# Process a local image

result = ocr.ocr("docs/images/demo_image/general_ocr_002.png")

# Each element contains the bounding box and recognized text tuple

for line in result:
    box, (text, score) = line
    print(f"Box: {box}, Text: '{text}', Score: {score:.3f}")

In this configuration, PaddleOCR first runs the detection model to generate candidate boxes, then feeds cropped regions to the recognition model. The output structure follows the [bounding_box, (text, confidence)] format documented in the repository README.

Key Detection Parameters Explained

Tuning these parameters in paddleocr/_models/_text_detection.py directly affects accuracy and inference speed.

Parameter Function Typical Range
limit_side_len Maximum dimension (width or height) after resizing; larger values improve small-text detection but increase compute. 640–960
limit_type Determines whether limit_side_len constrains the max (longer) or min (shorter) image side. max or min
thresh Pixel-level confidence threshold on the probability map; higher values reduce false positives. 0.2–0.4
box_thresh Minimum average confidence inside a detected box; filters low-quality regions. 0.5–0.7
unclip_ratio Multiplier expanding the shrunk polygon back to the full text region; increase for dense text. 1.5–2.5
input_shape Static (C, H, W) tensor shape for optimized inference engines like TensorRT or ONNX. (3, 640, 640)

As implemented in ppocr/postprocess/picodet_postprocess.py, the post-processor first applies thresh to binarize the probability map, then averages pixel scores within each connected component to enforce box_thresh, and finally dilates the contour by unclip_ratio to recover the full text boundary.

Summary

  • PaddleOCR text detection is implemented through the TextDetection class, which combines TextDetectionMixin for configuration and PaddleXPredictorWrapper for inference.
  • The pipeline executes five distinct stages: model loading, image preprocessing, neural network inference, post-processing (thresholding and unclipping), and result formatting.
  • You can invoke detection via the paddleocr text_detection CLI for quick tasks or through the Python API by setting rec=False for stand-alone detection.
  • Critical tuning parameters—thresh, box_thresh, unclip_ratio, and limit_side_len—control the trade-off between detection recall and precision.
  • Source files paddleocr/_models/text_detection.py and ppocr/postprocess/picodet_postprocess.py contain the core logic for model execution and result refinement.

Frequently Asked Questions

How do I adjust the sensitivity of text detection in PaddleOCR?

Modify the thresh and box_thresh parameters when initializing the PaddleOCR class. Lower thresh values (e.g., 0.2) detect faint or low-contrast text but may introduce noise, while higher box_thresh values (e.g., 0.6) filter out uncertain detections.

What is the difference between rec=False and det=True in the Python API?

Setting rec=False disables the recognition model entirely, causing ocr.ocr() to return only bounding boxes and confidence scores. Setting det=True (the default) ensures the detection model runs; if both det=True and rec=True, the method returns recognized text alongside each box.

Can I run PaddleOCR text detection without GPU support?

Yes, the PaddleXPredictorWrapper automatically falls back to CPU execution if PaddlePaddle is installed without CUDA support. Specify --use_gpu false in CLI commands or ensure no GPU flag is passed in Python to force CPU inference.

Where is the bounding box expansion logic implemented?

The unclipping algorithm that expands detected boxes by the unclip_ratio multiplier resides in ppocr/postprocess/picodet_postprocess.py within the PicoDetPostProcess class. This post-processor converts the shrunk text regions predicted by the neural network back to the original text boundaries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →