How NekoImageGallery Extracts and Indexes Text from Images Using PaddleOCR

NekoImageGallery extracts text from images using a configurable PaddleOCR pipeline that preprocesses images, filters low-confidence results, and stores both raw text and BERT vectors for hybrid exact-match and semantic search.

The hv0905/nekoimagegallery repository implements a complete OCR search system that converts visual text into searchable data. This article examines the end-to-end flow from configuration through text extraction to vector database indexing, demonstrating how the application leverages PaddleOCR to make image content discoverable.

Configuration and Service Initialization

The OCR feature is controlled through OCRSearchSettings in app/config.py. Administrators enable the functionality by setting ocr_search.enable to True and select the backend via ocr_search.ocr_module, which defaults to "easypaddleocr" but can be switched to "paddleocr" for direct PaddleOCR integration.


# app/config.py

class OCRSearchSettings(BaseModel):
    enable: bool = True                  # Master toggle for OCR features

    ocr_module: str = "paddleocr"        # Backend selection

    ocr_language: list[str] = ["ch_sim", "en"]
    ocr_min_confidence: float = 1e-2     # Confidence threshold filtering

During application startup, the ServiceProvider class in app/Services/provider.py instantiates the appropriate OCR service based on these settings. When ocr_module is set to "paddleocr", the provider imports and creates a PaddleOCRService instance.


# app/Services/provider.py

if config.ocr_search.enable and (environment.local_indexing or config.admin_api_enable):
    match config.ocr_search.ocr_module:
        case "paddleocr":
            from .ocr_services import PaddleOCRService
            self.ocr_service = PaddleOCRService()

The OCR Pipeline: From Image Upload to Text Extraction

When an image is uploaded, the IndexService._prepare_image method in app/Services/index_service.py coordinates the extraction process. If OCR is enabled and not explicitly skipped, the method forwards the PIL Image object to the OCR service's ocr_interface method.

Image Preprocessing

Before text recognition, the OCRService base class optionally preprocesses images to ensure consistent input dimensions. The _image_preprocess method in app/Services/ocr_services.py resizes images exceeding 1024 pixels on any axis and pads them to a square 1024×1024 canvas, preventing distortion while standardizing the input for the OCR model.


# app/Services/ocr_services.py

def _image_preprocess(self, img: Image.Image) -> Image.Image:
    # Resizes to max 1024px and pads to 1024x1024 square

    max_size = 1024
    if max(img.size) > max_size:
        ratio = max_size / max(img.size)
        new_size = tuple(int(x * ratio) for x in img.size)
        img = img.resize(new_size, Image.Resampling.LANCZOS)
    # Padding logic to create square canvas...

    return padded_img

Text Extraction with PaddleOCR

The PaddleOCRService class initializes the underlying paddleocr library in its constructor with Chinese language support (lang="ch") and GPU acceleration when available. The ocr_interface method orchestrates the extraction by optionally preprocessing the image and delegating to _paddleocr_process.


# app/Services/ocr_services.py

class PaddleOCRService(OCRService):
    def __init__(self):
        super().__init__()
        import paddleocr
        self._paddle_ocr_module = paddleocr.PaddleOCR(
            lang="ch", use_angle_cls=True, use_gpu=self._device == "cuda"
        )

    def ocr_interface(self, img: Image.Image, need_preprocess=True) -> str:
        start_time = time()
        logger.info("Processing text with PaddleOCR...")
        res = self._paddleocr_process(
            self._image_preprocess(img) if need_preprocess else img
        )
        logger.success("OCR processed done. Time elapsed: {:.2f}s", time() - start_time)
        return res

Confidence Filtering

The _paddleocr_process method filters PaddleOCR's raw output based on the configured confidence threshold. Only words exceeding config.ocr_search.ocr_min_confidence (default 0.01) are concatenated into the final text string, eliminating noisy low-confidence detections.


# app/Services/ocr_services.py

def _paddleocr_process(self, img: Image.Image) -> str:
    ocr_result = self._paddle_ocr_module.ocr(np.array(img), cls=True)
    if ocr_result[0]:
        return "".join(
            itm[1][0] for itm in ocr_result[0]
            if itm[1][1] > config.ocr_search.ocr_min_confidence
        )
    return ""

Once extracted, the text flows through a dual-storage strategy enabling both exact-match and semantic search capabilities.

Storing Raw Text and Vectors

In app/Services/index_service.py, the _prepare_image method stores the extracted text in the MappedImage.ocr_text field. Simultaneously, it generates a dense vector representation using the TransformersService BERT model, storing this in text_contain_vector for semantic similarity search.


# app/Services/index_service.py

def _prepare_image(self, image: Image.Image, image_data: MappedImage, skip_ocr=False):
    # ... image vector extraction omitted ...

    if not skip_ocr and config.ocr_search.enable:
        image_data.ocr_text = self._ocr_service.ocr_interface(image)
        if image_data.ocr_text != "":
            image_data.text_contain_vector = self._transformers_service.get_bert_vector(
                image_data.ocr_text
            )
        else:
            image_data.ocr_text = None

Database Persistence

The MappedImage data model in app/Models/mapped_image.py defines both ocr_text and a computed ocr_text_lower property. When VectorDbContext.insert_items persists the record to Qdrant, it includes both the raw string and a lower-cased copy. The lower-cased version supports case-insensitive exact matching without requiring runtime transformations.

Querying Images by OCR Content

During search operations, VectorDbContext._filter_by_params in app/Services/vector_db_context.py constructs Qdrant filters. When the request includes an OCR text parameter, the system appends a MatchText condition against the ocr_text_lower field, enabling substring exact-match filtering within the vector database.


# app/Services/vector_db_context.py

if filter_param.ocr_text is not None:
    filters.append(models.FieldCondition(
        key="ocr_text_lower",
        match=models.MatchText(text=filter_param.ocr_text.lower())
    ))

This hybrid approach allows users to filter images by specific text strings while simultaneously performing semantic vector searches on the visual and textual content.

Summary

  • Configuration-driven architecture: Toggle OCR and select backends via OCRSearchSettings in app/config.py without code changes.
  • PaddleOCR integration: The PaddleOCRService class in app/Services/ocr_services.py initializes the engine with GPU support and Chinese language models.
  • Quality control: The ocr_min_confidence threshold filters out low-quality text detections during the _paddleocr_process stage.
  • Dual indexing strategy: Stores both raw OCR text (ocr_text) and BERT vectors (text_contain_vector) for hybrid search capabilities.
  • Case-insensitive exact matching: The ocr_text_lower field enables efficient substring filtering within Qdrant vector queries.

Frequently Asked Questions

How do I switch from EasyPaddleOCR to standard PaddleOCR?

Set ocr_module to "paddleocr" in your configuration file. The ServiceProvider in app/Services/provider.py automatically instantiates PaddleOCRService when this value is detected, loading the standard paddleocr Python package instead of the EasyPaddleOCR wrapper.

What preprocessing is applied to images before OCR?

The system resizes images exceeding 1024 pixels on any dimension and pads them to a 1024×1024 square canvas. This standardization occurs in OCRService._image_preprocess within app/Services/ocr_services.py, ensuring consistent input sizes for the PaddleOCR model regardless of source image dimensions.

How does the system handle low-confidence OCR detections?

Words with confidence scores below config.ocr_search.ocr_min_confidence (default 0.01) are automatically discarded during the _paddleocr_process method. Only high-confidence text concatenations are returned for indexing, preventing database pollution from ambiguous characters or background noise.

Can I search for images containing specific text phrases?

Yes. The VectorDbContext._filter_by_params method implements a MatchText filter on the ocr_text_lower field. When you provide an ocr_text parameter in your search query, the system performs case-insensitive exact substring matching against indexed OCR content while maintaining the vector similarity search on visual features.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →