Does pdf-inspector Support OCR for Image-Based PDFs? How the Library Detects Scan Pages

pdf-inspector does not perform OCR itself; instead, it analyzes each page and flags those requiring OCR, returning a list of 1-indexed page numbers and machine-readable reasons for downstream processing.

The firecrawl/pdf-inspector library provides intelligent OCR detection for PDF documents without embedding an actual OCR engine. By analyzing text operators, image counts, and font encodings, it identifies which pages need optical character recognition while leaving the heavy compute work to specialized engines like Tesseract or PaddleOCR.

How pdf-inspector Detects Pages Requiring OCR

The detection system operates at two levels: document classification and per-page analysis. This dual approach enables hybrid pipelines where clean text pages process normally while only flagged pages trigger expensive OCR operations.

Document-Level PDF Classification

In src/detector.rs (lines ~350-400), the library first classifies the entire document into one of four categories: TextBased, Scanned, ImageBased, or Mixed. This initial categorization helps downstream systems understand the overall document composition before processing individual pages. The detector examines global PDF properties including font dictionaries, image XObject counts, and text extraction viability.

Per-Page Heuristic Analysis

Following document classification, the detector performs granular inspection of each page in src/detector.rs (lines ~404-426). It evaluates text operator streams, embedded image densities, and font encoding validity. Pages containing undecodable fonts, excessive image coverage, or missing text operators get added to the pages_needing_ocr collection. This targeted approach ensures compute resources get allocated only to pages actually requiring OCR conversion.

Understanding the OCR Flagging Architecture

The library exposes OCR requirements through structured data types that communicate not just which pages need processing, but why they need it.

The PageOcrReasons Struct

Defined in src/lib.rs (lines 105-119), the PageOcrReasons struct records specific justifications for OCR flagging. The system categorizes pages using enum variants including scanned for pure image pages, suspected_garbled_text for encoding failures, and vector_text for certain graphical text representations. These granular reasons allow OCR engines to apply appropriate preprocessing—enhancing scanned images differently than correcting font encoding issues.

The PdfProcessResult Payload

The public API returns detection results through the PdfProcessResult struct declared in src/lib.rs (lines ~145-149). This payload contains two critical fields: pages_needing_ocr (a vector of 1-indexed page numbers) and ocr_reasons_by_page (a mapping of page indices to their specific OCR justifications). For CLI users, src/bin/detect_pdf.rs exposes this information through the detect-pdf command, while src/bin/pdf2md.rs includes the flags in JSON output when using the --json flag.

Implementing Hybrid OCR Pipelines with pdf-inspector

The library's design enables efficient document processing by separating detection from recognition. You can extract the OCR requirements and route only flagged pages to external engines.

Rust Implementation

Use the detect_pdf function to analyze documents without generating markdown:

use pdf_inspector::{detect_pdf, PdfOptions};

fn main() -> Result<(), pdf_inspector::PdfError> {
    // Fast detection only (no markdown generation)
    let result = detect_pdf("example.pdf")?;

    println!("PDF type: {:?}", result.pdf_type);
    println!("Pages needing OCR: {:?}", result.pages_needing_ocr);
    for reason in result.ocr_reasons_by_page {
        println!("Page {} → reasons: {:?}", reason.page, reason.reasons);
    }
    Ok(())
}

Python Wrapper Usage

The Python FFI bindings in src/python.rs expose identical functionality:

from pdf_inspector import detect_pdf

result = detect_pdf("example.pdf")
print("PDF type:", result.pdf_type)
print("Pages needing OCR:", result.pages_needing_ocr)
for r in result.ocr_reasons_by_page:
    print(f"Page {r.page} → reasons: {r.reasons}")

Integrating with External OCR Engines

Route flagged pages to Tesseract or similar engines while preserving clean text from other pages:

for page in result.pages_needing_ocr {
    // Path to a temporary PNG rendered from the PDF page
    let img_path = render_page_to_png("example.pdf", page);
    // Run Tesseract (or any OCR library) on the image
    let text = tesseract::recognize(&img_path)?;
    // … integrate `text` into your final markdown output …
}

This pattern, supported by the hybrid pipeline logic in src/lib.rs (lines ~640-680), ensures optimal performance by avoiding unnecessary OCR operations on machine-readable pages.

Summary

  • pdf-inspector detects but does not execute OCR, functioning as a preprocessor for OCR workflows.
  • The library returns 1-indexed page numbers requiring OCR via pages_needing_ocr in the PdfProcessResult struct.
  • Granular reasoning via PageOcrReasons distinguishes between scanned images, garbled text, and vector graphics.
  • Hybrid pipelines enable selective OCR processing, routing only flagged pages to external engines like Tesseract or PaddleOCR.
  • Core detection logic resides in src/detector.rs, while public API definitions and OCR constants live in src/lib.rs.

Frequently Asked Questions

Does pdf-inspector include a built-in OCR engine?

No, pdf-inspector does not contain OCR capabilities. According to the firecrawl/pdf-inspector source code, the library exclusively analyzes PDF structure and content to identify which pages require OCR, then delegates actual text recognition to downstream engines such as Tesseract, PaddleOCR, or GLM-OCR.

What OCR reasons does pdf-inspector return?

The library returns machine-readable justifications through the PageOcrReasons struct, including scanned for image-based pages, suspected_garbled_text for encoding failures, and vector_text for graphical text representations. These reasons appear in the ocr_reasons_by_page field of the detection result.

How do I integrate pdf-inspector with Tesseract or PaddleOCR?

First run detect_pdf() to obtain the pages_needing_ocr list. Then iterate through those specific page numbers, render each to an image format (like PNG), and pass the image data to your chosen OCR engine. Finally, merge the extracted text back into your document processing pipeline while preserving the text from pages that didn't require OCR.

Can pdf-inspector distinguish between scanned images and garbled text?

Yes, the detection system differentiates between true scans and pages with encoding issues. The PageOcrReasons enum specifically marks pages as scanned versus suspected_garbled_text, allowing you to apply different preprocessing strategies: image enhancement for scans versus font substitution or encoding repair for garbled text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →