How pdf-inspector Handles Scanned Documents: Detection and OCR Classification

pdf-inspector classifies scanned PDFs by sampling page content for text operators and, upon detecting image-only data, flags the document for OCR while skipping the standard text-extraction pipeline.

The firecrawl/pdf-inspector repository provides a Rust-based engine for analyzing PDF content structure before extraction. When the library encounters documents containing only rasterized images without extractable text operators, it employs a specialized detection heuristic to prevent empty extraction attempts and signal the need for optical character recognition.

The PdfType Classification Enum

At the core of scanned document detection lies the PdfType enum defined in src/detector.rs. This classification system categorizes every analyzed PDF into one of four distinct types based on content composition:

pub enum PdfType {
    TextBased,   // extractable text (Tj/TJ operators)
    Scanned,     // images only, no text operators
    ImageBased,  // mostly images, little or no text
    Mixed,       // a blend of text and image-heavy pages
}

The Scanned variant specifically represents PDFs where pages contain only images with zero extractable text operators, while ImageBased covers documents with predominantly visual content but minimal vector text elements.

Detection Heuristics in src/detector.rs

The detection engine implemented in src/detector.rs analyzes a configurable sample of pages (defaulting to 8 sampled pages) to count specific content signals:

  • Text operators (Tj/TJ operators in PDF content streams)
  • Page images (rasterized content)
  • Vector text (text rendered as vector paths)

When the algorithm encounters zero text operators across the sampled pages but detects the presence of images or vector text, it triggers the scanned document classification logic:

else if pages_with_text == 0 && (pages_with_images > 0 || pages_with_vector_text > 0) {
    ocr_recommended = true;
    if total_text_ops == 0 && pages_with_vector_text == 0 {
        (PdfType::Scanned, 0.95)
    } else {
        (PdfType::ImageBased, 0.8)
    }
}

This logic assigns a confidence score of 0.95 for pure Scanned documents (no text operators and no vector text) and 0.8 for ImageBased documents. The system also sets ocr_recommended = true and marks all pages (range 1..=total_pages) as requiring OCR processing with the reason "scanned".

Pipeline Short-Circuit in src/lib.rs

Once classified, the main processing pipeline in src/lib.rs checks the document type before attempting text extraction. For scanned content, the engine performs an early exit to avoid futile extraction attempts:

if matches!(pdf_type, PdfType::Scanned | PdfType::ImageBased) {
    return Ok(PdfProcessResult {
        pdf_type,
        markdown: None,
        // … other fields …
    });
}

Because scanned PDFs contain no extractable text, this short-circuit returns a PdfProcessResult with markdown: None and includes the complete list of pages needing OCR processing.

CLI Tools and Usage Examples

The repository provides two command-line utilities that demonstrate this behavior:

Detecting Scanned PDFs:

detect-pdf my_scan.pdf

# → {"pdf_type":"Scanned","ocr_recommended":true,"pages_needing_ocr":[1,2,3],"confidence":0.95}

Attempting Markdown Extraction:

pdf2md my_scan.pdf --json

# → {"markdown":null,"pdf_type":"Scanned","pages_needing_ocr":[1,2,3],"ocr_recommended":true}

The pdf2md tool produces empty Markdown output for pure scans, while detect-pdf explicitly reports the classification and OCR recommendation.

Programmatic Detection with the Rust API

You can implement scanned PDF detection directly in Rust applications using the library's public API:

use pdf_inspector::{detect_pdf_type, PdfType};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = detect_pdf_type("my_scan.pdf")?;
    assert_eq!(result.pdf_type, PdfType::Scanned);
    assert!(result.ocr_recommended);
    Ok(())
}

This approach allows applications to route documents through OCR pipelines before attempting secondary text extraction.

Summary

  • pdf-inspector detects scanned documents in src/detector.rs by sampling pages for text operators and images.
  • Documents with zero text operators but containing images are classified as PdfType::Scanned with 95% confidence.
  • The processing pipeline in src/lib.rs short-circuits extraction for scanned PDFs, returning markdown: None.
  • Every page of a scanned document is flagged with ocr_recommended: true and listed in pages_needing_ocr.
  • The CLI tools detect-pdf and pdf2md surface this classification to users through JSON output and exit behavior.

Frequently Asked Questions

How does pdf-inspector distinguish between Scanned and ImageBased PDFs?

According to the source code in src/detector.rs, a document is classified as Scanned only when total_text_ops == 0 and pages_with_vector_text == 0, indicating pure rasterized images without any text elements. If vector text is present but page-based text operators are absent, the document receives the ImageBased classification with a lower confidence score of 0.8.

What output does pdf2md produce for a scanned document?

The pdf2md CLI tool returns a JSON result where the markdown field is set to null and ocr_recommended is true. The output includes a complete array of pages_needing_ocr spanning all pages in the document, signaling that the content requires optical character recognition before text extraction can succeed.

Can I adjust the number of pages sampled for scanned PDF detection?

Yes. While the default configuration samples 8 pages to balance accuracy and performance, the detection engine accepts a configurable parameter to adjust the sample size. Increasing the sample count improves detection accuracy for large documents with mixed content types, though this requires modifying the detection call parameters in the Rust API.

Does pdf-inspector perform OCR on scanned documents automatically?

No. The library detects and flags scanned documents for OCR but does not perform the actual optical character recognition. When PdfType::Scanned is detected, the pipeline returns early with an OCR recommendation, leaving the actual text recognition to external OCR engines or downstream processing pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →