# How pdf-inspector Handles Scanned Documents: Detection and OCR Classification

> Discover how pdf-inspector efficiently processes scanned documents. It detects image-only pages, flags them for OCR, and optimizes extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-04

---

**pdf-inspector classifies scanned PDFs by sampling page content for text operators and, upon detecting image-only data, flags the document for OCR while skipping the standard text-extraction pipeline.**

The firecrawl/pdf-inspector repository provides a Rust-based engine for analyzing PDF content structure before extraction. When the library encounters documents containing only rasterized images without extractable text operators, it employs a specialized detection heuristic to prevent empty extraction attempts and signal the need for optical character recognition.

## The PdfType Classification Enum

At the core of scanned document detection lies the **`PdfType`** enum defined in **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)**. This classification system categorizes every analyzed PDF into one of four distinct types based on content composition:

```rust
pub enum PdfType {
    TextBased,   // extractable text (Tj/TJ operators)
    Scanned,     // images only, no text operators
    ImageBased,  // mostly images, little or no text
    Mixed,       // a blend of text and image-heavy pages
}

```

The **`Scanned`** variant specifically represents PDFs where pages contain only images with zero extractable text operators, while **`ImageBased`** covers documents with predominantly visual content but minimal vector text elements.

## Detection Heuristics in src/detector.rs

The detection engine implemented in **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** analyzes a configurable sample of pages (defaulting to **8 sampled pages**) to count specific content signals:

- **Text operators** (`Tj`/`TJ` operators in PDF content streams)
- **Page images** (rasterized content)
- **Vector text** (text rendered as vector paths)

When the algorithm encounters zero text operators across the sampled pages but detects the presence of images or vector text, it triggers the scanned document classification logic:

```rust
else if pages_with_text == 0 && (pages_with_images > 0 || pages_with_vector_text > 0) {
    ocr_recommended = true;
    if total_text_ops == 0 && pages_with_vector_text == 0 {
        (PdfType::Scanned, 0.95)
    } else {
        (PdfType::ImageBased, 0.8)
    }
}

```

This logic assigns a **confidence score of 0.95** for pure `Scanned` documents (no text operators and no vector text) and **0.8** for `ImageBased` documents. The system also sets **`ocr_recommended = true`** and marks **all pages** (range `1..=total_pages`) as requiring OCR processing with the reason `"scanned"`.

## Pipeline Short-Circuit in src/lib.rs

Once classified, the main processing pipeline in **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** checks the document type before attempting text extraction. For scanned content, the engine performs an early exit to avoid futile extraction attempts:

```rust
if matches!(pdf_type, PdfType::Scanned | PdfType::ImageBased) {
    return Ok(PdfProcessResult {
        pdf_type,
        markdown: None,
        // … other fields …
    });
}

```

Because **scanned PDFs contain no extractable text**, this short-circuit returns a `PdfProcessResult` with `markdown: None` and includes the complete list of pages needing OCR processing.

## CLI Tools and Usage Examples

The repository provides two command-line utilities that demonstrate this behavior:

**Detecting Scanned PDFs:**

```bash
detect-pdf my_scan.pdf

# → {"pdf_type":"Scanned","ocr_recommended":true,"pages_needing_ocr":[1,2,3],"confidence":0.95}

```

**Attempting Markdown Extraction:**

```bash
pdf2md my_scan.pdf --json

# → {"markdown":null,"pdf_type":"Scanned","pages_needing_ocr":[1,2,3],"ocr_recommended":true}

```

The `pdf2md` tool produces empty Markdown output for pure scans, while `detect-pdf` explicitly reports the classification and OCR recommendation.

## Programmatic Detection with the Rust API

You can implement scanned PDF detection directly in Rust applications using the library's public API:

```rust
use pdf_inspector::{detect_pdf_type, PdfType};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = detect_pdf_type("my_scan.pdf")?;
    assert_eq!(result.pdf_type, PdfType::Scanned);
    assert!(result.ocr_recommended);
    Ok(())
}

```

This approach allows applications to route documents through OCR pipelines before attempting secondary text extraction.

## Summary

- **pdf-inspector** detects scanned documents in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) by sampling pages for text operators and images.
- Documents with **zero text operators** but containing images are classified as **`PdfType::Scanned`** with 95% confidence.
- The processing pipeline in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) **short-circuits extraction** for scanned PDFs, returning `markdown: None`.
- Every page of a scanned document is flagged with **`ocr_recommended: true`** and listed in `pages_needing_ocr`.
- The CLI tools `detect-pdf` and `pdf2md` surface this classification to users through JSON output and exit behavior.

## Frequently Asked Questions

### How does pdf-inspector distinguish between Scanned and ImageBased PDFs?

According to the source code in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), a document is classified as **`Scanned`** only when `total_text_ops == 0` and `pages_with_vector_text == 0`, indicating pure rasterized images without any text elements. If vector text is present but page-based text operators are absent, the document receives the **`ImageBased`** classification with a lower confidence score of 0.8.

### What output does pdf2md produce for a scanned document?

The `pdf2md` CLI tool returns a JSON result where the `markdown` field is set to `null` and `ocr_recommended` is `true`. The output includes a complete array of `pages_needing_ocr` spanning all pages in the document, signaling that the content requires optical character recognition before text extraction can succeed.

### Can I adjust the number of pages sampled for scanned PDF detection?

Yes. While the default configuration samples **8 pages** to balance accuracy and performance, the detection engine accepts a configurable parameter to adjust the sample size. Increasing the sample count improves detection accuracy for large documents with mixed content types, though this requires modifying the detection call parameters in the Rust API.

### Does pdf-inspector perform OCR on scanned documents automatically?

No. The library **detects and flags** scanned documents for OCR but does not perform the actual optical character recognition. When `PdfType::Scanned` is detected, the pipeline returns early with an OCR recommendation, leaving the actual text recognition to external OCR engines or downstream processing pipelines.