# How the pdf-inspector Classifier Differentiates Mixed, Scanned, and Image-Based PDF Types

> Learn how the pdf-inspector classifier differentiates Mixed, Scanned, and ImageBased PDF types by analyzing text operators, image presence, and vector-text cues. Discover the heuristics behind PDF categorization.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-06

---

**The pdf-inspector classifier distinguishes PDF types by analyzing text operators, image presence, template-image detection, and vector-text cues across sampled pages, then applies document-wide heuristics to categorize files as `Mixed`, `Scanned`, or `ImageBased`.**

The **pdf-inspector** library, developed by Firecrawl, provides robust PDF type detection for Rust applications. Understanding how its classifier differentiates between `Mixed`, `Scanned`, and `ImageBased` PDF types helps developers choose the right extraction strategy—whether standard text extraction or full OCR is needed.

## The Three PDF Types Explained

The `PdfType` enum is defined in **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** at lines 12–23:

```rust
pub enum PdfType {
    TextBased,
    Scanned,
    ImageBased,
    Mixed,
}

```

Each variant represents a distinct content pattern:

| Type | Description | Typical Source |
|------|-------------|--------------|
| **TextBased** | Native digital text with no images | Generated by word processors |
| **Scanned** | Pure raster images, no extractable text | Flatbed or sheet-fed scanners |
| **ImageBased** | Images with some text operators (garbled or vector outlines) | OCR failures, complex layouts |
| **Mixed** | Combination of text pages and image pages | Scanned documents with OCR'd sections |

## Three-Phase Classification Process

The classifier operates in sequential phases, each implemented in **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)**.

### Phase 1: Page-Level Analysis

The `analyze_page_content` function inspects individual pages for four key signals:

- **Text operators** (`Tj`/`TJ`) — indicates native PDF text
- **Image count** — raster images present on the page
- **Template-image flag** — detects repeating background images (watermarks, letterheads)
- **Vector-text flag** — identifies text rendered as vector paths rather than text operators

Results are stored in a `PageAnalysis` struct for aggregation.

### Phase 2: Document-Wide Heuristics

The `detect_from_document` function (lines 317–334) computes ratios across all sampled pages:

- `pages_with_text` — pages containing extractable text operators
- `pages_with_images` — pages with raster images
- `text_ratio` — proportion of text-bearing pages
- `pages_with_vector_text` — pages with vector-based text

### Phase 3: OCR Target Identification

After classification, the system enumerates which specific pages require OCR. This reinforces the same signals used for type detection.

## Core Decision Logic in `detect_from_document`

The classification algorithm follows a priority-ordered conditional structure:

```rust
// src/detector.rs – classification excerpt (lines 317–334)
if has_template_images && pages_with_text > 0 {
    (PdfType::Mixed, …)
} else if text_ratio >= config.text_page_ratio_threshold {
    (PdfType::TextBased, …)
} else if pages_with_text == 0 && (pages_with_images > 0 ||
                                   pages_with_vector_text > 0) {
    // No extractable text, but images or vector-text present
    // Distinguish Scanned vs ImageBased here
    if total_text_ops == 0 && pages_with_vector_text == 0 {
        (PdfType::Scanned, 0.95)      // pure image PDFs
    } else {
        (PdfType::ImageBased, 0.8)    // images plus some text operators
    }
} else if pages_with_text > 0 &&
          (pages_with_images > 0 || pages_with_vector_text > 0) {
    (PdfType::Mixed, 0.7)             // both text and images present
}
// additional fallback branches...

```

## Critical Distinction: Mixed vs. ImageBased PDF Types

The difference between `Mixed` and `ImageBased` hinges on **two orthogonal signals**:

### Presence of Extractable Text

- `pages_with_text > 0` indicates at least one page has readable text operators
- When combined with `pages_with_images > 0` or `pages_with_vector_text > 0`, the classifier selects **`PdfType::Mixed`**

### Absence of Extractable Text with Operator Artifacts

- `pages_with_text == 0` means no page has usable text
- If **any** text operators remain (`total_text_ops > 0`) or vector-text exists (`pages_with_vector_text > 0`), the classifier chooses **`PdfType::ImageBased`** rather than pure `Scanned`

**Key insight**: `ImageBased` PDFs contain "ghost" text—operator sequences that don't render as readable text. This commonly occurs with failed OCR, corrupted encodings, or text converted to vector outlines.

## Practical Code Examples

### Rust API Usage

```rust
use pdf_inspector::detect_pdf_type;

fn main() -> Result<(), pdf_inspector::PdfError> {
    // Detect PDF type from file path
    let result = detect_pdf_type("document.pdf")?;
    
    match result.pdf_type {
        pdf_inspector::PdfType::Mixed => {
            println!("Mixed PDF: apply OCR to image pages only");
        }
        pdf_inspector::PdfType::ImageBased => {
            println!("Image-based PDF: full OCR required");
        }
        pdf_inspector::PdfType::Scanned => {
            println!("Scanned PDF: full OCR with high confidence");
        }
        _ => {}
    }
    
    println!("Confidence: {:.2}", result.confidence);
    Ok(())
}

```

### CLI Usage

The `detect-pdf` binary in **[`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs)** (lines 98–104) wraps the same detection routine:

```bash

# Detect and classify a PDF file

$ detect-pdf mixed-document.pdf

# → MIXED (text + images, OCR recommended on image pages)

$ detect-pdf corrupted-ocr.pdf  

# → IMAGE-BASED (mostly images, OCR may help)

$ detect-pdf pure-scan.pdf

# → SCANNED (requires full OCR)

```

## Configuration Thresholds

The classifier uses configurable thresholds from `PdfDetectionConfig`:

| Parameter | Default | Purpose |
|-----------|---------|---------|
| `text_page_ratio_threshold` | 0.9 | Minimum text page ratio for `TextBased` |
| Sampling rate | Page-based | Pages analyzed for performance |

Adjust these in **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** to tune sensitivity for specific document populations.

## Summary

- **Mixed PDF type** is selected when **extractable text exists on some pages** AND **images or vector-text exist on any pages**—enabling selective OCR on image-only pages.

- **ImageBased PDF type** applies when **no extractable text exists** BUT **text operators or vector-text artifacts remain**—indicating failed or corrupted text layers.

- **Scanned PDF type** requires **zero text operators and zero vector-text**—pure raster images.

- The distinction lives in **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** lines 317–334, with the `PdfType` enum defined at lines 12–23.

## Frequently Asked Questions

### What causes a PDF to be classified as ImageBased instead of Scanned?

ImageBased classification occurs when `pages_with_text == 0` but `total_text_ops > 0` or `pages_with_vector_text > 0`. This pattern indicates the PDF contains text operator sequences that don't form readable text—common with OCR failures, encoding corruption, or text converted to vector paths. Pure Scanned PDFs have no text operators whatsoever.

### Why does the classifier detect template images separately?

Template images—repeating backgrounds like letterheads or watermarks—are flagged specially because they don't represent page-specific content. The `has_template_images` check (branch ① in the decision logic) combined with `pages_with_text > 0` immediately triggers Mixed classification, avoiding misidentification of corporate documents as image-heavy.

### Can I adjust the confidence thresholds for classification?

Yes. The `detect_from_document` function accepts a `PdfDetectionConfig` parameter with adjustable thresholds including `text_page_ratio_threshold`. Modify these values in **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** or pass custom configuration through the public API in **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)**.

### How does pdf-inspector handle mixed documents with very few text pages?

Documents where `text_ratio` falls below `config.text_page_ratio_threshold` but `pages_with_text > 0` with accompanying images route to branch ④, yielding `PdfType::Mixed` with 0.7 confidence. For borderline cases, inspect the `per_page_ocr_list` in detection results to identify which specific pages need OCR processing.