How the pdf-inspector Classifier Differentiates Mixed, Scanned, and Image-Based PDF Types

The pdf-inspector classifier distinguishes PDF types by analyzing text operators, image presence, template-image detection, and vector-text cues across sampled pages, then applies document-wide heuristics to categorize files as Mixed, Scanned, or ImageBased.

The pdf-inspector library, developed by Firecrawl, provides robust PDF type detection for Rust applications. Understanding how its classifier differentiates between Mixed, Scanned, and ImageBased PDF types helps developers choose the right extraction strategy—whether standard text extraction or full OCR is needed.

The Three PDF Types Explained

The PdfType enum is defined in src/detector.rs at lines 12–23:

pub enum PdfType {
    TextBased,
    Scanned,
    ImageBased,
    Mixed,
}

Each variant represents a distinct content pattern:

Type Description Typical Source
TextBased Native digital text with no images Generated by word processors
Scanned Pure raster images, no extractable text Flatbed or sheet-fed scanners
ImageBased Images with some text operators (garbled or vector outlines) OCR failures, complex layouts
Mixed Combination of text pages and image pages Scanned documents with OCR'd sections

Three-Phase Classification Process

The classifier operates in sequential phases, each implemented in src/detector.rs.

Phase 1: Page-Level Analysis

The analyze_page_content function inspects individual pages for four key signals:

  • Text operators (Tj/TJ) — indicates native PDF text
  • Image count — raster images present on the page
  • Template-image flag — detects repeating background images (watermarks, letterheads)
  • Vector-text flag — identifies text rendered as vector paths rather than text operators

Results are stored in a PageAnalysis struct for aggregation.

Phase 2: Document-Wide Heuristics

The detect_from_document function (lines 317–334) computes ratios across all sampled pages:

  • pages_with_text — pages containing extractable text operators
  • pages_with_images — pages with raster images
  • text_ratio — proportion of text-bearing pages
  • pages_with_vector_text — pages with vector-based text

Phase 3: OCR Target Identification

After classification, the system enumerates which specific pages require OCR. This reinforces the same signals used for type detection.

Core Decision Logic in detect_from_document

The classification algorithm follows a priority-ordered conditional structure:

// src/detector.rs – classification excerpt (lines 317–334)
if has_template_images && pages_with_text > 0 {
    (PdfType::Mixed, …)
} else if text_ratio >= config.text_page_ratio_threshold {
    (PdfType::TextBased, …)
} else if pages_with_text == 0 && (pages_with_images > 0 ||
                                   pages_with_vector_text > 0) {
    // No extractable text, but images or vector-text present
    // Distinguish Scanned vs ImageBased here
    if total_text_ops == 0 && pages_with_vector_text == 0 {
        (PdfType::Scanned, 0.95)      // pure image PDFs
    } else {
        (PdfType::ImageBased, 0.8)    // images plus some text operators
    }
} else if pages_with_text > 0 &&
          (pages_with_images > 0 || pages_with_vector_text > 0) {
    (PdfType::Mixed, 0.7)             // both text and images present
}
// additional fallback branches...

Critical Distinction: Mixed vs. ImageBased PDF Types

The difference between Mixed and ImageBased hinges on two orthogonal signals:

Presence of Extractable Text

  • pages_with_text > 0 indicates at least one page has readable text operators
  • When combined with pages_with_images > 0 or pages_with_vector_text > 0, the classifier selects PdfType::Mixed

Absence of Extractable Text with Operator Artifacts

  • pages_with_text == 0 means no page has usable text
  • If any text operators remain (total_text_ops > 0) or vector-text exists (pages_with_vector_text > 0), the classifier chooses PdfType::ImageBased rather than pure Scanned

Key insight: ImageBased PDFs contain "ghost" text—operator sequences that don't render as readable text. This commonly occurs with failed OCR, corrupted encodings, or text converted to vector outlines.

Practical Code Examples

Rust API Usage

use pdf_inspector::detect_pdf_type;

fn main() -> Result<(), pdf_inspector::PdfError> {
    // Detect PDF type from file path
    let result = detect_pdf_type("document.pdf")?;
    
    match result.pdf_type {
        pdf_inspector::PdfType::Mixed => {
            println!("Mixed PDF: apply OCR to image pages only");
        }
        pdf_inspector::PdfType::ImageBased => {
            println!("Image-based PDF: full OCR required");
        }
        pdf_inspector::PdfType::Scanned => {
            println!("Scanned PDF: full OCR with high confidence");
        }
        _ => {}
    }
    
    println!("Confidence: {:.2}", result.confidence);
    Ok(())
}

CLI Usage

The detect-pdf binary in src/bin/detect_pdf.rs (lines 98–104) wraps the same detection routine:


# Detect and classify a PDF file

$ detect-pdf mixed-document.pdf

# → MIXED (text + images, OCR recommended on image pages)

$ detect-pdf corrupted-ocr.pdf  

# → IMAGE-BASED (mostly images, OCR may help)

$ detect-pdf pure-scan.pdf

# → SCANNED (requires full OCR)

Configuration Thresholds

The classifier uses configurable thresholds from PdfDetectionConfig:

Parameter Default Purpose
text_page_ratio_threshold 0.9 Minimum text page ratio for TextBased
Sampling rate Page-based Pages analyzed for performance

Adjust these in src/detector.rs to tune sensitivity for specific document populations.

Summary

  • Mixed PDF type is selected when extractable text exists on some pages AND images or vector-text exist on any pages—enabling selective OCR on image-only pages.

  • ImageBased PDF type applies when no extractable text exists BUT text operators or vector-text artifacts remain—indicating failed or corrupted text layers.

  • Scanned PDF type requires zero text operators and zero vector-text—pure raster images.

  • The distinction lives in src/detector.rs lines 317–334, with the PdfType enum defined at lines 12–23.

Frequently Asked Questions

What causes a PDF to be classified as ImageBased instead of Scanned?

ImageBased classification occurs when pages_with_text == 0 but total_text_ops > 0 or pages_with_vector_text > 0. This pattern indicates the PDF contains text operator sequences that don't form readable text—common with OCR failures, encoding corruption, or text converted to vector paths. Pure Scanned PDFs have no text operators whatsoever.

Why does the classifier detect template images separately?

Template images—repeating backgrounds like letterheads or watermarks—are flagged specially because they don't represent page-specific content. The has_template_images check (branch ① in the decision logic) combined with pages_with_text > 0 immediately triggers Mixed classification, avoiding misidentification of corporate documents as image-heavy.

Can I adjust the confidence thresholds for classification?

Yes. The detect_from_document function accepts a PdfDetectionConfig parameter with adjustable thresholds including text_page_ratio_threshold. Modify these values in src/detector.rs or pass custom configuration through the public API in src/lib.rs.

How does pdf-inspector handle mixed documents with very few text pages?

Documents where text_ratio falls below config.text_page_ratio_threshold but pages_with_text > 0 with accompanying images route to branch ④, yielding PdfType::Mixed with 0.7 confidence. For borderline cases, inspect the per_page_ocr_list in detection results to identify which specific pages need OCR processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →