How the pdf-inspector Classifier Differentiates Mixed, Scanned, and Image-Based PDF Types
The pdf-inspector classifier distinguishes PDF types by analyzing text operators, image presence, template-image detection, and vector-text cues across sampled pages, then applies document-wide heuristics to categorize files as Mixed, Scanned, or ImageBased.
The pdf-inspector library, developed by Firecrawl, provides robust PDF type detection for Rust applications. Understanding how its classifier differentiates between Mixed, Scanned, and ImageBased PDF types helps developers choose the right extraction strategy—whether standard text extraction or full OCR is needed.
The Three PDF Types Explained
The PdfType enum is defined in src/detector.rs at lines 12–23:
pub enum PdfType {
TextBased,
Scanned,
ImageBased,
Mixed,
}
Each variant represents a distinct content pattern:
| Type | Description | Typical Source |
|---|---|---|
| TextBased | Native digital text with no images | Generated by word processors |
| Scanned | Pure raster images, no extractable text | Flatbed or sheet-fed scanners |
| ImageBased | Images with some text operators (garbled or vector outlines) | OCR failures, complex layouts |
| Mixed | Combination of text pages and image pages | Scanned documents with OCR'd sections |
Three-Phase Classification Process
The classifier operates in sequential phases, each implemented in src/detector.rs.
Phase 1: Page-Level Analysis
The analyze_page_content function inspects individual pages for four key signals:
- Text operators (
Tj/TJ) — indicates native PDF text - Image count — raster images present on the page
- Template-image flag — detects repeating background images (watermarks, letterheads)
- Vector-text flag — identifies text rendered as vector paths rather than text operators
Results are stored in a PageAnalysis struct for aggregation.
Phase 2: Document-Wide Heuristics
The detect_from_document function (lines 317–334) computes ratios across all sampled pages:
pages_with_text— pages containing extractable text operatorspages_with_images— pages with raster imagestext_ratio— proportion of text-bearing pagespages_with_vector_text— pages with vector-based text
Phase 3: OCR Target Identification
After classification, the system enumerates which specific pages require OCR. This reinforces the same signals used for type detection.
Core Decision Logic in detect_from_document
The classification algorithm follows a priority-ordered conditional structure:
// src/detector.rs – classification excerpt (lines 317–334)
if has_template_images && pages_with_text > 0 {
(PdfType::Mixed, …)
} else if text_ratio >= config.text_page_ratio_threshold {
(PdfType::TextBased, …)
} else if pages_with_text == 0 && (pages_with_images > 0 ||
pages_with_vector_text > 0) {
// No extractable text, but images or vector-text present
// Distinguish Scanned vs ImageBased here
if total_text_ops == 0 && pages_with_vector_text == 0 {
(PdfType::Scanned, 0.95) // pure image PDFs
} else {
(PdfType::ImageBased, 0.8) // images plus some text operators
}
} else if pages_with_text > 0 &&
(pages_with_images > 0 || pages_with_vector_text > 0) {
(PdfType::Mixed, 0.7) // both text and images present
}
// additional fallback branches...
Critical Distinction: Mixed vs. ImageBased PDF Types
The difference between Mixed and ImageBased hinges on two orthogonal signals:
Presence of Extractable Text
pages_with_text > 0indicates at least one page has readable text operators- When combined with
pages_with_images > 0orpages_with_vector_text > 0, the classifier selectsPdfType::Mixed
Absence of Extractable Text with Operator Artifacts
pages_with_text == 0means no page has usable text- If any text operators remain (
total_text_ops > 0) or vector-text exists (pages_with_vector_text > 0), the classifier choosesPdfType::ImageBasedrather than pureScanned
Key insight: ImageBased PDFs contain "ghost" text—operator sequences that don't render as readable text. This commonly occurs with failed OCR, corrupted encodings, or text converted to vector outlines.
Practical Code Examples
Rust API Usage
use pdf_inspector::detect_pdf_type;
fn main() -> Result<(), pdf_inspector::PdfError> {
// Detect PDF type from file path
let result = detect_pdf_type("document.pdf")?;
match result.pdf_type {
pdf_inspector::PdfType::Mixed => {
println!("Mixed PDF: apply OCR to image pages only");
}
pdf_inspector::PdfType::ImageBased => {
println!("Image-based PDF: full OCR required");
}
pdf_inspector::PdfType::Scanned => {
println!("Scanned PDF: full OCR with high confidence");
}
_ => {}
}
println!("Confidence: {:.2}", result.confidence);
Ok(())
}
CLI Usage
The detect-pdf binary in src/bin/detect_pdf.rs (lines 98–104) wraps the same detection routine:
# Detect and classify a PDF file
$ detect-pdf mixed-document.pdf
# → MIXED (text + images, OCR recommended on image pages)
$ detect-pdf corrupted-ocr.pdf
# → IMAGE-BASED (mostly images, OCR may help)
$ detect-pdf pure-scan.pdf
# → SCANNED (requires full OCR)
Configuration Thresholds
The classifier uses configurable thresholds from PdfDetectionConfig:
| Parameter | Default | Purpose |
|---|---|---|
text_page_ratio_threshold |
0.9 | Minimum text page ratio for TextBased |
| Sampling rate | Page-based | Pages analyzed for performance |
Adjust these in src/detector.rs to tune sensitivity for specific document populations.
Summary
-
Mixed PDF type is selected when extractable text exists on some pages AND images or vector-text exist on any pages—enabling selective OCR on image-only pages.
-
ImageBased PDF type applies when no extractable text exists BUT text operators or vector-text artifacts remain—indicating failed or corrupted text layers.
-
Scanned PDF type requires zero text operators and zero vector-text—pure raster images.
-
The distinction lives in
src/detector.rslines 317–334, with thePdfTypeenum defined at lines 12–23.
Frequently Asked Questions
What causes a PDF to be classified as ImageBased instead of Scanned?
ImageBased classification occurs when pages_with_text == 0 but total_text_ops > 0 or pages_with_vector_text > 0. This pattern indicates the PDF contains text operator sequences that don't form readable text—common with OCR failures, encoding corruption, or text converted to vector paths. Pure Scanned PDFs have no text operators whatsoever.
Why does the classifier detect template images separately?
Template images—repeating backgrounds like letterheads or watermarks—are flagged specially because they don't represent page-specific content. The has_template_images check (branch ① in the decision logic) combined with pages_with_text > 0 immediately triggers Mixed classification, avoiding misidentification of corporate documents as image-heavy.
Can I adjust the confidence thresholds for classification?
Yes. The detect_from_document function accepts a PdfDetectionConfig parameter with adjustable thresholds including text_page_ratio_threshold. Modify these values in src/detector.rs or pass custom configuration through the public API in src/lib.rs.
How does pdf-inspector handle mixed documents with very few text pages?
Documents where text_ratio falls below config.text_page_ratio_threshold but pages_with_text > 0 with accompanying images route to branch ④, yielding PdfType::Mixed with 0.7 confidence. For borderline cases, inspect the per_page_ocr_list in detection results to identify which specific pages need OCR processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →