How pdf-inspector Detects PDF Type: TextBased, Scanned, Mixed, and ImageBased Classification
pdf-inspector detects PDF type by sampling pages and analyzing content streams for text operators, images, fonts, and vector graphics, then aggregating per-page metrics to classify documents as TextBased, Scanned, Mixed, or ImageBased.
The open-source Rust library firecrawl/pdf-inspector provides robust PDF classification through a three-phase pipeline implemented in src/detector.rs. This system powers both the library API exposed in src/lib.rs and the detect-pdf CLI binary found in src/bin/detect_pdf.rs.
Overview of the Detection Pipeline
The detection process follows three distinct phases: page sampling, content analysis, and hierarchical classification. By default, the system uses a Sample(8) strategy that evenly distributes up to eight pages across the document, including the first and last pages. This approach balances speed with accuracy for large documents.
The core detection logic resides in src/detector.rs, where the DetectionConfig struct defines tunable thresholds and the PdfTypeResult struct returns the final classification with confidence scores and OCR recommendations.
Phase 1 – Page Sampling Strategy
pdf-inspector selects pages for analysis according to a configurable ScanStrategy enum. The distribute_pages function ensures the first and last pages are always included, with remaining indices spaced evenly throughout the document.
Available strategies include:
Sample(N)– Evenly distributes N pages (default: 8)Pages([...])– Analyzes specific page numbers onlyFull– Scans every page in the documentEarlyExit– Stops early if classification confidence reaches threshold
For a 100-page PDF using Sample(8), the sampled indices might be [1, 13, 25, 37, 49, 61, 73, 85, 100].
Phase 2 – Content Stream Analysis
For each sampled page, the analyze_page_content function parses the PDF content stream and populates a PageAnalysis struct with granular metrics.
Text Detection via Operators
The system counts Tj and TJ text-showing operators to determine text presence. A page qualifies as text-rich when text_operator_count exceeds the min_text_ops_per_page threshold (default: 3). The analysis also tracks unique_text_chars and unique_alphanum_chars to distinguish meaningful text from decorative glyphs.
Image and Template Detection
The detector identifies images by monitoring Do (Draw Object) operators. It specifically flags template images—single full-page background images that indicate scanned document backgrounds—through the has_template_image boolean.
Vector Text Identification
When a page contains vector graphics masquerading as text, the system detects this through path operation density. If path operations exceed 1,000 while unique alphanumerics remain below 30, the has_vector_text flag activates, indicating text rendered as curves rather than font glyphs.
Font Resolution and Decodability
After resolving font names to their underlying ObjectIds while respecting PDF resource inheritance, the analyzer checks three critical font properties:
has_identity_h_no_tounicode– Identity-H/V fonts lacking ToUnicode CMapshas_only_type3_fonts– All fonts are Type3 without Unicode mappingshas_decodable_text_fonts– At least one font supports Unicode extraction
These flags determine whether text is extractable or requires OCR.
Phase 3 – Classification Logic
The classification logic (implemented around lines 10,030–10,335 in src/detector.rs) follows a strict decision hierarchy:
- Template-image check – If
has_template_imagesis true andpages_with_text > 0, classify asMixed(OCR recommended) - Text ratio evaluation – If
text_ratio >= text_page_ratio_threshold(default: 0.6), classify asTextBased - Image-only detection – If no text exists but images or vector text are present, classify as
Scanned(orImageBasedif vector text detected) - Mixed content – If text coexists with images or vector graphics, classify as
Mixed - Fallback – Default to
TextBasedfor edge cases
A secondary "newspaper-layout" heuristic (lines 10,050–10,074) may upgrade ocr_recommended for dense, multi-column PDFs even when the primary type is TextBased.
Configuration and Thresholds
The DetectionConfig struct allows customization of detection behavior:
pub struct DetectionConfig {
pub strategy: ScanStrategy, // Page selection method
pub min_text_ops_per_page: u32, // Minimum Tj/TJ ops for text pages (default: 3)
pub text_page_ratio_threshold: f32 // Ratio for TextBased classification (default: 0.6)
}
Default settings favor speed while maintaining accuracy for typical document layouts.
Using the Detection API
Rust Library Usage
Import the detection functions from src/lib.rs to classify PDFs programmatically:
use pdf_inspector::{detect_pdf_type, DetectionConfig, ScanStrategy};
// Basic detection with defaults (Sample 8 pages)
let result = detect_pdf_type("document.pdf")?;
println!("Type: {:?}, Confidence: {}", result.pdf_type, result.confidence);
// Custom configuration – analyze only first 3 pages with stricter thresholds
let config = DetectionConfig {
strategy: ScanStrategy::Pages(vec![1, 2, 3]),
min_text_ops_per_page: 5,
text_page_ratio_threshold: 0.7,
};
let result = pdf_inspector::detect_pdf_type_with_config("document.pdf", config)?;
CLI Binary Usage
The detect-pdf binary in src/bin/detect_pdf.rs provides command-line access:
# Basic human-readable output
detect-pdf path/to/document.pdf
# JSON output for integration
detect-pdf --json path/to/document.pdf
# Custom page sampling
detect-pdf --strategy=pages=1,10,20 path/to/document.pdf
Accessing OCR Recommendations
The PdfTypeResult struct provides actionable OCR guidance:
if result.ocr_recommended {
for page in &result.pages_needing_ocr {
let reasons = &result.ocr_reasons_by_page[page];
println!("Page {} needs OCR: {:?}", page, reasons);
}
}
Scanned and ImageBased PDFs return all pages in pages_needing_ocr, while Mixed PDFs return only specific pages requiring OCR due to undecodable fonts or image content.
Summary
- pdf-inspector detects PDF type through a three-phase pipeline: page sampling, content stream analysis, and hierarchical classification.
- The system analyzes text operators (
Tj/TJ), image objects, vector graphics, and font properties insrc/detector.rs. - Default
Sample(8)strategy balances performance with accuracy by examining evenly distributed pages including first and last. - Classification depends on configurable thresholds:
min_text_ops_per_page(default 3) andtext_page_ratio_threshold(default 0.6). - Results include specific OCR recommendations via
pages_needing_ocrandocr_reasons_by_pagefor targeted processing.
Frequently Asked Questions
What sampling strategy does pdf-inspector use by default?
By default, pdf-inspector uses ScanStrategy::Sample(8), which selects up to eight evenly distributed pages across the document while always including the first and last pages. This strategy provides representative analysis without processing every page, making it suitable for large documents. You can override this with Full (all pages), Pages([...]) (specific pages), or EarlyExit (stop when confident).
How does pdf-inspector distinguish between TextBased and Scanned PDFs?
pdf-inspector distinguishes TextBased from Scanned PDFs by counting text-showing operators (Tj/TJ) and calculating the ratio of text-rich pages. If the text_ratio meets or exceeds the text_page_ratio_threshold (default 0.6), the document is TextBased. If no text operators are found but images or vector text exist, the document is classified as Scanned or ImageBased, triggering OCR recommendations for all pages.
What triggers the Mixed PDF classification?
The Mixed classification triggers in three scenarios: when template images coexist with text content, when the text ratio falls below the threshold but some text exists alongside images, or when specific pages contain undecodable fonts or vector text masquerading as content. Mixed documents return ocr_recommended: true with specific pages_needing_ocr identified for targeted OCR processing.
Can I customize the detection thresholds?
Yes, the DetectionConfig struct allows full customization of detection parameters. You can adjust min_text_ops_per_page to require more text operators before a page counts as text-rich, modify text_page_ratio_threshold to change the classification sensitivity, or specify exact pages via ScanStrategy::Pages. Pass your custom configuration to detect_pdf_type_with_config available in src/lib.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →