ScanStrategy Options in pdf‑inspector: When to Use EarlyExit, Full, Sample, and Pages
The ScanStrategy enum in firecrawl/pdf‑inspector controls which pages are analyzed during PDF classification, offering four variants—EarlyExit, Full, Sample(N), and Pages(Vec)—that trade speed against accuracy depending on your pipeline requirements.
The pdf‑inspector crate classifies PDF documents by sampling page content to determine if they are text-based, mixed, or scanned. The ScanStrategy options defined in src/detector.rs let you customize this sampling behavior to optimize for speed, accuracy, or specific page targeting.
ScanStrategy Variants
EarlyExit
Scans pages sequentially and aborts immediately when encountering a page without text operators (Tj/TJ). This is the fastest option for confirming text-based documents.
Use EarlyExit in fast-path pipelines where you only need to confirm a document is TextBased. It quickly routes pure-text PDFs to fast extraction engines while sending any document containing a non-text page to OCR paths.
Full
Inspects every page in the document regardless of content, building a complete picture of text distribution.
Choose Full when you need reliable distinction between Mixed and Scanned PDFs. This is essential when downstream processing must decide whether OCR is mandatory or optional, as implemented in detect_from_document within src/detector.rs.
Sample(N)
Selects N evenly distributed pages across the document (first, last, and middle pages) rather than scanning sequentially.
Ideal for huge PDFs with hundreds of pages where full scanning would be too slow. The trade-off is slight precision loss, but sampled pages usually provide sufficient signal for classification.
Pages(Vec)
Accepts an explicit vector of 1-indexed page numbers to examine, ignoring all other pages.
Use this when you already know which pages likely contain text—such as when you have a table of contents—or when limiting scanning to specific pages for diagnostic purposes.
How Detection Applies ScanStrategy
During detection, the detect_from_document function in src/detector.rs (lines 181-199) converts your chosen ScanStrategy into two concrete values:
sample_indices— The exact page numbers that will be examined.allow_early_exit— A boolean flag set totrueforEarlyExitandfalsefor other strategies, telling the scanner whether it may stop early.
The detector walks these indices, counts text operators, and returns a PdfTypeResult containing confidence scores, OCR recommendations, and page-level reasons.
Code Examples
Default Sampling (Sample 8)
use pdf_inspector::detect_pdf_type;
let result = detect_pdf_type("report.pdf")?;
println!("PDF type: {:?}, confidence: {}", result.pdf_type, result.confidence);
The default configuration uses Sample(8), which inspects eight evenly distributed pages. This is defined in the Default implementation for DetectionConfig in src/detector.rs (lines 80-88).
Fast Routing with EarlyExit
use pdf_inspector::{DetectionConfig, ScanStrategy, detect_pdf_type_with_config};
let config = DetectionConfig {
strategy: ScanStrategy::EarlyExit,
..Default::default()
};
let res = detect_pdf_type_with_config("large_manual.pdf", config)?;
println!("{:?}", res);
The detector stops at the first page lacking text operators, making the check very quick.
Precise Mixed vs Scanned Classification
let config = DetectionConfig {
strategy: ScanStrategy::Full,
..Default::default()
};
let res = pdf_inspector::detect_pdf_type_with_config("mixed_content.pdf", config)?;
println!("Pages with text: {}", res.pages_with_text);
All pages are examined, guaranteeing the most accurate classification.
Custom Sampling for Large Documents
let config = DetectionConfig {
strategy: ScanStrategy::Sample(4),
..Default::default()
};
let res = pdf_inspector::detect_pdf_type_with_config("huge_book.pdf", config)?;
println!("Sampled pages: {}", res.pages_sampled);
Useful for very large PDFs where speed is critical.
Targeting Specific Pages
let config = DetectionConfig {
strategy: ScanStrategy::Pages(vec![1, 2, 10]),
..Default::default()
};
let res = pdf_inspector::detect_pdf_type_with_config("custom_selection.pdf", config)?;
println!("{:?}", res);
Only pages 1, 2, and 10 are inspected.
Key Source Files
src/detector.rs— Defines theScanStrategyenum,DetectionConfigstruct, and the coredetect_from_documentalgorithm.src/lib.rs— Exposes the public API includingdetect_pdf_typeanddetect_pdf_type_with_config.src/bin/detect_pdf.rs— CLI example demonstrating strategy toggling via command-line flags.
Summary
- EarlyExit stops at the first non-text page for maximum speed when confirming text-based documents.
- Full scans every page to accurately distinguish between Mixed and Scanned PDFs.
- Sample(N) checks N evenly distributed pages, balancing speed and accuracy for large documents.
- Pages(Vec) targets specific 1-indexed pages when you know exactly where to look.
- The strategy is passed through
DetectionConfigtodetect_pdf_type_with_config, affecting thesample_indicesandallow_early_exitvalues insrc/detector.rs.
Frequently Asked Questions
What is the default ScanStrategy in pdf‑inspector?
The default strategy is Sample(8), which inspects eight evenly distributed pages. This provides a reasonable balance between classification accuracy and performance without requiring explicit configuration, as implemented in the Default trait implementation for DetectionConfig in src/detector.rs.
How does EarlyExit improve performance?
EarlyExit improves performance by aborting the scan as soon as it encounters a page lacking text operators (Tj/TJ). For text-based PDFs where all pages contain extractable text, this means scanning only a few pages instead of the entire document, significantly reducing processing time in fast-path pipelines.
Can I combine multiple ScanStrategy options?
No, you must select a single variant when constructing DetectionConfig. However, you can achieve similar results by using Pages(Vec<u32>) to specify exactly which pages to sample, effectively combining the specificity of manual selection with the efficiency of sampling.
Where is the ScanStrategy enum defined?
The ScanStrategy enum is defined in src/detector.rs at lines 25-40 in the firecrawl/pdf‑inspector repository. This file also contains the logic that interprets each variant into sample_indices and allow_early_exit parameters used by the detection algorithm.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →