Optimal ScanStrategy Configurations for PDF Detection Performance in pdf‑inspector
The optimal ScanStrategy depends on your speed‑accuracy trade‑off: use Sample(8) for balanced performance, Full for maximum accuracy, EarlyExit for fastest text‑heavy pipelines, and Pages for custom page‑level control.
pdf‑inspector classifies PDFs as TextBased, Scanned, ImageBased, or Mixed by detecting text operators (Tj/TJ) across selected pages. The ScanStrategy enum in [src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L25-L40) controls which pages get examined, directly impacting detection performance. Choosing the right ScanStrategy configuration ensures you get reliable PDF type detection without wasting compute on unnecessary page analysis.
Understanding the Four ScanStrategy Variants
EarlyExit
EarlyExit scans pages sequentially from page 1 and stops immediately upon finding any non‑text page.
- Speed: Fastest for text‑heavy PDFs
- Risk: Misclassifies documents with image‑only covers (e.g., annual reports, scanned books)
- Best for: Pipelines that route TextBased PDFs to fast extractors and treat everything else as "needs OCR"
In [detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L25-L40), the early‑exit logic lives inside the core detection loop, breaking as soon as text_ops_count falls below the threshold on any single page.
Full
Full examines every page without early termination.
- Speed: Slowest, proportional to document length
- Accuracy: Highest—reliably distinguishes Mixed from Scanned by computing exact text‑page ratios
- Best for: Precision‑critical workflows, pre‑OCR classification, compliance scanning
The Full strategy calls the detection routine on all pages and aggregates results in [detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L80-L113), bypassing any short‑circuit logic.
Sample(N)
Sample(N) selects N evenly‑distributed pages using the distribute_pages helper at detector.rs:80-88 and detector.rs:80-113.
- Speed: Very fast—bounded by N regardless of document size
- Accuracy: Robust for most real‑world PDFs; edge cases include highly heterogeneous documents where sampled pages miss text clusters
- Distribution guarantee: Always includes first page, last page, and spread intermediates
The default DetectionConfig::default() uses Sample(8):
DetectionConfig {
strategy: ScanStrategy::Sample(8),
min_text_ops_per_page: 3,
text_page_ratio_threshold: 0.6,
}
Pages(Vec)
Pages(Vec<u32>) scans only explicitly listed 1‑indexed page numbers.
- Speed: Depends on list length
- Accuracy: Deterministic—what you specify is what gets checked
- Best for: Custom workflows where callers already know representative pages (e.g., page 1 for cover, page 2 for content)
Recommended ScanStrategy Configurations by Workload
| Goal | Configuration | When to use |
|---|---|---|
| Balanced speed‑accuracy | ScanStrategy::Sample(8) (default) |
General production workloads |
| Maximum speed on large PDFs | ScanStrategy::Sample(4) |
Documents >500 pages, archival batches |
| Definitive classification | ScanStrategy::Full |
OCR routing, compliance, forensic analysis |
| Fast text‑biased routing | ScanStrategy::EarlyExit |
Known‑clean document streams, search‑indexing pipelines |
| Custom page inspection | ScanStrategy::Pages(vec![1, 5, 10, last]) |
Template‑driven processing, user‑specified sampling |
Implementing Custom ScanStrategy Configurations
Override defaults through the public API in [src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):
use pdf_inspector::detector::{DetectionConfig, ScanStrategy};
// Fast sampling for large‑volume processing
let fast_cfg = DetectionConfig {
strategy: ScanStrategy::Sample(6),
..Default::default()
};
// Exhaustive scan for critical documents
let accurate_cfg = DetectionConfig {
strategy: ScanStrategy::Full,
..Default::default()
};
// Apply configuration
let result = pdf_inspector::detect_pdf_type_with_config("document.pdf", fast_cfg)?;
CLI Configuration Without Code Changes
The binary at [src/bin/detect_pdf.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) exposes ScanStrategy through the --strategy flag:
# Default sampling
detect-pdf --strategy sample:8 input.pdf
# Full document scan
detect-pdf --strategy full input.pdf
# Early exit mode
detect-pdf --strategy early-exit input.pdf
# Specific pages
detect-pdf --strategy pages:1,3,5,10 input.pdf
Key Implementation Files
| File | Relevance to ScanStrategy configuration |
|---|---|
[src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) |
Defines ScanStrategy enum, distribute_pages helper, and detect_from_document algorithm |
[src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) |
Public API: detect_pdf_type, detect_pdf_type_with_config |
[src/bin/detect_pdf.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) |
CLI argument parsing for --strategy |
Summary
Sample(N)with N=8 provides the best default ScanStrategy configuration for most production workloads.Fulleliminates sampling error when distinguishing Mixed from Scanned PDFs is critical.EarlyExitdelivers fastest detection for text‑dominated pipelines but risks misclassification on cover‑page anomalies.Pagesoffers deterministic control for custom integration scenarios.- Tune
min_text_ops_per_pageandtext_page_ratio_thresholdalongside ScanStrategy for complete detection performance optimization.
Frequently Asked Questions
How does Sample(N) choose which pages to scan?
The distribute_pages function in [detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L80-113) ensures first page, last page, and evenly‑spaced intermediate pages are selected. For N=8 on a 100‑page document, pages 1, 15, 29, 43, 57, 71, 85, and 100 get examined.
When should I avoid EarlyExit strategy?
Avoid EarlyExit when processing documents with image‑only front matter—annual reports, scanned books, or marketing PDFs with decorative covers. These trigger false‑positive "non‑text" classification on page 1 despite text‑heavy subsequent pages.
Can I combine Sample with custom page lists?
Not directly—the ScanStrategy enum variants are mutually exclusive per detector.rs#L25-40. Implement custom hybrid logic by using Pages and computing your own distribution, or run multiple detection passes with different strategies.
What performance difference exists between Sample(4) and Sample(8)?
Both execute in O(N) time where N is the sample count, not document length. Sample(4) halves page I/O and operator parsing versus Sample(8). For a 500‑page PDF, expect ~2× speedup with marginal accuracy reduction on heterogeneous documents.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →