ScanStrategy Options in pdf‑inspector: When to Use EarlyExit, Full, Sample, and Pages

The ScanStrategy enum in firecrawl/pdf‑inspector controls which pages are analyzed during PDF classification, offering four variants—EarlyExit, Full, Sample(N), and Pages(Vec)—that trade speed against accuracy depending on your pipeline requirements.

The pdf‑inspector crate classifies PDF documents by sampling page content to determine if they are text-based, mixed, or scanned. The ScanStrategy options defined in src/detector.rs let you customize this sampling behavior to optimize for speed, accuracy, or specific page targeting.

ScanStrategy Variants

EarlyExit

Scans pages sequentially and aborts immediately when encountering a page without text operators (Tj/TJ). This is the fastest option for confirming text-based documents.

Use EarlyExit in fast-path pipelines where you only need to confirm a document is TextBased. It quickly routes pure-text PDFs to fast extraction engines while sending any document containing a non-text page to OCR paths.

Full

Inspects every page in the document regardless of content, building a complete picture of text distribution.

Choose Full when you need reliable distinction between Mixed and Scanned PDFs. This is essential when downstream processing must decide whether OCR is mandatory or optional, as implemented in detect_from_document within src/detector.rs.

Sample(N)

Selects N evenly distributed pages across the document (first, last, and middle pages) rather than scanning sequentially.

Ideal for huge PDFs with hundreds of pages where full scanning would be too slow. The trade-off is slight precision loss, but sampled pages usually provide sufficient signal for classification.

Pages(Vec)

Accepts an explicit vector of 1-indexed page numbers to examine, ignoring all other pages.

Use this when you already know which pages likely contain text—such as when you have a table of contents—or when limiting scanning to specific pages for diagnostic purposes.

How Detection Applies ScanStrategy

During detection, the detect_from_document function in src/detector.rs (lines 181-199) converts your chosen ScanStrategy into two concrete values:

  1. sample_indices — The exact page numbers that will be examined.
  2. allow_early_exit — A boolean flag set to true for EarlyExit and false for other strategies, telling the scanner whether it may stop early.

The detector walks these indices, counts text operators, and returns a PdfTypeResult containing confidence scores, OCR recommendations, and page-level reasons.

Code Examples

Default Sampling (Sample 8)

use pdf_inspector::detect_pdf_type;

let result = detect_pdf_type("report.pdf")?;
println!("PDF type: {:?}, confidence: {}", result.pdf_type, result.confidence);

The default configuration uses Sample(8), which inspects eight evenly distributed pages. This is defined in the Default implementation for DetectionConfig in src/detector.rs (lines 80-88).

Fast Routing with EarlyExit

use pdf_inspector::{DetectionConfig, ScanStrategy, detect_pdf_type_with_config};

let config = DetectionConfig {
    strategy: ScanStrategy::EarlyExit,
    ..Default::default()
};

let res = detect_pdf_type_with_config("large_manual.pdf", config)?;
println!("{:?}", res);

The detector stops at the first page lacking text operators, making the check very quick.

Precise Mixed vs Scanned Classification

let config = DetectionConfig {
    strategy: ScanStrategy::Full,
    ..Default::default()
};

let res = pdf_inspector::detect_pdf_type_with_config("mixed_content.pdf", config)?;
println!("Pages with text: {}", res.pages_with_text);

All pages are examined, guaranteeing the most accurate classification.

Custom Sampling for Large Documents

let config = DetectionConfig {
    strategy: ScanStrategy::Sample(4),
    ..Default::default()
};

let res = pdf_inspector::detect_pdf_type_with_config("huge_book.pdf", config)?;
println!("Sampled pages: {}", res.pages_sampled);

Useful for very large PDFs where speed is critical.

Targeting Specific Pages

let config = DetectionConfig {
    strategy: ScanStrategy::Pages(vec![1, 2, 10]),
    ..Default::default()
};

let res = pdf_inspector::detect_pdf_type_with_config("custom_selection.pdf", config)?;
println!("{:?}", res);

Only pages 1, 2, and 10 are inspected.

Key Source Files

  • src/detector.rs — Defines the ScanStrategy enum, DetectionConfig struct, and the core detect_from_document algorithm.
  • src/lib.rs — Exposes the public API including detect_pdf_type and detect_pdf_type_with_config.
  • src/bin/detect_pdf.rs — CLI example demonstrating strategy toggling via command-line flags.

Summary

  • EarlyExit stops at the first non-text page for maximum speed when confirming text-based documents.
  • Full scans every page to accurately distinguish between Mixed and Scanned PDFs.
  • Sample(N) checks N evenly distributed pages, balancing speed and accuracy for large documents.
  • Pages(Vec) targets specific 1-indexed pages when you know exactly where to look.
  • The strategy is passed through DetectionConfig to detect_pdf_type_with_config, affecting the sample_indices and allow_early_exit values in src/detector.rs.

Frequently Asked Questions

What is the default ScanStrategy in pdf‑inspector?

The default strategy is Sample(8), which inspects eight evenly distributed pages. This provides a reasonable balance between classification accuracy and performance without requiring explicit configuration, as implemented in the Default trait implementation for DetectionConfig in src/detector.rs.

How does EarlyExit improve performance?

EarlyExit improves performance by aborting the scan as soon as it encounters a page lacking text operators (Tj/TJ). For text-based PDFs where all pages contain extractable text, this means scanning only a few pages instead of the entire document, significantly reducing processing time in fast-path pipelines.

Can I combine multiple ScanStrategy options?

No, you must select a single variant when constructing DetectionConfig. However, you can achieve similar results by using Pages(Vec<u32>) to specify exactly which pages to sample, effectively combining the specificity of manual selection with the efficiency of sampling.

Where is the ScanStrategy enum defined?

The ScanStrategy enum is defined in src/detector.rs at lines 25-40 in the firecrawl/pdf‑inspector repository. This file also contains the logic that interprets each variant into sample_indices and allow_early_exit parameters used by the detection algorithm.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →