Optimal ScanStrategy Configurations for PDF Detection Performance in pdf‑inspector

The optimal ScanStrategy depends on your speed‑accuracy trade‑off: use Sample(8) for balanced performance, Full for maximum accuracy, EarlyExit for fastest text‑heavy pipelines, and Pages for custom page‑level control.

pdf‑inspector classifies PDFs as TextBased, Scanned, ImageBased, or Mixed by detecting text operators (Tj/TJ) across selected pages. The ScanStrategy enum in [src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L25-L40) controls which pages get examined, directly impacting detection performance. Choosing the right ScanStrategy configuration ensures you get reliable PDF type detection without wasting compute on unnecessary page analysis.

Understanding the Four ScanStrategy Variants

EarlyExit

EarlyExit scans pages sequentially from page 1 and stops immediately upon finding any non‑text page.

  • Speed: Fastest for text‑heavy PDFs
  • Risk: Misclassifies documents with image‑only covers (e.g., annual reports, scanned books)
  • Best for: Pipelines that route TextBased PDFs to fast extractors and treat everything else as "needs OCR"

In [detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L25-L40), the early‑exit logic lives inside the core detection loop, breaking as soon as text_ops_count falls below the threshold on any single page.

Full

Full examines every page without early termination.

  • Speed: Slowest, proportional to document length
  • Accuracy: Highest—reliably distinguishes Mixed from Scanned by computing exact text‑page ratios
  • Best for: Precision‑critical workflows, pre‑OCR classification, compliance scanning

The Full strategy calls the detection routine on all pages and aggregates results in [detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L80-L113), bypassing any short‑circuit logic.

Sample(N)

Sample(N) selects N evenly‑distributed pages using the distribute_pages helper at detector.rs:80-88 and detector.rs:80-113.

  • Speed: Very fast—bounded by N regardless of document size
  • Accuracy: Robust for most real‑world PDFs; edge cases include highly heterogeneous documents where sampled pages miss text clusters
  • Distribution guarantee: Always includes first page, last page, and spread intermediates

The default DetectionConfig::default() uses Sample(8):

DetectionConfig {
    strategy: ScanStrategy::Sample(8),
    min_text_ops_per_page: 3,
    text_page_ratio_threshold: 0.6,
}

Pages(Vec)

Pages(Vec<u32>) scans only explicitly listed 1‑indexed page numbers.

  • Speed: Depends on list length
  • Accuracy: Deterministic—what you specify is what gets checked
  • Best for: Custom workflows where callers already know representative pages (e.g., page 1 for cover, page 2 for content)
Goal Configuration When to use
Balanced speed‑accuracy ScanStrategy::Sample(8) (default) General production workloads
Maximum speed on large PDFs ScanStrategy::Sample(4) Documents >500 pages, archival batches
Definitive classification ScanStrategy::Full OCR routing, compliance, forensic analysis
Fast text‑biased routing ScanStrategy::EarlyExit Known‑clean document streams, search‑indexing pipelines
Custom page inspection ScanStrategy::Pages(vec![1, 5, 10, last]) Template‑driven processing, user‑specified sampling

Implementing Custom ScanStrategy Configurations

Override defaults through the public API in [src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):

use pdf_inspector::detector::{DetectionConfig, ScanStrategy};

// Fast sampling for large‑volume processing
let fast_cfg = DetectionConfig {
    strategy: ScanStrategy::Sample(6),
    ..Default::default()
};

// Exhaustive scan for critical documents
let accurate_cfg = DetectionConfig {
    strategy: ScanStrategy::Full,
    ..Default::default()
};

// Apply configuration
let result = pdf_inspector::detect_pdf_type_with_config("document.pdf", fast_cfg)?;

CLI Configuration Without Code Changes

The binary at [src/bin/detect_pdf.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) exposes ScanStrategy through the --strategy flag:


# Default sampling

detect-pdf --strategy sample:8 input.pdf

# Full document scan

detect-pdf --strategy full input.pdf

# Early exit mode

detect-pdf --strategy early-exit input.pdf

# Specific pages

detect-pdf --strategy pages:1,3,5,10 input.pdf

Key Implementation Files

File Relevance to ScanStrategy configuration
[src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) Defines ScanStrategy enum, distribute_pages helper, and detect_from_document algorithm
[src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) Public API: detect_pdf_type, detect_pdf_type_with_config
[src/bin/detect_pdf.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) CLI argument parsing for --strategy

Summary

  • Sample(N) with N=8 provides the best default ScanStrategy configuration for most production workloads.
  • Full eliminates sampling error when distinguishing Mixed from Scanned PDFs is critical.
  • EarlyExit delivers fastest detection for text‑dominated pipelines but risks misclassification on cover‑page anomalies.
  • Pages offers deterministic control for custom integration scenarios.
  • Tune min_text_ops_per_page and text_page_ratio_threshold alongside ScanStrategy for complete detection performance optimization.

Frequently Asked Questions

How does Sample(N) choose which pages to scan?

The distribute_pages function in [detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L80-113) ensures first page, last page, and evenly‑spaced intermediate pages are selected. For N=8 on a 100‑page document, pages 1, 15, 29, 43, 57, 71, 85, and 100 get examined.

When should I avoid EarlyExit strategy?

Avoid EarlyExit when processing documents with image‑only front matter—annual reports, scanned books, or marketing PDFs with decorative covers. These trigger false‑positive "non‑text" classification on page 1 despite text‑heavy subsequent pages.

Can I combine Sample with custom page lists?

Not directly—the ScanStrategy enum variants are mutually exclusive per detector.rs#L25-40. Implement custom hybrid logic by using Pages and computing your own distribution, or run multiple detection passes with different strategies.

What performance difference exists between Sample(4) and Sample(8)?

Both execute in O(N) time where N is the sample count, not document length. Sample(4) halves page I/O and operator parsing versus Sample(8). For a 500‑page PDF, expect ~2× speedup with marginal accuracy reduction on heterogeneous documents.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →