# Optimal ScanStrategy Configurations for PDF Detection Performance in pdf‑inspector

> Discover optimal ScanStrategy configurations for pdf-inspector detection performance. Choose Sample for balance, Full for accuracy, EarlyExit for speed, or Pages for custom control.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: performance
- Published: 2026-08-06

---

**The optimal ScanStrategy depends on your speed‑accuracy trade‑off: use `Sample(8)` for balanced performance, `Full` for maximum accuracy, `EarlyExit` for fastest text‑heavy pipelines, and `Pages` for custom page‑level control.**

pdf‑inspector classifies PDFs as **TextBased**, **Scanned**, **ImageBased**, or **Mixed** by detecting text operators (`Tj/TJ`) across selected pages. The **`ScanStrategy`** enum in [[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L25-L40) controls which pages get examined, directly impacting detection performance. Choosing the right ScanStrategy configuration ensures you get reliable PDF type detection without wasting compute on unnecessary page analysis.

## Understanding the Four ScanStrategy Variants

### EarlyExit

**`EarlyExit`** scans pages sequentially from page 1 and stops immediately upon finding any non‑text page.

- **Speed:** Fastest for text‑heavy PDFs
- **Risk:** Misclassifies documents with image‑only covers (e.g., annual reports, scanned books)
- **Best for:** Pipelines that route TextBased PDFs to fast extractors and treat everything else as "needs OCR"

In [[`detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L25-L40), the early‑exit logic lives inside the core detection loop, breaking as soon as `text_ops_count` falls below the threshold on any single page.

### Full

**`Full`** examines every page without early termination.

- **Speed:** Slowest, proportional to document length
- **Accuracy:** Highest—reliably distinguishes **Mixed** from **Scanned** by computing exact text‑page ratios
- **Best for:** Precision‑critical workflows, pre‑OCR classification, compliance scanning

The `Full` strategy calls the detection routine on all pages and aggregates results in [[`detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L80-L113), bypassing any short‑circuit logic.

### Sample(N)

**`Sample(N)`** selects **N** evenly‑distributed pages using the `distribute_pages` helper at [`detector.rs:80-88`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L80-L88) and [`detector.rs:80-113`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L80-L113).

- **Speed:** Very fast—bounded by N regardless of document size
- **Accuracy:** Robust for most real‑world PDFs; edge cases include highly heterogeneous documents where sampled pages miss text clusters
- **Distribution guarantee:** Always includes first page, last page, and spread intermediates

The default `DetectionConfig::default()` uses `Sample(8)`:

```rust
DetectionConfig {
    strategy: ScanStrategy::Sample(8),
    min_text_ops_per_page: 3,
    text_page_ratio_threshold: 0.6,
}

```

### Pages(Vec<u32>)

**`Pages(Vec<u32>)`** scans only explicitly listed 1‑indexed page numbers.

- **Speed:** Depends on list length
- **Accuracy:** Deterministic—what you specify is what gets checked
- **Best for:** Custom workflows where callers already know representative pages (e.g., page 1 for cover, page 2 for content)

## Recommended ScanStrategy Configurations by Workload

| Goal | Configuration | When to use |
|------|-------------|-------------|
| Balanced speed‑accuracy | `ScanStrategy::Sample(8)` (default) | General production workloads |
| Maximum speed on large PDFs | `ScanStrategy::Sample(4)` | Documents >500 pages, archival batches |
| Definitive classification | `ScanStrategy::Full` | OCR routing, compliance, forensic analysis |
| Fast text‑biased routing | `ScanStrategy::EarlyExit` | Known‑clean document streams, search‑indexing pipelines |
| Custom page inspection | `ScanStrategy::Pages(vec![1, 5, 10, last])` | Template‑driven processing, user‑specified sampling |

## Implementing Custom ScanStrategy Configurations

Override defaults through the public API in [[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):

```rust
use pdf_inspector::detector::{DetectionConfig, ScanStrategy};

// Fast sampling for large‑volume processing
let fast_cfg = DetectionConfig {
    strategy: ScanStrategy::Sample(6),
    ..Default::default()
};

// Exhaustive scan for critical documents
let accurate_cfg = DetectionConfig {
    strategy: ScanStrategy::Full,
    ..Default::default()
};

// Apply configuration
let result = pdf_inspector::detect_pdf_type_with_config("document.pdf", fast_cfg)?;

```

## CLI Configuration Without Code Changes

The binary at [[`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) exposes ScanStrategy through the `--strategy` flag:

```bash

# Default sampling

detect-pdf --strategy sample:8 input.pdf

# Full document scan

detect-pdf --strategy full input.pdf

# Early exit mode

detect-pdf --strategy early-exit input.pdf

# Specific pages

detect-pdf --strategy pages:1,3,5,10 input.pdf

```

## Key Implementation Files

| File | Relevance to ScanStrategy configuration |
|------|----------------------------------------|
| [[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | Defines `ScanStrategy` enum, `distribute_pages` helper, and `detect_from_document` algorithm |
| [[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API: `detect_pdf_type`, `detect_pdf_type_with_config` |
| [[`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) | CLI argument parsing for `--strategy` |

## Summary

- **`Sample(N)`** with N=8 provides the best default ScanStrategy configuration for most production workloads.
- **`Full`** eliminates sampling error when distinguishing Mixed from Scanned PDFs is critical.
- **`EarlyExit`** delivers fastest detection for text‑dominated pipelines but risks misclassification on cover‑page anomalies.
- **`Pages`** offers deterministic control for custom integration scenarios.
- Tune `min_text_ops_per_page` and `text_page_ratio_threshold` alongside ScanStrategy for complete detection performance optimization.

## Frequently Asked Questions

### How does `Sample(N)` choose which pages to scan?

The `distribute_pages` function in [[`detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L80-113) ensures first page, last page, and evenly‑spaced intermediate pages are selected. For N=8 on a 100‑page document, pages 1, 15, 29, 43, 57, 71, 85, and 100 get examined.

### When should I avoid `EarlyExit` strategy?

Avoid `EarlyExit` when processing documents with image‑only front matter—annual reports, scanned books, or marketing PDFs with decorative covers. These trigger false‑positive "non‑text" classification on page 1 despite text‑heavy subsequent pages.

### Can I combine `Sample` with custom page lists?

Not directly—the `ScanStrategy` enum variants are mutually exclusive per [`detector.rs#L25-40`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs#L25-L40). Implement custom hybrid logic by using `Pages` and computing your own distribution, or run multiple detection passes with different strategies.

### What performance difference exists between `Sample(4)` and `Sample(8)`?

Both execute in O(N) time where N is the sample count, not document length. `Sample(4)` halves page I/O and operator parsing versus `Sample(8)`. For a 500‑page PDF, expect ~2× speedup with marginal accuracy reduction on heterogeneous documents.