# ScanStrategy Options in pdf‑inspector: When to Use EarlyExit, Full, Sample, and Pages

> Explore pdf-inspector ScanStrategy options like EarlyExit, Full, Sample, and Pages. Learn when to use each strategy to optimize PDF analysis speed and accuracy for your project.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-08

---

**The `ScanStrategy` enum in firecrawl/pdf‑inspector controls which pages are analyzed during PDF classification, offering four variants—EarlyExit, Full, Sample(N), and Pages(Vec<u32>)—that trade speed against accuracy depending on your pipeline requirements.**

The `pdf‑inspector` crate classifies PDF documents by sampling page content to determine if they are text-based, mixed, or scanned. The **ScanStrategy** options defined in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) let you customize this sampling behavior to optimize for speed, accuracy, or specific page targeting.

## ScanStrategy Variants

### EarlyExit

Scans pages sequentially and aborts immediately when encountering a page without text operators (`Tj`/`TJ`). This is the fastest option for confirming text-based documents.

Use `EarlyExit` in fast-path pipelines where you only need to confirm a document is **TextBased**. It quickly routes pure-text PDFs to fast extraction engines while sending any document containing a non-text page to OCR paths.

### Full

Inspects every page in the document regardless of content, building a complete picture of text distribution.

Choose `Full` when you need reliable distinction between **Mixed** and **Scanned** PDFs. This is essential when downstream processing must decide whether OCR is mandatory or optional, as implemented in `detect_from_document` within [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs).

### Sample(N)

Selects **N** evenly distributed pages across the document (first, last, and middle pages) rather than scanning sequentially.

Ideal for huge PDFs with hundreds of pages where full scanning would be too slow. The trade-off is slight precision loss, but sampled pages usually provide sufficient signal for classification.

### Pages(Vec<u32>)

Accepts an explicit vector of 1-indexed page numbers to examine, ignoring all other pages.

Use this when you already know which pages likely contain text—such as when you have a table of contents—or when limiting scanning to specific pages for diagnostic purposes.

## How Detection Applies ScanStrategy

During detection, the `detect_from_document` function in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) (lines 181-199) converts your chosen `ScanStrategy` into two concrete values:

1. **`sample_indices`** — The exact page numbers that will be examined.
2. **`allow_early_exit`** — A boolean flag set to `true` for `EarlyExit` and `false` for other strategies, telling the scanner whether it may stop early.

The detector walks these indices, counts text operators, and returns a `PdfTypeResult` containing confidence scores, OCR recommendations, and page-level reasons.

## Code Examples

### Default Sampling (Sample 8)

```rust
use pdf_inspector::detect_pdf_type;

let result = detect_pdf_type("report.pdf")?;
println!("PDF type: {:?}, confidence: {}", result.pdf_type, result.confidence);

```

The default configuration uses `Sample(8)`, which inspects eight evenly distributed pages. This is defined in the `Default` implementation for `DetectionConfig` in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) (lines 80-88).

### Fast Routing with EarlyExit

```rust
use pdf_inspector::{DetectionConfig, ScanStrategy, detect_pdf_type_with_config};

let config = DetectionConfig {
    strategy: ScanStrategy::EarlyExit,
    ..Default::default()
};

let res = detect_pdf_type_with_config("large_manual.pdf", config)?;
println!("{:?}", res);

```

The detector stops at the first page lacking text operators, making the check very quick.

### Precise Mixed vs Scanned Classification

```rust
let config = DetectionConfig {
    strategy: ScanStrategy::Full,
    ..Default::default()
};

let res = pdf_inspector::detect_pdf_type_with_config("mixed_content.pdf", config)?;
println!("Pages with text: {}", res.pages_with_text);

```

All pages are examined, guaranteeing the most accurate classification.

### Custom Sampling for Large Documents

```rust
let config = DetectionConfig {
    strategy: ScanStrategy::Sample(4),
    ..Default::default()
};

let res = pdf_inspector::detect_pdf_type_with_config("huge_book.pdf", config)?;
println!("Sampled pages: {}", res.pages_sampled);

```

Useful for very large PDFs where speed is critical.

### Targeting Specific Pages

```rust
let config = DetectionConfig {
    strategy: ScanStrategy::Pages(vec![1, 2, 10]),
    ..Default::default()
};

let res = pdf_inspector::detect_pdf_type_with_config("custom_selection.pdf", config)?;
println!("{:?}", res);

```

Only pages 1, 2, and 10 are inspected.

## Key Source Files

- **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** — Defines the `ScanStrategy` enum, `DetectionConfig` struct, and the core `detect_from_document` algorithm.
- **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** — Exposes the public API including `detect_pdf_type` and `detect_pdf_type_with_config`.
- **[`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs)** — CLI example demonstrating strategy toggling via command-line flags.

## Summary

- **EarlyExit** stops at the first non-text page for maximum speed when confirming text-based documents.
- **Full** scans every page to accurately distinguish between Mixed and Scanned PDFs.
- **Sample(N)** checks N evenly distributed pages, balancing speed and accuracy for large documents.
- **Pages(Vec<u32>)** targets specific 1-indexed pages when you know exactly where to look.
- The strategy is passed through `DetectionConfig` to `detect_pdf_type_with_config`, affecting the `sample_indices` and `allow_early_exit` values in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs).

## Frequently Asked Questions

### What is the default ScanStrategy in pdf‑inspector?

The default strategy is `Sample(8)`, which inspects eight evenly distributed pages. This provides a reasonable balance between classification accuracy and performance without requiring explicit configuration, as implemented in the `Default` trait implementation for `DetectionConfig` in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs).

### How does EarlyExit improve performance?

`EarlyExit` improves performance by aborting the scan as soon as it encounters a page lacking text operators (`Tj`/`TJ`). For text-based PDFs where all pages contain extractable text, this means scanning only a few pages instead of the entire document, significantly reducing processing time in fast-path pipelines.

### Can I combine multiple ScanStrategy options?

No, you must select a single variant when constructing `DetectionConfig`. However, you can achieve similar results by using `Pages(Vec<u32>)` to specify exactly which pages to sample, effectively combining the specificity of manual selection with the efficiency of sampling.

### Where is the ScanStrategy enum defined?

The `ScanStrategy` enum is defined in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) at lines 25-40 in the firecrawl/pdf‑inspector repository. This file also contains the logic that interprets each variant into `sample_indices` and `allow_early_exit` parameters used by the detection algorithm.