# How to Configure Firecrawl PDF‑Inspector for Specific Use Cases: Detection, Speed, and Extraction Modes

> Master Firecrawl PDF-Inspector configuration for your specific needs. Learn to optimize detection, speed, and extraction modes with tailored settings for peak performance.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Configure Firecrawl PDF‑Inspector by adjusting `DetectionConfig` for page‑sampling strategies and `PdfOptions` for pipeline modes, OCR thresholds, page filtering, and encrypted PDF handling.**

The Firecrawl PDF‑Inspector is a Rust‑based modular pipeline that separates **detection**, **extraction**, and **markdown generation** into configurable stages. This article shows how to tailor the library for speed‑critical batch processing, precise document classification, selective page extraction, and encrypted file handling—using the actual source structures found in `firecrawl/pdf‑inspector`.

## Core Configuration Structures

PDF‑Inspector's flexibility centers on two main types defined in the source code:

- **`DetectionConfig`** ([`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), lines 69‑78) — Controls how PDFs are inspected for text‑based vs. scanned content
- **`PdfOptions`** ([`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), lines 164‑180) — Builder for end‑to‑end run settings including mode, detection config, page filters, and passwords

The pipeline's execution mode is determined by `ProcessMode` ([`src/process_mode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/process_mode.rs), lines 7‑15), an enum with three variants:

```rust
pub enum ProcessMode {
    DetectOnly,  // Fast detection only, no extraction
    Full,        // Detection → extraction → markdown (default)
    Analyze,     // Detection + layout analysis without markdown
}

```

## Configuring Detection Behavior with DetectionConfig

The `DetectionConfig` struct defines three critical thresholds:

```rust
pub struct DetectionConfig {
    pub strategy: ScanStrategy,           // Which pages to scan
    pub min_text_ops_per_page: u32,       // Minimum Tj/TJ operators for "text" page
    pub text_page_ratio_threshold: f32,   // Text pages / total pages for TextBased
}

```

### ScanStrategy Variants for Speed vs. Accuracy

| Strategy | Implementation | Best For |
|----------|---------------|----------|
| `EarlyExit` | Stops at first non‑text page | Quickly rejecting clearly scanned documents |
| `Full` | Scans every page | Precise Mixed/Scanned classification (legal contracts, mixed reports) |
| `Sample(N)` | Samples N evenly‑distributed pages | Large PDFs where speed matters |
| `Pages(vec)` | Scans only specified 1‑indexed pages | Known relevant pages (cover, index, appendix) |

**Default values** (defined in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), lines 80‑89):

- `ScanStrategy::Sample(8)`
- `min_text_ops_per_page: 3`
- `text_page_ratio_threshold: 0.6`

These defaults provide balanced accuracy for typical documents without scanning every page.

## Configuring Pipeline Execution with PdfOptions

`PdfOptions` uses a fluent builder pattern to assemble complete run configurations. The struct fields (from [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), lines 164‑180) include:

- `mode: ProcessMode` — Pipeline stage selection
- `detection: DetectionConfig` — Detection parameters
- `markdown: MarkdownOptions` — Output formatting (Full mode only)
- `page_filter: Option<Vec<PageNum>>` — Subset of pages to process
- `password: Option<String>` — Decryption password (with custom `Debug` to prevent logging)

Builder methods are implemented in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) at lines 30‑59.

## Use Case: Default Fast Detection

For quick classification without extraction, use the `detect_pdf` convenience function:

```rust
use pdf_inspector::detect_pdf;

let result = detect_pdf("report.pdf")?;
println!("Detected type: {:?}", result.pdf_type);

```

This internally calls `process_pdf_with_options` with `PdfOptions::detect_only()`, which sets `mode = ProcessMode::DetectOnly` and default detection config (source: [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), lines 73‑78).

## Use Case: Speed‑Focused Sampling for Large PDFs

Reduce detection overhead by sampling fewer pages:

```rust
use pdf_inspector::{PdfOptions, DetectionConfig, ScanStrategy};

let config = DetectionConfig {
    strategy: ScanStrategy::Sample(4), // Only 4 pages examined
    ..DetectionConfig::default()
};

let opts = PdfOptions::new().detection(config);
let result = pdf_inspector::process_pdf_with_options("huge.pdf", opts)?;

println!("Pages sampled: {}", result.pages_sampled);

```

**Performance impact:** Detection becomes O(1) regardless of total page count. The four sampled pages are distributed across the document (first, last, and intermediate pages) to improve representativeness.

## Use Case: Precise Mixed/Scanned Classification

Force complete page analysis for documents requiring accurate type detection:

```rust
use pdf_inspector::{PdfOptions, DetectionConfig, ScanStrategy, ProcessMode};

let config = DetectionConfig {
    strategy: ScanStrategy::Full,        // Every page scanned
    min_text_ops_per_page: 5,            // Stricter text threshold
    text_page_ratio_threshold: 0.7,      // Higher bar for TextBased
    ..DetectionConfig::default()
};

let opts = PdfOptions::new()
    .mode(ProcessMode::Full)
    .detection(config);

let result = pdf_inspector::process_pdf_with_options("contract.pdf", opts)?;
println!("PDF type: {:?}", result.pdf_type);

```

**When to use:** Legal documents, academic papers with scanned figures, or reports mixing generated and raster content. The `Full` strategy prevents early‑exit misclassification when a single non‑text page appears early in the document.

## Use Case: Selective Page Processing

Extract only specific pages without parsing the entire file:

```rust
use pdf_inspector::PdfOptions;

let opts = PdfOptions::new()
    .pages([1, 10, 20]); // 1‑indexed page numbers

let result = pdf_inspector::process_pdf_with_options("book.pdf", opts)?;
println!("Extracted {} pages", result.pages.len());

```

The `pages` method accepts any `IntoIterator<Item = PageNum>` and propagates the filter through the extraction pipeline ([`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)).

## Use Case: Encrypted PDF Handling

Provide decryption passwords safely:

```rust
use pdf_inspector::{PdfOptions, ProcessMode};

let opts = PdfOptions::new()
    .password("my‑pw")
    .mode(ProcessMode::Full);

let result = pdf_inspector::process_pdf_with_options("encrypted.pdf", opts)?;
println!("Decrypted and extracted {} pages", result.pages.len());

```

**Security note:** The password field uses a custom `Debug` implementation that masks the actual value, preventing accidental exposure in logs.

## Complete Configuration Example

Combine multiple options for production workflows:

```rust
use pdf_inspector::{
    ProcessMode, PdfOptions, DetectionConfig, ScanStrategy, MarkdownOptions,
};

let detection_cfg = DetectionConfig {
    strategy: ScanStrategy::Sample(6),   // Faster for 300‑page reports
    min_text_ops_per_page: 4,
    text_page_ratio_threshold: 0.65,
    ..Default::default()
};

let markdown_cfg = MarkdownOptions::default(); // Customize in src/markdown/mod.rs

let opts = PdfOptions::new()
    .mode(ProcessMode::Full)
    .detection(detection_cfg)
    .markdown(markdown_cfg)
    .pages(1..=10);                       // First ten pages only

let result = pdf_inspector::process_pdf_with_options("annual_report.pdf", opts)?;
println!("Completed: {} markdown blocks", result.markdown.len());

```

## Key Source Files Reference

| File | Purpose | Lines |
|------|---------|-------|
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | `DetectionConfig`, `ScanStrategy`, detection algorithm | 69‑89 |
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | `PdfOptions` builder, public API entry points | 30‑59, 73‑78, 164‑180 |
| [`src/process_mode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/process_mode.rs) | `ProcessMode` enum definition | 7‑15 |
| [`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs) | `MarkdownOptions` and output formatting | Module root |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Extraction orchestration respecting page filters | Module root |
| [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs) | Table detection strategies | Module root |

## Summary

- **Use `ProcessMode::DetectOnly`** for fastest classification without content extraction
- **Adjust `ScanStrategy`** to trade speed for accuracy: `Sample(N)` for large files, `Full` for precise classification, `Pages(vec)` for selective analysis
- **Tune OCR thresholds** via `min_text_ops_per_page` and `text_page_ratio_threshold` to match your document types
- **Apply `pages()` and `password()`** methods on `PdfOptions` for targeted processing and secure decryption
- **Prefer `PdfOptions::new()` builder** over raw struct initialization for forward compatibility

## Frequently Asked Questions

### What is the fastest way to classify PDF type without extracting content?

Use `pdf_inspector::detect_pdf("file.pdf")` or explicitly configure `PdfOptions::new().mode(ProcessMode::DetectOnly)`. This skips extraction and markdown generation entirely, running only the detection algorithm with default `DetectionConfig`.

### How do I process a 1000‑page PDF without scanning every page?

Set `DetectionConfig { strategy: ScanStrategy::Sample(4), ..Default::default() }` and pass it to `PdfOptions::detection()`. The detector examines only four evenly‑spaced pages regardless of total page count, making runtime effectively constant.

### When should I use ProcessMode::Analyze instead of Full?

Use `Analyze` when you need structural information—table boundaries, column positions, reading order—without the overhead of markdown serialization. This mode runs detection plus layout analysis but skips the final markdown generation step, saving CPU for downstream custom formatters.

### How do I prevent passwords from appearing in logs?

The `PdfOptions` struct implements a custom `Debug` trait that masks the `password` field. When you call `.password("secret")`, the value is stored for decryption but displays as redacted in any `{:?}` formatting or logging output.