# How to Configure pdf-inspector for Specific PDF Analysis Tasks

> Configure pdf-inspector for specific PDF analysis tasks using the PdfOptions builder. Tailor processing, detection, filtering, decryption, and formatting for precise PDF pipeline control.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-04

---

**Use the `PdfOptions` builder in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) to chain configuration methods that control processing mode, detection strategy, page filtering, password decryption, and markdown formatting, enabling precise tailoring of the PDF pipeline without redundant I/O.**

The `pdf-inspector` crate from firecrawl/pdf-inspector centralizes every behavioral knob in the **`PdfOptions`** builder. By chaining methods before calling `process_pdf_with_options()`, you can tailor the entire pipeline—detection, extraction, and markdown rendering—to match specific workflow requirements, from rapid classification to deep content analysis.

## Understanding the PdfOptions Builder Pattern

All configuration in `pdf-inspector` flows through the `PdfOptions` struct defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs). This builder pattern ensures that both detection and extraction phases share the same parsed `lopdf::Document`, eliminating redundant I/O operations regardless of how many options you set (see the architecture diagram in [`README.md`](https://github.com/firecrawl/pdf-inspector/blob/main/README.md) lines 66-78).

The typical entry point is `process_pdf_with_options(path, options)`, where `options` is a `PdfOptions` instance configured via method chaining.

## Configuring Processing Modes

The **processing mode** determines how far the pipeline executes. Set this via `PdfOptions::mode()` (defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) lines 30-34).

### Available Process Modes

- **`ProcessMode::Full`** (default) — Runs complete detection, extraction, and markdown generation.
- **`ProcessMode::Analyze`** — Performs detection plus layout analysis (tables, columns) without generating markdown output.
- **`ProcessMode::DetectOnly`** — Executes fast metadata-only detection, ideal for routing decisions when you only need to know if OCR is required.

```rust
use pdf_inspector::{PdfOptions, ProcessMode};

// Fast routing: detect PDF type without extraction
let result = pdf_inspector::process_pdf_with_options(
    "document.pdf",
    PdfOptions::new().mode(ProcessMode::DetectOnly)
);
// result.pdf_type indicates TextBased, Scanned, or Mixed

```

## Optimizing Detection Strategies

Control how aggressively the engine scans pages using **`DetectionConfig`** and **`ScanStrategy`** (defined in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) lines 25-40 and 71-78).

### ScanStrategy Variants

- **`ScanStrategy::EarlyExit`** (default) — Stops on the first non-text page. Optimal for pipelines routing TextBased PDFs to fast extraction paths.
- **`ScanStrategy::Full`** — Scans every page, providing accurate distinction between Mixed and Scanned documents.
- **`ScanStrategy::Sample(n)`** — Samples *n* evenly spaced pages, balancing speed and accuracy for huge PDFs.
- **`ScanStrategy::Pages(vec)`** — Scans only user-specified pages.

```rust
use pdf_inspector::{PdfOptions, DetectionConfig, ScanStrategy, ProcessMode};

// Large PDF analysis: sample 6 pages for detection, then full extraction
let detect_cfg = DetectionConfig {
    strategy: ScanStrategy::Sample(6),
    ..Default::default()
};

let opts = PdfOptions::new()
    .detection(detect_cfg)
    .mode(ProcessMode::Full);

let result = pdf_inspector::process_pdf_with_options("big.pdf", opts)?;

```

## Filtering Pages and Handling Encryption

### Selective Page Extraction

When you need only specific pages, use `PdfOptions::pages([...])`. This populates `page_filter: Option<HashSet<u32>>` (see [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) lines 48-52), ensuring extraction runs only on the requested subset.

```rust
// Extract only pages 2, 3, and 4
let opts = PdfOptions::new()
    .pages([2, 3, 4])
    .mode(ProcessMode::Full);

let result = pdf_inspector::process_pdf_with_options("financials.pdf", opts)?;

```

### Password-Protected PDFs

For encrypted documents, use `PdfOptions::password("secret")`. The implementation includes a custom `Debug` trait (lines 90-100 in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)) that redacts the password from logs to prevent secret leakage.

```rust
let opts = PdfOptions::new()
    .password("open-sesame")
    .mode(ProcessMode::Full);

let secure = pdf_inspector::process_pdf_with_options("encrypted.pdf", opts)?;

```

## Customizing Markdown Output

Fine-tune markdown generation via **`MarkdownOptions`** (located in [`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs) and consumed in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)). Pass this to `PdfOptions::markdown()` to control formatting flags such as `compact`, `raw`, and `pages`.

```rust
use pdf_inspector::{PdfOptions, MarkdownOptions, ProcessMode};

let markdown_cfg = MarkdownOptions {
    compact: true,
    raw: false,
    pages: true,
    ..Default::default()
};

let opts = PdfOptions::new()
    .markdown(markdown_cfg)
    .mode(ProcessMode::Full);

```

## Complete Configuration Examples

Combine multiple options to build task-specific pipelines:

**Fast classification for routing decisions:**

```rust
let opts = PdfOptions::new().mode(ProcessMode::DetectOnly);

```

**Full extraction with custom markdown on a password-protected document:**

```rust
let opts = PdfOptions::new()
    .password("my-pass")
    .mode(ProcessMode::Full)
    .markdown(MarkdownOptions::default());

```

**High-performance analysis of massive PDFs:**

```rust
let opts = PdfOptions::new()
    .detection(DetectionConfig {
        strategy: ScanStrategy::Sample(4),
        ..Default::default()
    })
    .mode(ProcessMode::Full);

```

## Summary

- **`PdfOptions`** in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) serves as the central configuration hub using a fluent builder API.
- **Processing modes** (`Full`, `Analyze`, `DetectOnly`) control pipeline depth without re-parsing the document.
- **ScanStrategy** variants (`EarlyExit`, `Full`, `Sample`, `Pages`) optimize detection speed for different document sizes and accuracy requirements.
- **Page filtering** via `pages([...])` restricts extraction to specific page numbers.
- **Password handling** automatically redacts secrets from debug output while enabling decryption.
- **MarkdownOptions** in [`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs) allow granular control over output formatting.

## Frequently Asked Questions

### How do I extract content from only specific pages?

Use the `pages()` method on `PdfOptions` with an array of page numbers. This sets the internal `page_filter: Option<HashSet<u32>>` field (defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) lines 48-52), ensuring the extractor processes only the requested pages while still sharing the same underlying `lopdf::Document` instance.

### What is the difference between ProcessMode::Analyze and ProcessMode::Full?

`ProcessMode::Analyze` runs detection plus layout analysis—including table and column detection—without generating markdown output, making it suitable for structural analysis. `ProcessMode::Full` executes the complete pipeline including markdown generation via the conversion engine in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs).

### How does pdf-inspector handle password-protected PDFs?

Pass the password via `PdfOptions::password("your-password")`. The crate decrypts the document using the provided credentials while implementing a custom `Debug` trait (lines 90-100 in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)) that masks the password field to prevent accidental exposure in logs or panic traces.

### Which ScanStrategy should I use for large PDFs?

For large documents where CPU time is constrained, use `ScanStrategy::Sample(n)` to scan *n* evenly spaced pages, or `ScanStrategy::EarlyExit` to stop at the first non-text page if you only need to identify TextBased PDFs. These strategies are defined in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) (lines 25-40) and wrapped by `DetectionConfig` (lines 71-78).