How to Use the `PdfOptions` Builder to Configure ProcessMode, Page Filters, and Passwords in pdf-inspector

The PdfOptions builder in pdf-inspector provides a fluent, chainable interface for configuring processing pipelines, page subsets, and decryption passwords through methods like mode(), pages(), and password().

The PdfOptions struct in firecrawl/pdf-inspector serves as the central configuration object for all high-level PDF processing operations. Following the classic builder pattern, each method consumes self, updates an internal field, and returns the modified struct—enabling method chaining for concise, readable configuration.

Understanding the PdfOptions Builder Pattern

The builder's fields and default values are defined in src/lib.rs at lines 176–187. The concrete implementation methods follow immediately after the impl Default block, providing type-safe configuration for three core dimensions: processing depth, page scope, and document encryption.

Builder method Configures Default
mode(ProcessMode) Pipeline stage (Full, DetectOnly, Analyze) ProcessMode::Full
pages(iter) 1-indexed page subset as HashSet<u32> None (all pages)
password(pw) User password for encrypted PDFs None (empty password)

The password field is intentionally obscured in debug output—displayed as "[REDACTED]"—protecting sensitive credentials in logs.

Configuring ProcessMode with mode()

The mode() method sets how far the processing pipeline runs. As implemented in src/lib.rs lines 30–34, it accepts any variant of the ProcessMode enum (defined in src/process_mode.rs):

use pdf_inspector::{PdfOptions, ProcessMode};

// Run complete extraction pipeline
let opts = PdfOptions::new().mode(ProcessMode::Full);

// Stop early after metadata detection only
let opts = PdfOptions::new().mode(ProcessMode::DetectOnly);

ProcessMode::Full executes the complete pipeline including content extraction. ProcessMode::DetectOnly halts after document type identification, useful for quick classification without parsing overhead.

Filtering Pages with pages()

The pages() method restricts processing to specific 1-indexed pages. Implemented at lines 48–52 in src/lib.rs, it accepts any iterator yielding u32 values—arrays, vectors, or ranges—and converts them to an internal HashSet<u32> for deduplication:

use pdf_inspector::PdfOptions;

// Process specific pages only
let opts = PdfOptions::new().pages([2, 5, 10]);

// Process first 10 pages using range syntax
let opts = PdfOptions::new().pages(1..=10);

// Process from a Vec
let subset: Vec<u32> = vec![1, 3, 7];
let opts = PdfOptions::new().pages(subset);

The page filter is passed to load_document_from_path_with_password() in src/detector.rs, limiting which pages the core loader examines.

Decrypting PDFs with password()

For password-protected PDFs, the password() method supplies decryption credentials. As shown in lines 54–57 of src/lib.rs:

use pdf_inspector::PdfOptions;

let opts = PdfOptions::new().password("my-secret");

When process_pdf_with_options() executes, the password becomes:

let (doc, page_count) = load_document_from_path_with_password(
    &path,
    options.password.as_deref(),
)?;

The as_deref() call converts Option<String> to Option<&str>, matching the loader's signature.

Complete Configuration Examples

Example 1: Full Extraction on Select Pages of Encrypted PDF

use pdf_inspector::{PdfOptions, ProcessMode};

let opts = PdfOptions::new()
    .mode(ProcessMode::Full)          // Complete pipeline
    .pages([2, 5])                    // Pages 2 and 5 only
    .password("my-secret");           // Decrypt first

let result = pdf_inspector::process_pdf_with_options("protected.pdf", opts)?;
println!("{:?}", result);

Example 2: Detection-Only with Page Limit

use pdf_inspector::{PdfOptions, ProcessMode};

let opts = PdfOptions::new()
    .mode(ProcessMode::DetectOnly)   // Metadata detection only
    .pages(1..=10);                   // Optional: limit scan scope

let detection = pdf_inspector::process_pdf_with_options("big.pdf", opts)?;
println!("Detected type: {:?}", detection.pdf_type);

Key Source Files Reference

File Purpose
src/lib.rs Defines PdfOptions struct, defaults, and builder methods (mode, pages, password) at lines 176–187 and surrounding implementation blocks
src/process_mode.rs ProcessMode enum definition with pipeline stage variants
src/detector.rs Consumes PdfOptions in load_document_from_path_with_password()
src/bin/pdf2md.rs CLI demonstrating real-world builder usage from command-line arguments

Summary

  • Method chaining: Each PdfOptions method returns self, enabling fluent .mode().pages().password() sequences
  • Flexible page input: pages() accepts any IntoIterator<Item = u32>—arrays, ranges, vectors
  • Security-conscious: Passwords are redacted in Debug output to prevent credential leakage
  • Pipeline control: ProcessMode determines whether detection, analysis, or full extraction runs
  • Direct integration: Builder values flow directly into load_document_from_path_with_password() for document loading

Frequently Asked Questions

What happens if I don't call mode() on PdfOptions?

The builder defaults to ProcessMode::Full through its Default implementation. Without explicit configuration, the complete extraction pipeline runs—equivalent to calling .mode(ProcessMode::Full).

Can I use pages() with overlapping or out-of-order values?

Yes. The implementation converts any iterator to a HashSet<u32>, automatically deduplicating duplicates and ignoring order. Pages 5, 2, 5 become {2, 5}; processing occurs in PDF-native page order regardless of input sequence.

Does password() accept empty strings for owner-only encryption?

Yes. Passing .password("") or omitting the call entirely both result in None for the internal Option<String>, triggering the empty password path in the PDF loader. For protected PDFs with non-empty user passwords, the correct value must be supplied.

Is PdfOptions reusable across multiple PDFs?

No—each builder method consumes self. To process multiple files with identical settings, call PdfOptions::new() and re-chain the configuration for each operation, or clone a configured instance with .clone() before passing to processing functions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →