# How to Configure Custom PDF Processing Options in pdf-inspector

> Learn to configure custom PDF processing options in pdf-inspector. Use the fluent builder API and process_pdf_with_options for advanced control over your PDF analysis.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-08

---

**Configure custom PDF processing options in pdf-inspector by constructing a `PdfOptions` struct using its fluent builder API and passing it to `process_pdf_with_options` defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs).**

The `pdf-inspector` crate from the Firecrawl organization provides a flexible pipeline for analyzing and converting PDF documents. To customize how the library processes files—ranging from simple type detection to full Markdown extraction—you configure the central `PdfOptions` struct before invoking the processing functions.

## Understanding the `PdfOptions` Builder Pattern

The **`PdfOptions`** struct, located in **src/lib.rs**, serves as the configuration hub for the entire PDF processing pipeline. It employs a consuming builder pattern where each method takes `self` by value, updates a specific field, and returns `self` to enable fluent method chaining.

### Core Configuration Fields

Each `PdfOptions` instance controls the pipeline through these key fields:

- **`mode`** – Controls how far the pipeline runs using the `ProcessMode` enum (`Full`, `Analyze`, or `DetectOnly`).
- **`detection`** – A `DetectionConfig` struct that fine-tunes the PDF type detector, including scan strategies and confidence thresholds.
- **`markdown`** – A `MarkdownOptions` struct that customizes the final Markdown output, such as base font sizes and header/footer stripping.
- **`page_filter`** – An optional `HashSet<u32>` that restricts processing to specific 1-indexed pages.
- **`password`** – An optional string for decrypting password-protected PDFs.

### Available Builder Methods

The builder implementation in **src/lib.rs** (lines 176-184) exposes these fluent methods:

| Method | Purpose |
|--------|---------|
| `new()` | Creates an instance with defaults (`ProcessMode::Full`). |
| `detect_only()` | Shortcut for detection-only runs (`ProcessMode::DetectOnly`). |
| `mode(ProcessMode)` | Sets the processing stage. |
| `detection(DetectionConfig)` | Supplies custom detection configuration. |
| `markdown(MarkdownOptions)` | Overrides Markdown formatting options. |
| `pages(iter)` | Limits processing to specified 1-indexed pages. |
| `password(str)` | Sets the decryption password. |

## Selecting the Processing Mode

The **`ProcessMode`** enum, defined in **src/process_mode.rs**, determines which stages execute:

- **`Full`** – Runs detection, content extraction, and Markdown generation.
- **`Analyze`** – Runs detection and extraction only, skipping Markdown conversion.
- **`DetectOnly`** – Performs only PDF type detection without content extraction.

## Customizing Detection and Output Formats

Beyond basic mode selection, you can fine-tune specific pipeline stages using configuration structs passed to the builder.

### Detection Configuration

The **`DetectionConfig`** struct from **src/detector.rs** allows you to adjust:

- **`scan_strategy`** – Choose between `ScanStrategy::Fast` for performance or more comprehensive scanning.
- **`confidence_threshold`** – Set the minimum confidence level for type detection (e.g., `0.6`).

### Markdown Formatting Options

The **`MarkdownOptions`** struct from **src/markdown/mod.rs** controls text output:

- **`strip_headers_footers`** – Boolean to remove page headers and footers.
- **`base_font_size`** – Optional float to normalize text sizing.

## Implementing Custom PDF Processing

Once configured, pass the `PdfOptions` instance to **`process_pdf_with_options`**, the entry point at **src/lib.rs** lines 80-86. For in-memory processing, use **`process_pdf_mem_with_options`** with a `&[u8]` buffer instead of a file path.

```rust
use pdf_inspector::{
    PdfOptions, ProcessMode, DetectionConfig, ScanStrategy,
    MarkdownOptions, process_pdf_with_options
};

fn main() -> Result<(), pdf_inspector::PdfError> {
    // Initialize with defaults and chain configurations
    let opts = PdfOptions::new()
        // Set analysis-only mode
        .mode(ProcessMode::Analyze)
        // Configure detection for speed
        .detection({
            let mut cfg = DetectionConfig::default();
            cfg.scan_strategy = ScanStrategy::Fast;
            cfg.confidence_threshold = 0.6;
            cfg
        })
        // Customize Markdown output
        .markdown({
            let mut md = MarkdownOptions::default();
            md.strip_headers_footers = true;
            md.base_font_size = Some(11.0);
            md
        })
        // Process only specific pages (1-indexed)
        .pages([1, 3, 5]);
        // Optional: .password("secret123")

    // Execute the pipeline
    let result = process_pdf_with_options("document.pdf", opts)?;
    
    println!("Detected type: {:?}", result.pdf_type);
    if let Some(md) = result.markdown {
        println!("Generated Markdown:\n{}", md);
    }
    
    Ok(())
}

```

## Summary

- **`PdfOptions`** in **src/lib.rs** is the central configuration struct for customizing PDF processing in the firecrawl/pdf-inspector repository.
- Use the **fluent builder API** (methods like `mode()`, `detection()`, `markdown()`, and `pages()`) to chain configuration options without intermediate variables.
- Select the appropriate **`ProcessMode`** (`Full`, `Analyze`, or `DetectOnly`) to control pipeline depth according to **src/process_mode.rs**.
- Pass the configured options to **`process_pdf_with_options`** or **`process_pdf_mem_with_options`** to execute processing with your custom settings.

## Frequently Asked Questions

### How do I process only specific pages of a PDF?

Use the **`pages()`** builder method on `PdfOptions`, passing an iterator of 1-indexed page numbers (e.g., `opts.pages([1, 3, 5])`). The pipeline will skip all other pages, improving performance for large documents according to the implementation in **src/lib.rs**.

### What is the difference between `ProcessMode::Analyze` and `ProcessMode::Full`?

**`Analyze`** runs detection and content extraction but skips Markdown generation, returning structured data about the PDF content. **`Full`** executes the complete pipeline including final Markdown conversion, as implemented in the processing logic at **src/lib.rs**.

### How do I handle password-protected PDFs?

Chain the **`password()`** method when building your `PdfOptions` instance, providing the decryption string (e.g., `opts.password("secret123")`). The library will attempt decryption before processing the document contents.

### Can I use `PdfOptions` with in-memory PDF data?

Yes. Instead of `process_pdf_with_options`, call **`process_pdf_mem_with_options`** and pass a `&[u8]` buffer along with your configured `PdfOptions`. This avoids filesystem I/O when working with downloaded or generated PDF bytes.