How to Configure Custom PDF Processing Options in pdf-inspector

Configure custom PDF processing options in pdf-inspector by constructing a PdfOptions struct using its fluent builder API and passing it to process_pdf_with_options defined in src/lib.rs.

The pdf-inspector crate from the Firecrawl organization provides a flexible pipeline for analyzing and converting PDF documents. To customize how the library processes files—ranging from simple type detection to full Markdown extraction—you configure the central PdfOptions struct before invoking the processing functions.

Understanding the PdfOptions Builder Pattern

The PdfOptions struct, located in src/lib.rs, serves as the configuration hub for the entire PDF processing pipeline. It employs a consuming builder pattern where each method takes self by value, updates a specific field, and returns self to enable fluent method chaining.

Core Configuration Fields

Each PdfOptions instance controls the pipeline through these key fields:

  • mode – Controls how far the pipeline runs using the ProcessMode enum (Full, Analyze, or DetectOnly).
  • detection – A DetectionConfig struct that fine-tunes the PDF type detector, including scan strategies and confidence thresholds.
  • markdown – A MarkdownOptions struct that customizes the final Markdown output, such as base font sizes and header/footer stripping.
  • page_filter – An optional HashSet<u32> that restricts processing to specific 1-indexed pages.
  • password – An optional string for decrypting password-protected PDFs.

Available Builder Methods

The builder implementation in src/lib.rs (lines 176-184) exposes these fluent methods:

Method Purpose
new() Creates an instance with defaults (ProcessMode::Full).
detect_only() Shortcut for detection-only runs (ProcessMode::DetectOnly).
mode(ProcessMode) Sets the processing stage.
detection(DetectionConfig) Supplies custom detection configuration.
markdown(MarkdownOptions) Overrides Markdown formatting options.
pages(iter) Limits processing to specified 1-indexed pages.
password(str) Sets the decryption password.

Selecting the Processing Mode

The ProcessMode enum, defined in src/process_mode.rs, determines which stages execute:

  • Full – Runs detection, content extraction, and Markdown generation.
  • Analyze – Runs detection and extraction only, skipping Markdown conversion.
  • DetectOnly – Performs only PDF type detection without content extraction.

Customizing Detection and Output Formats

Beyond basic mode selection, you can fine-tune specific pipeline stages using configuration structs passed to the builder.

Detection Configuration

The DetectionConfig struct from src/detector.rs allows you to adjust:

  • scan_strategy – Choose between ScanStrategy::Fast for performance or more comprehensive scanning.
  • confidence_threshold – Set the minimum confidence level for type detection (e.g., 0.6).

Markdown Formatting Options

The MarkdownOptions struct from src/markdown/mod.rs controls text output:

  • strip_headers_footers – Boolean to remove page headers and footers.
  • base_font_size – Optional float to normalize text sizing.

Implementing Custom PDF Processing

Once configured, pass the PdfOptions instance to process_pdf_with_options, the entry point at src/lib.rs lines 80-86. For in-memory processing, use process_pdf_mem_with_options with a &[u8] buffer instead of a file path.

use pdf_inspector::{
    PdfOptions, ProcessMode, DetectionConfig, ScanStrategy,
    MarkdownOptions, process_pdf_with_options
};

fn main() -> Result<(), pdf_inspector::PdfError> {
    // Initialize with defaults and chain configurations
    let opts = PdfOptions::new()
        // Set analysis-only mode
        .mode(ProcessMode::Analyze)
        // Configure detection for speed
        .detection({
            let mut cfg = DetectionConfig::default();
            cfg.scan_strategy = ScanStrategy::Fast;
            cfg.confidence_threshold = 0.6;
            cfg
        })
        // Customize Markdown output
        .markdown({
            let mut md = MarkdownOptions::default();
            md.strip_headers_footers = true;
            md.base_font_size = Some(11.0);
            md
        })
        // Process only specific pages (1-indexed)
        .pages([1, 3, 5]);
        // Optional: .password("secret123")

    // Execute the pipeline
    let result = process_pdf_with_options("document.pdf", opts)?;
    
    println!("Detected type: {:?}", result.pdf_type);
    if let Some(md) = result.markdown {
        println!("Generated Markdown:\n{}", md);
    }
    
    Ok(())
}

Summary

  • PdfOptions in src/lib.rs is the central configuration struct for customizing PDF processing in the firecrawl/pdf-inspector repository.
  • Use the fluent builder API (methods like mode(), detection(), markdown(), and pages()) to chain configuration options without intermediate variables.
  • Select the appropriate ProcessMode (Full, Analyze, or DetectOnly) to control pipeline depth according to src/process_mode.rs.
  • Pass the configured options to process_pdf_with_options or process_pdf_mem_with_options to execute processing with your custom settings.

Frequently Asked Questions

How do I process only specific pages of a PDF?

Use the pages() builder method on PdfOptions, passing an iterator of 1-indexed page numbers (e.g., opts.pages([1, 3, 5])). The pipeline will skip all other pages, improving performance for large documents according to the implementation in src/lib.rs.

What is the difference between ProcessMode::Analyze and ProcessMode::Full?

Analyze runs detection and content extraction but skips Markdown generation, returning structured data about the PDF content. Full executes the complete pipeline including final Markdown conversion, as implemented in the processing logic at src/lib.rs.

How do I handle password-protected PDFs?

Chain the password() method when building your PdfOptions instance, providing the decryption string (e.g., opts.password("secret123")). The library will attempt decryption before processing the document contents.

Can I use PdfOptions with in-memory PDF data?

Yes. Instead of process_pdf_with_options, call process_pdf_mem_with_options and pass a &[u8] buffer along with your configured PdfOptions. This avoids filesystem I/O when working with downloaded or generated PDF bytes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →