How to Configure Custom PDF Processing Options in pdf-inspector
Configure custom PDF processing options in pdf-inspector by constructing a PdfOptions struct using its fluent builder API and passing it to process_pdf_with_options defined in src/lib.rs.
The pdf-inspector crate from the Firecrawl organization provides a flexible pipeline for analyzing and converting PDF documents. To customize how the library processes files—ranging from simple type detection to full Markdown extraction—you configure the central PdfOptions struct before invoking the processing functions.
Understanding the PdfOptions Builder Pattern
The PdfOptions struct, located in src/lib.rs, serves as the configuration hub for the entire PDF processing pipeline. It employs a consuming builder pattern where each method takes self by value, updates a specific field, and returns self to enable fluent method chaining.
Core Configuration Fields
Each PdfOptions instance controls the pipeline through these key fields:
mode– Controls how far the pipeline runs using theProcessModeenum (Full,Analyze, orDetectOnly).detection– ADetectionConfigstruct that fine-tunes the PDF type detector, including scan strategies and confidence thresholds.markdown– AMarkdownOptionsstruct that customizes the final Markdown output, such as base font sizes and header/footer stripping.page_filter– An optionalHashSet<u32>that restricts processing to specific 1-indexed pages.password– An optional string for decrypting password-protected PDFs.
Available Builder Methods
The builder implementation in src/lib.rs (lines 176-184) exposes these fluent methods:
| Method | Purpose |
|---|---|
new() |
Creates an instance with defaults (ProcessMode::Full). |
detect_only() |
Shortcut for detection-only runs (ProcessMode::DetectOnly). |
mode(ProcessMode) |
Sets the processing stage. |
detection(DetectionConfig) |
Supplies custom detection configuration. |
markdown(MarkdownOptions) |
Overrides Markdown formatting options. |
pages(iter) |
Limits processing to specified 1-indexed pages. |
password(str) |
Sets the decryption password. |
Selecting the Processing Mode
The ProcessMode enum, defined in src/process_mode.rs, determines which stages execute:
Full– Runs detection, content extraction, and Markdown generation.Analyze– Runs detection and extraction only, skipping Markdown conversion.DetectOnly– Performs only PDF type detection without content extraction.
Customizing Detection and Output Formats
Beyond basic mode selection, you can fine-tune specific pipeline stages using configuration structs passed to the builder.
Detection Configuration
The DetectionConfig struct from src/detector.rs allows you to adjust:
scan_strategy– Choose betweenScanStrategy::Fastfor performance or more comprehensive scanning.confidence_threshold– Set the minimum confidence level for type detection (e.g.,0.6).
Markdown Formatting Options
The MarkdownOptions struct from src/markdown/mod.rs controls text output:
strip_headers_footers– Boolean to remove page headers and footers.base_font_size– Optional float to normalize text sizing.
Implementing Custom PDF Processing
Once configured, pass the PdfOptions instance to process_pdf_with_options, the entry point at src/lib.rs lines 80-86. For in-memory processing, use process_pdf_mem_with_options with a &[u8] buffer instead of a file path.
use pdf_inspector::{
PdfOptions, ProcessMode, DetectionConfig, ScanStrategy,
MarkdownOptions, process_pdf_with_options
};
fn main() -> Result<(), pdf_inspector::PdfError> {
// Initialize with defaults and chain configurations
let opts = PdfOptions::new()
// Set analysis-only mode
.mode(ProcessMode::Analyze)
// Configure detection for speed
.detection({
let mut cfg = DetectionConfig::default();
cfg.scan_strategy = ScanStrategy::Fast;
cfg.confidence_threshold = 0.6;
cfg
})
// Customize Markdown output
.markdown({
let mut md = MarkdownOptions::default();
md.strip_headers_footers = true;
md.base_font_size = Some(11.0);
md
})
// Process only specific pages (1-indexed)
.pages([1, 3, 5]);
// Optional: .password("secret123")
// Execute the pipeline
let result = process_pdf_with_options("document.pdf", opts)?;
println!("Detected type: {:?}", result.pdf_type);
if let Some(md) = result.markdown {
println!("Generated Markdown:\n{}", md);
}
Ok(())
}
Summary
PdfOptionsin src/lib.rs is the central configuration struct for customizing PDF processing in the firecrawl/pdf-inspector repository.- Use the fluent builder API (methods like
mode(),detection(),markdown(), andpages()) to chain configuration options without intermediate variables. - Select the appropriate
ProcessMode(Full,Analyze, orDetectOnly) to control pipeline depth according to src/process_mode.rs. - Pass the configured options to
process_pdf_with_optionsorprocess_pdf_mem_with_optionsto execute processing with your custom settings.
Frequently Asked Questions
How do I process only specific pages of a PDF?
Use the pages() builder method on PdfOptions, passing an iterator of 1-indexed page numbers (e.g., opts.pages([1, 3, 5])). The pipeline will skip all other pages, improving performance for large documents according to the implementation in src/lib.rs.
What is the difference between ProcessMode::Analyze and ProcessMode::Full?
Analyze runs detection and content extraction but skips Markdown generation, returning structured data about the PDF content. Full executes the complete pipeline including final Markdown conversion, as implemented in the processing logic at src/lib.rs.
How do I handle password-protected PDFs?
Chain the password() method when building your PdfOptions instance, providing the decryption string (e.g., opts.password("secret123")). The library will attempt decryption before processing the document contents.
Can I use PdfOptions with in-memory PDF data?
Yes. Instead of process_pdf_with_options, call process_pdf_mem_with_options and pass a &[u8] buffer along with your configured PdfOptions. This avoids filesystem I/O when working with downloaded or generated PDF bytes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →