How to Configure pdf-inspector for Specific PDF Analysis Tasks
Use the PdfOptions builder in src/lib.rs to chain configuration methods that control processing mode, detection strategy, page filtering, password decryption, and markdown formatting, enabling precise tailoring of the PDF pipeline without redundant I/O.
The pdf-inspector crate from firecrawl/pdf-inspector centralizes every behavioral knob in the PdfOptions builder. By chaining methods before calling process_pdf_with_options(), you can tailor the entire pipeline—detection, extraction, and markdown rendering—to match specific workflow requirements, from rapid classification to deep content analysis.
Understanding the PdfOptions Builder Pattern
All configuration in pdf-inspector flows through the PdfOptions struct defined in src/lib.rs. This builder pattern ensures that both detection and extraction phases share the same parsed lopdf::Document, eliminating redundant I/O operations regardless of how many options you set (see the architecture diagram in README.md lines 66-78).
The typical entry point is process_pdf_with_options(path, options), where options is a PdfOptions instance configured via method chaining.
Configuring Processing Modes
The processing mode determines how far the pipeline executes. Set this via PdfOptions::mode() (defined in src/lib.rs lines 30-34).
Available Process Modes
ProcessMode::Full(default) — Runs complete detection, extraction, and markdown generation.ProcessMode::Analyze— Performs detection plus layout analysis (tables, columns) without generating markdown output.ProcessMode::DetectOnly— Executes fast metadata-only detection, ideal for routing decisions when you only need to know if OCR is required.
use pdf_inspector::{PdfOptions, ProcessMode};
// Fast routing: detect PDF type without extraction
let result = pdf_inspector::process_pdf_with_options(
"document.pdf",
PdfOptions::new().mode(ProcessMode::DetectOnly)
);
// result.pdf_type indicates TextBased, Scanned, or Mixed
Optimizing Detection Strategies
Control how aggressively the engine scans pages using DetectionConfig and ScanStrategy (defined in src/detector.rs lines 25-40 and 71-78).
ScanStrategy Variants
ScanStrategy::EarlyExit(default) — Stops on the first non-text page. Optimal for pipelines routing TextBased PDFs to fast extraction paths.ScanStrategy::Full— Scans every page, providing accurate distinction between Mixed and Scanned documents.ScanStrategy::Sample(n)— Samples n evenly spaced pages, balancing speed and accuracy for huge PDFs.ScanStrategy::Pages(vec)— Scans only user-specified pages.
use pdf_inspector::{PdfOptions, DetectionConfig, ScanStrategy, ProcessMode};
// Large PDF analysis: sample 6 pages for detection, then full extraction
let detect_cfg = DetectionConfig {
strategy: ScanStrategy::Sample(6),
..Default::default()
};
let opts = PdfOptions::new()
.detection(detect_cfg)
.mode(ProcessMode::Full);
let result = pdf_inspector::process_pdf_with_options("big.pdf", opts)?;
Filtering Pages and Handling Encryption
Selective Page Extraction
When you need only specific pages, use PdfOptions::pages([...]). This populates page_filter: Option<HashSet<u32>> (see src/lib.rs lines 48-52), ensuring extraction runs only on the requested subset.
// Extract only pages 2, 3, and 4
let opts = PdfOptions::new()
.pages([2, 3, 4])
.mode(ProcessMode::Full);
let result = pdf_inspector::process_pdf_with_options("financials.pdf", opts)?;
Password-Protected PDFs
For encrypted documents, use PdfOptions::password("secret"). The implementation includes a custom Debug trait (lines 90-100 in src/lib.rs) that redacts the password from logs to prevent secret leakage.
let opts = PdfOptions::new()
.password("open-sesame")
.mode(ProcessMode::Full);
let secure = pdf_inspector::process_pdf_with_options("encrypted.pdf", opts)?;
Customizing Markdown Output
Fine-tune markdown generation via MarkdownOptions (located in src/markdown/mod.rs and consumed in src/markdown/convert.rs). Pass this to PdfOptions::markdown() to control formatting flags such as compact, raw, and pages.
use pdf_inspector::{PdfOptions, MarkdownOptions, ProcessMode};
let markdown_cfg = MarkdownOptions {
compact: true,
raw: false,
pages: true,
..Default::default()
};
let opts = PdfOptions::new()
.markdown(markdown_cfg)
.mode(ProcessMode::Full);
Complete Configuration Examples
Combine multiple options to build task-specific pipelines:
Fast classification for routing decisions:
let opts = PdfOptions::new().mode(ProcessMode::DetectOnly);
Full extraction with custom markdown on a password-protected document:
let opts = PdfOptions::new()
.password("my-pass")
.mode(ProcessMode::Full)
.markdown(MarkdownOptions::default());
High-performance analysis of massive PDFs:
let opts = PdfOptions::new()
.detection(DetectionConfig {
strategy: ScanStrategy::Sample(4),
..Default::default()
})
.mode(ProcessMode::Full);
Summary
PdfOptionsinsrc/lib.rsserves as the central configuration hub using a fluent builder API.- Processing modes (
Full,Analyze,DetectOnly) control pipeline depth without re-parsing the document. - ScanStrategy variants (
EarlyExit,Full,Sample,Pages) optimize detection speed for different document sizes and accuracy requirements. - Page filtering via
pages([...])restricts extraction to specific page numbers. - Password handling automatically redacts secrets from debug output while enabling decryption.
- MarkdownOptions in
src/markdown/mod.rsallow granular control over output formatting.
Frequently Asked Questions
How do I extract content from only specific pages?
Use the pages() method on PdfOptions with an array of page numbers. This sets the internal page_filter: Option<HashSet<u32>> field (defined in src/lib.rs lines 48-52), ensuring the extractor processes only the requested pages while still sharing the same underlying lopdf::Document instance.
What is the difference between ProcessMode::Analyze and ProcessMode::Full?
ProcessMode::Analyze runs detection plus layout analysis—including table and column detection—without generating markdown output, making it suitable for structural analysis. ProcessMode::Full executes the complete pipeline including markdown generation via the conversion engine in src/markdown/convert.rs.
How does pdf-inspector handle password-protected PDFs?
Pass the password via PdfOptions::password("your-password"). The crate decrypts the document using the provided credentials while implementing a custom Debug trait (lines 90-100 in src/lib.rs) that masks the password field to prevent accidental exposure in logs or panic traces.
Which ScanStrategy should I use for large PDFs?
For large documents where CPU time is constrained, use ScanStrategy::Sample(n) to scan n evenly spaced pages, or ScanStrategy::EarlyExit to stop at the first non-text page if you only need to identify TextBased PDFs. These strategies are defined in src/detector.rs (lines 25-40) and wrapped by DetectionConfig (lines 71-78).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →