How to Configure Firecrawl PDF‑Inspector for Specific Use Cases: Detection, Speed, and Extraction Modes
Configure Firecrawl PDF‑Inspector by adjusting DetectionConfig for page‑sampling strategies and PdfOptions for pipeline modes, OCR thresholds, page filtering, and encrypted PDF handling.
The Firecrawl PDF‑Inspector is a Rust‑based modular pipeline that separates detection, extraction, and markdown generation into configurable stages. This article shows how to tailor the library for speed‑critical batch processing, precise document classification, selective page extraction, and encrypted file handling—using the actual source structures found in firecrawl/pdf‑inspector.
Core Configuration Structures
PDF‑Inspector's flexibility centers on two main types defined in the source code:
DetectionConfig(src/detector.rs, lines 69‑78) — Controls how PDFs are inspected for text‑based vs. scanned contentPdfOptions(src/lib.rs, lines 164‑180) — Builder for end‑to‑end run settings including mode, detection config, page filters, and passwords
The pipeline's execution mode is determined by ProcessMode (src/process_mode.rs, lines 7‑15), an enum with three variants:
pub enum ProcessMode {
DetectOnly, // Fast detection only, no extraction
Full, // Detection → extraction → markdown (default)
Analyze, // Detection + layout analysis without markdown
}
Configuring Detection Behavior with DetectionConfig
The DetectionConfig struct defines three critical thresholds:
pub struct DetectionConfig {
pub strategy: ScanStrategy, // Which pages to scan
pub min_text_ops_per_page: u32, // Minimum Tj/TJ operators for "text" page
pub text_page_ratio_threshold: f32, // Text pages / total pages for TextBased
}
ScanStrategy Variants for Speed vs. Accuracy
| Strategy | Implementation | Best For |
|---|---|---|
EarlyExit |
Stops at first non‑text page | Quickly rejecting clearly scanned documents |
Full |
Scans every page | Precise Mixed/Scanned classification (legal contracts, mixed reports) |
Sample(N) |
Samples N evenly‑distributed pages | Large PDFs where speed matters |
Pages(vec) |
Scans only specified 1‑indexed pages | Known relevant pages (cover, index, appendix) |
Default values (defined in src/detector.rs, lines 80‑89):
ScanStrategy::Sample(8)min_text_ops_per_page: 3text_page_ratio_threshold: 0.6
These defaults provide balanced accuracy for typical documents without scanning every page.
Configuring Pipeline Execution with PdfOptions
PdfOptions uses a fluent builder pattern to assemble complete run configurations. The struct fields (from src/lib.rs, lines 164‑180) include:
mode: ProcessMode— Pipeline stage selectiondetection: DetectionConfig— Detection parametersmarkdown: MarkdownOptions— Output formatting (Full mode only)page_filter: Option<Vec<PageNum>>— Subset of pages to processpassword: Option<String>— Decryption password (with customDebugto prevent logging)
Builder methods are implemented in src/lib.rs at lines 30‑59.
Use Case: Default Fast Detection
For quick classification without extraction, use the detect_pdf convenience function:
use pdf_inspector::detect_pdf;
let result = detect_pdf("report.pdf")?;
println!("Detected type: {:?}", result.pdf_type);
This internally calls process_pdf_with_options with PdfOptions::detect_only(), which sets mode = ProcessMode::DetectOnly and default detection config (source: src/lib.rs, lines 73‑78).
Use Case: Speed‑Focused Sampling for Large PDFs
Reduce detection overhead by sampling fewer pages:
use pdf_inspector::{PdfOptions, DetectionConfig, ScanStrategy};
let config = DetectionConfig {
strategy: ScanStrategy::Sample(4), // Only 4 pages examined
..DetectionConfig::default()
};
let opts = PdfOptions::new().detection(config);
let result = pdf_inspector::process_pdf_with_options("huge.pdf", opts)?;
println!("Pages sampled: {}", result.pages_sampled);
Performance impact: Detection becomes O(1) regardless of total page count. The four sampled pages are distributed across the document (first, last, and intermediate pages) to improve representativeness.
Use Case: Precise Mixed/Scanned Classification
Force complete page analysis for documents requiring accurate type detection:
use pdf_inspector::{PdfOptions, DetectionConfig, ScanStrategy, ProcessMode};
let config = DetectionConfig {
strategy: ScanStrategy::Full, // Every page scanned
min_text_ops_per_page: 5, // Stricter text threshold
text_page_ratio_threshold: 0.7, // Higher bar for TextBased
..DetectionConfig::default()
};
let opts = PdfOptions::new()
.mode(ProcessMode::Full)
.detection(config);
let result = pdf_inspector::process_pdf_with_options("contract.pdf", opts)?;
println!("PDF type: {:?}", result.pdf_type);
When to use: Legal documents, academic papers with scanned figures, or reports mixing generated and raster content. The Full strategy prevents early‑exit misclassification when a single non‑text page appears early in the document.
Use Case: Selective Page Processing
Extract only specific pages without parsing the entire file:
use pdf_inspector::PdfOptions;
let opts = PdfOptions::new()
.pages([1, 10, 20]); // 1‑indexed page numbers
let result = pdf_inspector::process_pdf_with_options("book.pdf", opts)?;
println!("Extracted {} pages", result.pages.len());
The pages method accepts any IntoIterator<Item = PageNum> and propagates the filter through the extraction pipeline (src/extractor/mod.rs).
Use Case: Encrypted PDF Handling
Provide decryption passwords safely:
use pdf_inspector::{PdfOptions, ProcessMode};
let opts = PdfOptions::new()
.password("my‑pw")
.mode(ProcessMode::Full);
let result = pdf_inspector::process_pdf_with_options("encrypted.pdf", opts)?;
println!("Decrypted and extracted {} pages", result.pages.len());
Security note: The password field uses a custom Debug implementation that masks the actual value, preventing accidental exposure in logs.
Complete Configuration Example
Combine multiple options for production workflows:
use pdf_inspector::{
ProcessMode, PdfOptions, DetectionConfig, ScanStrategy, MarkdownOptions,
};
let detection_cfg = DetectionConfig {
strategy: ScanStrategy::Sample(6), // Faster for 300‑page reports
min_text_ops_per_page: 4,
text_page_ratio_threshold: 0.65,
..Default::default()
};
let markdown_cfg = MarkdownOptions::default(); // Customize in src/markdown/mod.rs
let opts = PdfOptions::new()
.mode(ProcessMode::Full)
.detection(detection_cfg)
.markdown(markdown_cfg)
.pages(1..=10); // First ten pages only
let result = pdf_inspector::process_pdf_with_options("annual_report.pdf", opts)?;
println!("Completed: {} markdown blocks", result.markdown.len());
Key Source Files Reference
| File | Purpose | Lines |
|---|---|---|
src/detector.rs |
DetectionConfig, ScanStrategy, detection algorithm |
69‑89 |
src/lib.rs |
PdfOptions builder, public API entry points |
30‑59, 73‑78, 164‑180 |
src/process_mode.rs |
ProcessMode enum definition |
7‑15 |
src/markdown/mod.rs |
MarkdownOptions and output formatting |
Module root |
src/extractor/mod.rs |
Extraction orchestration respecting page filters | Module root |
src/tables/mod.rs |
Table detection strategies | Module root |
Summary
- Use
ProcessMode::DetectOnlyfor fastest classification without content extraction - Adjust
ScanStrategyto trade speed for accuracy:Sample(N)for large files,Fullfor precise classification,Pages(vec)for selective analysis - Tune OCR thresholds via
min_text_ops_per_pageandtext_page_ratio_thresholdto match your document types - Apply
pages()andpassword()methods onPdfOptionsfor targeted processing and secure decryption - Prefer
PdfOptions::new()builder over raw struct initialization for forward compatibility
Frequently Asked Questions
What is the fastest way to classify PDF type without extracting content?
Use pdf_inspector::detect_pdf("file.pdf") or explicitly configure PdfOptions::new().mode(ProcessMode::DetectOnly). This skips extraction and markdown generation entirely, running only the detection algorithm with default DetectionConfig.
How do I process a 1000‑page PDF without scanning every page?
Set DetectionConfig { strategy: ScanStrategy::Sample(4), ..Default::default() } and pass it to PdfOptions::detection(). The detector examines only four evenly‑spaced pages regardless of total page count, making runtime effectively constant.
When should I use ProcessMode::Analyze instead of Full?
Use Analyze when you need structural information—table boundaries, column positions, reading order—without the overhead of markdown serialization. This mode runs detection plus layout analysis but skips the final markdown generation step, saving CPU for downstream custom formatters.
How do I prevent passwords from appearing in logs?
The PdfOptions struct implements a custom Debug trait that masks the password field. When you call .password("secret"), the value is stored for decryption but displays as redacted in any {:?} formatting or logging output.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →